straw/rust/strawcore/src/runtime.rs
Cobb e4a8751942 reliability + security fix sprint (vc=91)
Straw full-audit fix batch (2026-07-04). Confirmed HIGH/MED findings:

- Updater self-brick: drop the fdroid.sulkta.com leaf+E7 SPKI pins (LE
  rotates the leaf every ~90d + intermediates from a pool → a routine
  renewal was guaranteed to miss both pins and silently kill the only
  in-app update path). System-CA validation + the install-time APK
  signature gate remain. Also distinguish "index unreachable" from "up to
  date", add a 30s callTimeout + bounded index read, guard the API-26
  NotificationChannel, and setOnlyAlertOnce.
- Request POST_NOTIFICATIONS at runtime so A13+ update/media notifications
  actually post (was declared but never requested → silent no-op).
- isDebuggable=false on the shipped debug variant (closes ADB run-as dump
  of watch/search history + subs; no package-id cutover).
- Crash fixes: distinctBy(url) on moreFromChannel + related (duplicate
  LazyColumn key); back-handling driven by live nav depth so it survives
  moveTaskToBack on A12+; PlaylistsStore per-store caps + off-main
  hydration (hostile import ANR); SponsorBlock staleness fence via
  NowPlaying.refreshIfCurrent.
- CacheCap.nearest() snaps up to the nearest finite cap (defaults no longer
  resolve to Unlimited/100k on fresh installs).
- RYD/SB FFI shims wrap the uniffi call (swallow a core panic per contract).
- Rust wrapper: release panic="unwind" (was "abort", which defeated UniFFI
  catch_unwind → any core panic = whole-app SIGABRT) + a 60s wall-clock
  timeout on every blocking extractor call (unblocks the caller, bounds
  thread pileup). Pairs with the strawcore-core reliability fix.
2026-07-04 07:09:54 -07:00

142 lines
6.2 KiB
Rust

// Runtime bootstrap. Called once from Kotlin's StrawApp.onCreate via
// init_logging(), and again before every strawcore call. Wires the
// strawcore-core Downloader + Localization singleton so the extractor
// has an HTTP client to use.
//
// the prior shape used `Once::call_once` and
// silently swallowed errors. If the FIRST call ran while the network
// stack wasn't ready (cold boot in airplane mode, SELinux denial on
// first TLS init, transient resolver failure), the Once slot was
// consumed, NewPipe::init_full never ran, and every subsequent
// search/streamInfo/channelInfo returned DownloaderMissing for the
// rest of the process lifetime.
//
// New shape: use an AtomicBool to track success. Only "success" closes
// the door. On failure we retry — rate-limited so a persistently-broken
// network doesn't hammer reqwest::Client::new() on every call.
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::sync::{Arc, Mutex};
use std::time::{Duration, SystemTime, UNIX_EPOCH};
use strawcore_core::downloader::ReqwestDownloader;
use strawcore_core::localization::{ContentCountry, Localization};
use strawcore_core::newpipe::NewPipe;
use crate::error::StrawcoreError;
static INITIALIZED: AtomicBool = AtomicBool::new(false);
static LAST_ATTEMPT_MS: AtomicU64 = AtomicU64::new(0);
// Guards the actual init attempt so concurrent calls don't all try
// to build the downloader in parallel; serial retry is the goal.
static INIT_LOCK: Mutex<()> = Mutex::new(());
/// Min ms between retries when init has failed. 5s — enough that a
/// hot loop of failed searches doesn't pin a CPU on reqwest setup,
/// short enough that a user who toggled airplane mode off recovers
/// within one tap.
const RETRY_BACKOFF_MS: u64 = 5_000;
fn now_ms() -> u64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0)
}
pub fn ensure_initialized() {
// Fast path: already initialized. Single Acquire load.
if INITIALIZED.load(Ordering::Acquire) {
return;
}
// Backoff check BEFORE the lock — a recent failure shouldn't
// make N concurrent callers queue on a mutex they'll all skip
// out of anyway.
let last = LAST_ATTEMPT_MS.load(Ordering::Acquire);
let now = now_ms();
if last != 0 && now.saturating_sub(last) < RETRY_BACKOFF_MS {
return;
}
// try_lock — if another thread is already mid-init, return
// immediately rather than block. The caller will get
// DownloaderMissing once from the extractor and recover on
// the next user action; the alternative (blocking N tokio
// workers for the full duration of a slow init) freezes the
// UI. was the regression on round-5's
// mutex-first ordering.
let _guard = match INIT_LOCK.try_lock() {
Ok(g) => g,
Err(_) => return,
};
// Re-check under the lock — another thread may have just succeeded.
if INITIALIZED.load(Ordering::Acquire) {
return;
}
match ReqwestDownloader::new() {
Ok(dl) => {
NewPipe::init_full(
Arc::new(dl),
Localization::default(),
ContentCountry::default(),
);
INITIALIZED.store(true, Ordering::Release);
// Clear LAST_ATTEMPT_MS so a future hypothetical
// re-init path (none today) wouldn't see cooldown
// bleed from this success.
LAST_ATTEMPT_MS.store(0, Ordering::Release);
log::info!("strawcore-core: downloader + localization initialized");
}
Err(e) => {
// Stamp the timestamp on FAILURE only, so the next
// caller within RETRY_BACKOFF_MS skips, but a successful
// attempt isn't reflected in the backoff state.
LAST_ATTEMPT_MS.store(now, Ordering::Release);
log::error!("strawcore-core: downloader init failed (will retry on next call)");
let _ = e;
}
}
}
/// Wall-clock ceiling on a single blocking extractor call. Core bounds each
/// HTTP request to 30s and (since the 2026-07-04 reliability fix) the JS
/// deobfuscation runtime to a 5s interrupt, so a legitimate multi-round-trip
/// extraction on a slow connection stays comfortably under this; anything
/// past it is a wedge (a pathological player.js, a black-holed CDN) and we'd
/// rather surface a timeout to the Kotlin caller than spin forever.
const EXTRACT_TIMEOUT: Duration = Duration::from_secs(60);
/// Run a blocking strawcore-core call on the tokio blocking pool, bounded by
/// [`EXTRACT_TIMEOUT`]. Replaces the raw `spawn_blocking(..).await??` pattern
/// at every FFI entry point so a hung extraction unblocks the Kotlin suspend
/// instead of (a) spinning the spinner forever and (b) piling detached
/// blocking threads toward tokio's 512-thread cap, after which *all*
/// spawn_blocking-based FFI calls would queue behind the wedged ones for the
/// process lifetime.
///
/// Caveat (documented, not a bug): `spawn_blocking` tasks are not abortable,
/// so on timeout the detached thread keeps running until the core call
/// returns on its own — core's per-request (30s) + JS-interrupt (5s) bounds
/// cap how long that is. The win here is unblocking the *caller* and bounding
/// retry pileup, not reclaiming the thread instantly.
pub(crate) async fn run_extract<T, E, F>(what: &'static str, f: F) -> Result<T, StrawcoreError>
where
F: FnOnce() -> Result<T, E> + Send + 'static,
T: Send + 'static,
E: Send + 'static,
StrawcoreError: From<E>,
{
match tokio::time::timeout(EXTRACT_TIMEOUT, tokio::task::spawn_blocking(f)).await {
// Timed out waiting on the blocking task.
Err(_elapsed) => Err(StrawcoreError::Extractor {
msg: format!("timeout: {what} exceeded {}s", EXTRACT_TIMEOUT.as_secs()),
}),
// Blocking task finished within the deadline but the join failed —
// panic (only under a non-abort profile) or cancellation.
Ok(Err(join)) => Err(StrawcoreError::Extractor {
msg: format!("join: {join}"),
}),
// Blocking task returned; propagate its own error via the existing
// `From<ExtractionError>` mapping.
Ok(Ok(inner)) => inner.map_err(StrawcoreError::from),
}
}