DriftBot AI 2.0: Ready to analyze 10-year disclosures across 50+ filers.
⚑ Run Radar Scan πŸ” Audit Outliers πŸ“Š SQL Benchmark
10-Year Document Forensics (2016–2025) β€’ 50+ Leading Entities

Detect when documents
change meaning over time

A general-purpose semantic drift detection engine for large versioned document corpora. Tracking high-dimensional cluster trajectories, Wasserstein distributional shifts, and evidence-grounded materiality across 10 years of SEC 10-K disclosures.

SEC EDGAR REPOSITORY β€’ 10-YEAR FORENSIC PIPELINE
ACTIVE TELEMETRY
STAGE 01 β€’ INGESTION
SEC Item 1A / Item 7
500+ annual filings parsed across 2016–2025. Immutable bronze cache with SHA verification.
data.sec.gov β€’ 8 req/s
STAGE 02 β€’ EMBEDDING
Neural Latent Space
384-dimensional dense sentence embeddings using local bge-small-en-v1.5 with L2 normalization.
UMAP + HDBSCAN Manifold
STAGE 03 β€’ FORENSICS
Wasserstein Drift Engine
Distributional transport shift with 60-iteration permutation tests & DuckDB-Wasm in-browser queries.
$0 Server Cost β€’ Client Wasm
⚑ AI & CHIP FOUNDRY DEPENDENCY
NVIDIA (NVDA) β€’ FY2025
Advanced CoWoS packaging wafer concentration & HBM3e memory stack bottlenecks.
πŸ›‘οΈ KERNEL & SENSOR RESILIENCY
CrowdStrike (CRWD) β€’ FY2024
Kernel-level sensor driver update validation & global multi-tenant outage exposure.
🌐 CLOUD PRIVACY & SAFETY
Apple (AAPL) β€’ FY2024
Private cloud compute nodes & foundation model safety guardrail liabilities.

Interactive Latent Space Trajectory (2016–2025)

Real-time UMAP projection of sentence embeddings across 10 fiscal disclosure years. Drag the slider to observe multi-modal cluster drift.

10-Year Scrub: 2025
AI Infrastructure & Model Safety
Global Semiconductor Foundries & CoWoS
Cloud Privacy, Sovereignty & Zero Trust
Export Controls & Trade Sanctions
Workforce & Supply Chain Fragility
500+
10-K Filings Processed
50+
Corpus Entities (11 Sectors)
10 Years
Temporal Horizon (2016–2025)
$0.00
Monthly Runtime Cost

Forensic Architecture Pipeline

Interactive view of our modular document-to-drift analysis pipeline. Click any stage for engineering details.

STAGE 01
SEC EDGAR Ingestion
Rate-limited (8 req/s) retrieval with immutable Bronze checkpoints across 10 years (2016-2025).
STAGE 02
Format-Aware Parsing
Handles both legacy plain HTML (2016-2018) and modern iXBRL tags with TOC avoidance.
STAGE 03
Vector Embeddings
Local bge-small-en-v1.5 representations with L2 normalization and disk caching.
STAGE 04
UMAP + Outlier Discovery
Global density clustering with sanitized prose filters to isolate genuine black-swan risks.
STAGE 05
Wasserstein Drift Tests
Multi-modal distribution distance and vectorized permutation tests (p-values).
STAGE 06
Evidence Grounding
Local LLM synthesis citing exact source chunks; validated into Gold Parquet tables.
STAGE 01: SEC EDGAR Ingestion (2016–2025)
Enforces an 8 req/s monotonic rate limit to respect data.sec.gov rules. Pulls raw 10-K primary documents into immutable bronze folders with SHA-verified metadata sidecars across 10 fiscal years.
Cross-Corpus Semantic Topology & Systemic Contagion

Interactive Knowledge Graph Network

Discover hidden cross-corporate dependencies and systemic contagion across 10 years of SEC disclosures. Click or hover any node to inspect connected filers, drift intensities, and verbatim excerpt diffs.

Cluster Lens:
Corporate Filers (e.g. NVDA, AAPL)
AI & Compute
Semiconductor Foundries
Infrastructure Resiliency
Export Bans & Sanctions
Drag to reposition β€’ Click to inspect node
THEMATIC RISK MANIFOLD

AI Infrastructure & Compute

6 Connected Filers
πŸ’‘ What This Connection Means:

Hyperscalers and device makers are simultaneously disclosing severe dependencies on frontier foundation models and custom accelerator clusters, creating an industry-wide compute chokepoint.

Connected Entities & Risk Vectors:
πŸ•ΈοΈ

1. Spot Systemic Contagion

Filers sharing the same risk node face correlated operational vulnerabilities (e.g. NVIDIA, Apple, and Broadcom sharing Taiwan packaging dependencies).

🎯

2. Identify Critical Risk Hubs

Larger nodes represent high-centrality disclosure manifolds that bridge multiple distinct industry sectors.

πŸ”¬

3. Jump to Verbatim Evidence

Click any node or link in the inspector to directly open the side-by-side Before/After 10-K diffs grounding that connection.

Top Disclosed Drift Events (2016–2025)

Ranked by Materiality Score with Wasserstein distribution divergence. Click any row for forensic before/after verification.

Entity Theme Cluster Year Classification Ξ” Intensity Centroid Drift Wasserstein (p-val) Materiality Score

Company 10-Year Trajectory

Longitudinal evolution of thematic risk disclosures (2016–2025).

DuckDB-Wasm Interactive Sandbox

Execute real SQL queries directly on 10-year analytical Parquet files in your browser. Zero backend, zero server latency.

Example Queries:
System Philosophy & Technical Architecture

About DriftLens

A general-purpose semantic drift detection engine designed to detect when the meaning of recurring sections in versioned document corpora changes over time. Demonstrated on 10 years of SEC 10-K filings across 50+ leading companies.

πŸ”

What DriftLens Is

A document intelligence system combining bi-encoder embeddings, unsupervised manifold clustering, Wasserstein distribution distance, and local LLM evidence grounding to identify substantive corporate disclosure changes.

πŸ›‘οΈ

What DriftLens Is NOT

It is not a stock picker, trading bot, or generic RAG demo. SEC filings were chosen because they are legally standardized, versioned annually, and freely accessible β€” serving as the ideal testbed for semantic drift.

⚑

Permanent $0/Month Architecture

All machine learning and LLM inference runs offline once during batch ETL. The live web application executes entirely client-side via DuckDB-Wasm reading compressed static Parquet tables. Zero server bills forever.

End-to-End Latent Vector Forensics Architecture

How raw SEC Form 10-K disclosures are ingested, parsed into semantic representations, clustered into manifold themes, and served client-side with zero server cost.

01 SEC Ingestion β€’ 8 req/s EDGAR rate limit β€’ 10 Years (2016-2025) β€’ Immutable Bronze cache β€’ SHA-256 sidecars data/bronze/*.html 02 Parser & AST β€’ Item 1A / 7 Extractor β€’ TOC link avoidance β€’ Dehyphenation & cleaning β€’ Sentence chunking data/silver/*.parquet 03 NLP & Manifolds β€’ bge-small-en-v1.5 β€’ UMAP Dimensionality β€’ HDBSCAN Cluster Lenses β€’ Centroid Vectors embeddings.npz 04 Drift Engine β€’ Wasserstein Distance β€’ Cosine Centroid Drift β€’ Calibrated Materiality β€’ LLM Evidence Diff data/gold/*.parquet 05 Client Wasm β€’ DuckDB-Wasm β€’ Browser memory β€’ Interactive SQL β€’ $0 Server cost GitHub Pages

Mathematical & Algorithmic Formulation

1. Calibrated Materiality Score:

Materiality = |Intensity[t] - Intensity[t-1]| Γ— ln(1 + Paragraph_Count[t])

Weights the magnitude of percentage shift by the depth of actual paragraph discussion, ensuring isolated 1-sentence tweaks do not outrank multi-page risk shifts.

2. Multi-Modal Wasserstein Distribution Distance:

Energy_Distance(P, Q) = 2 Β· E[||X - Y||] - E[||X - X'||] - E[||Y - Y'||]

Measures the true geometric transport distance between year N and year N-1 disclosure embedding clouds, verified via a 60-iteration permutation significance test (p < 0.05).

3. What You Can Do With This Engine:

  • Audit Longitudinal Risk Trajectories: Track how a company's cyber or supply chain stance shifted across 10 years.
  • Detect Novel Black Swans: Flag unprecedented outlier paragraphs before they become industry-wide clusters.
  • Run In-Browser SQL Forensics: Execute ad-hoc analytical queries across all gold tables with DuckDB-Wasm.
  • Verify Primary Evidence: Inspect side-by-side excerpt diffs grounding every synthesized claim in verbatim text.