Quant — how the whole system works
An in-house Screener.in: fundamentals for NSE + BSE companies, ingested from exchange filings with per-field provenance, reconciled into one trusted layer, and served to this UI, a documented API, and an AI formula endpoint. Numbers below are live.
What this is
Every listed company files results with the exchanges (XBRL data files + PDF statements). This system scrapes those filings, stores them immutably, parses every number exactly as filed, then reconciles them — picking one trusted value per line item with full provenance — and serves the result as company pages, a screener, an API and an AI-queryable formula engine. The design principle behind everything: raw is sacred, meaning is made once (in the reconciler), readers only read the serve layer.
Architecture — four layers
| quant-calendar-morning | 09:45 | refresh results calendar before the day's watch ticks |
| quant-watch-build/surge/tail | 10:00–22:30 | intraday: detect filed results → trigger the pipeline |
| quant-pipeline-daily | 20:30 | nightly: scrape → parse → reconcile → publish |
| quant-concalls | ~10×/day | poll Tijori → upsert serve_concall (denser 3–9pm) |
| news-poll | 07:00 | portfolio news feed (separate module) |
L0 RAW S3, immutable XBRL zips · PDFs · bhavcopies · calendar snapshots
L1 PARSED SQLite per module every fact AS FILED (xbrl_facts, pdf_parser)
L2 RECONCILED the truth-maker canonical_fact: one value per line + provenance,
versioning, NSE↔BSE verdicts, sector views
L3 SERVE Postgres (DB server) read models the UI / API / MCP query per requestNightly at 20:30 IST: cron runs the pipeline job — seed today’s results-calendar into the queue → per company: scrape NSE + BSE → parse XBRL → extract PDF → reconcile → copy new raw files to S3 → publish the serve layer to Postgres. Pages here read Postgres per request — no bulk data ever enters a serving path (the raw store is ~1.3 GB and growing; it stays with the pipeline).
Modules
| Module | Does | Status |
|---|---|---|
| entity_master | ISIN registry — identity is ISIN, never name/symbol; symbol/name history | live |
| calendars | NSE + BSE “results expected” merged on ISIN+date; seeds the nightly run | live |
| nse/bse_scraper | Result filings + attachments (PDF, XBRL) → immutable raw zone | live |
| xbrl_parser | XBRL → flat facts exactly as filed | live |
| pdf_parser | PDF → facts with confidence (text-layer + vision tiers) | extraction only |
| reconciler | Truth-maker: candidate resolution, NSE cumulative-label correction, NSE↔BSE equality, Q4 = FY − 9M, sector classification + normalized headline, restatement versioning | live |
| prices | Bhavcopy OHLCV; corp-action adjustment + event windows later | early |
| publish | After each run: SQLite → Postgres read models | live |
| quant-api | The serve API (documented, testable) | live |
| formula_mcp | AI formula DSL over reconciled keys — safe arithmetic, never SQL | live |
| concalls | Conference-call transcripts + AI summaries (Tijori); consumed straight into Postgres serve_concall, windowed poll ~10×/day + derived event tags | new |
| backfill | Full-market historical backfill | not built |
The pages
- Calendar live
- Who reports when — the merged NSE+BSE results calendar. Month grid + agenda (collapsible days, scoped to the selected month). This same calendar seeds the nightly pipeline.
- Results live
- The screener: one row per company — latest reported quarter with QoQ/YoY on revenue, EBITDA and PAT, margin deltas in bps, EPS. EBITDA = PBT + Depreciation + Finance Costs − Other Income (PBT mandatory); banks and insurers use their normalized headline operating profit. MCap fills as price + balance-sheet coverage grows.
- Filings live
- The ingestion ledger: every filing the scrapers discovered, with per-document state (XBRL / PDF fetched or not) and links to the source documents.
- Companies live
- All listed companies, searched server-side and paginated. Click through for identity, listings, history, latest price, filing counts and the as-filed statements — a tab per (exchange × scope), every number exactly as filed, with result-PDF links per column where available.
- Facts parked
- Parked. It browsed raw L1 facts, which no longer ship to the UI. Full provenance is preserved per canonical fact (candidates_json); this page returns as a provenance browser.
- Concalls partial
- Earnings-call intelligence: each call's one-line highlight, derived event tags (M&A / Capex / Buyback / …), transcript PDF and audio, plus Tijori's AI summary as collapsible sections. Server-side search + status filter, keyset-paginated. Consumed from Tijori — no LLM of ours.
- Formula MCP live
- Human UI for the AI formula engine: pick a company and period, write arithmetic over reconciled keys (net_profit/total_assets*100) and get the value plus the exact inputs bound. The same engine serves AI clients over MCP.
APIs & docs
| Surface | Base URL | Docs |
|---|---|---|
| quant-api — the data API | https://api.13-234-73-65.sslip.io | /docs (Scalar, testable) |
| formula MCP — REST | https://mcp.13-234-73-65.sslip.io | /docs (Swagger) · /llm.txt |
| formula MCP — MCP protocol | https://mcp.13-234-73-65.sslip.io/mcp | for Claude / Gemini MCP clients |
How to fetch — in three steps
- Open https://api.13-234-73-65.sslip.io/docs — every endpoint is listed and testable right in the browser (no key, no setup).
- Pick an endpoint below and just call it:
GET https://api.13-234-73-65.sslip.io/…→ plain JSON back. - Name a company by ISIN, NSE symbol, or BSE scrip — all three work wherever you see
{id}. Money is in rupees;yearis the fiscal-year end (2025 = FY 2024-25).
What to fetch
| Endpoint | What you get |
|---|---|
| /companies?q=tata | Search / list companies (paginated) |
| /company/{id} | One company — identity, listings, latest price, filing counts |
| /company/{id}/as-filed?kind=quarter | Every filed fact for that company — one tab per (exchange, scope) |
| /company/{id}/as-filed/metrics | Computed per quarter — EBITDA, margins, growth, valuation |
| /company/{id}/prices?range=1y | Daily close series + result-date markers |
| /company/{id}/concalls | That company's earnings-call transcripts + AI summaries |
| /screens/asfiled/metrics | The screener — latest quarter per company (revenue, EBITDA, PAT, margins); ?date= for "who reported today" |
| /concalls?status=recorded | Conference-call feed (newest first, keyset-paginated) |
| /calendar | Who reports when (merged NSE+BSE results calendar) |
| /filings | Every scraped filing + links to the source documents |
| /stats | Corpus counts + snapshot dates |
Example
curl "https://api.13-234-73-65.sslip.io/company/INE002A01018/as-filed/metrics?scope=Consolidated"
→ Reliance Industries' consolidated numbers per quarter — EBITDA, margins, growth, valuation — as JSON.
Data-quality guarantees
- Two-source verification. NSE↔BSE equality verified exhaustively: 99.47% of 136k comparable fact-pairs bitwise identical, zero disagreement on any core line — residual differences are exchange-template artifacts, recorded as conflicts, never silently merged.
- NSE cumulative-context correction. NSE routinely tags year-to-date values onto quarter contexts; the reconciler detects and corrects it.
- Restatements never overwrite. A changed value becomes a new version; as-first-reported history is preserved.
- “Not reported” is not zero. An all-zero filed context (a company that filed consolidated only annually) is dropped and flagged, not shown as a row of zeros.
- Validated against Screener.in line-by-line: ~80% of line×year checks match, net profit ~92%; the rest is Screener’s editorial re-presentation — a documented ceiling, not reconciler error.
What does not work yet — honestly
- PDF cross-verification (“the PDF thing”). PDF extraction works (with confidence tiers), but the PDF↔XBRL verify-state machine is not wired into the reconciler yet — canonical facts today come from XBRL only. Parked by decision; it is the next reconciler rung. Until then PDF-only fields (notes, detail schedules) are not served.
- Coverage: 4,524 companies with as-filed results of 9,029 listed — grows every drain. The reconciled layer above it is parked: frozen at 264 companies since 2026-07-21, so the UI and API read as-filed only.
- Facts page parked (returns as a provenance browser over candidates_json).
- Prices are v1 — 1,247 trading sessions ingested; no corp-action adjustment or event windows yet, so MCap in Results is mostly empty.
- Result-PDF links are sparse — only filings scraped recently carry PDF URLs; backfills as drains continue.
- Gen-insurance quarterly net profit can be null (sector mapping gap).
- A few pre-fix rows in Postgres (phantom zero-quarters for 15 companies) — already re-queued; they clean up on the next drain and the UI hides them meanwhile.
- No authentication on any surface — deliberate (“security scan once everything is built”); IAP is the planned fix.
Where things run
AWS Mumbai (ap-south-1), two servers. The app server runs the UI, the API, the formula MCP and the filing upload (Docker, behind Caddy), plus the pipeline jobs on cron — the nightly pipeline, tg-fast, the watcher, the calendar and concalls. The DB server runs Postgres 17. S3 bucket quant-raw-868967336660 holds raw + the nightly state backup. Engineering source of truth: GUIDE.md, PLAN.md and TODO.md in the repo.