Guide

Quant — how the whole system works

An in-house Screener.in: fundamentals for NSE + BSE companies, ingested from exchange filings with per-field provenance, reconciled into one trusted layer, and served to this UI, a documented API, and an AI formula endpoint. Numbers below are live.

Companies
All listed companies, searched server-side and paginated. Click through for identity, listings, history, latest price, filing counts and the as-filed statements — a tab per (exchange × scope), every number exactly as filed, with result-PDF links per column where available.
With results
4,524
NSE filings
20,368
Calendar events
4,301

What this is

Every listed company files results with the exchanges (XBRL data files + PDF statements). This system scrapes those filings, stores them immutably, parses every number exactly as filed, then reconciles them — picking one trusted value per line item with full provenance — and serves the result as company pages, a screener, an API and an AI-queryable formula engine. The design principle behind everything: raw is sacred, meaning is made once (in the reconciler), readers only read the serve layer.

Architecture — four layers

Cron jobs (app server · Asia/Kolkata)
quant-calendar-morning09:45refresh results calendar before the day's watch ticks
quant-watch-build/surge/tail10:00–22:30intraday: detect filed results → trigger the pipeline
quant-pipeline-daily20:30nightly: scrape → parse → reconcile → publish
quant-concalls~10×/daypoll Tijori → upsert serve_concall (denser 3–9pm)
news-poll07:00portfolio news feed (separate module)
Raw is sacred (S3 + SQLite, as-filed) → the reconciler makes meaning once → publish materializes read models into Postgres → every reader (UI · API · MCP) does dumb indexed SELECTs. Concalls is consume-only, so it upserts straight into Postgres, bypassing reconcile.
L0  RAW          S3, immutable         XBRL zips · PDFs · bhavcopies · calendar snapshots
L1  PARSED       SQLite per module     every fact AS FILED (xbrl_facts, pdf_parser)
L2  RECONCILED   the truth-maker       canonical_fact: one value per line + provenance,
                                       versioning, NSE↔BSE verdicts, sector views
L3  SERVE        Postgres (DB server)  read models the UI / API / MCP query per request

Nightly at 20:30 IST: cron runs the pipeline job — seed today’s results-calendar into the queue → per company: scrape NSE + BSE → parse XBRL → extract PDF → reconcile → copy new raw files to S3 → publish the serve layer to Postgres. Pages here read Postgres per request — no bulk data ever enters a serving path (the raw store is ~1.3 GB and growing; it stays with the pipeline).

Modules

ModuleDoesStatus
entity_masterISIN registry — identity is ISIN, never name/symbol; symbol/name historylive
calendarsNSE + BSE “results expected” merged on ISIN+date; seeds the nightly runlive
nse/bse_scraperResult filings + attachments (PDF, XBRL) → immutable raw zonelive
xbrl_parserXBRL → flat facts exactly as filedlive
pdf_parserPDF → facts with confidence (text-layer + vision tiers)extraction only
reconcilerTruth-maker: candidate resolution, NSE cumulative-label correction, NSE↔BSE equality, Q4 = FY − 9M, sector classification + normalized headline, restatement versioninglive
pricesBhavcopy OHLCV; corp-action adjustment + event windows laterearly
publishAfter each run: SQLite → Postgres read modelslive
quant-apiThe serve API (documented, testable)live
formula_mcpAI formula DSL over reconciled keys — safe arithmetic, never SQLlive
concallsConference-call transcripts + AI summaries (Tijori); consumed straight into Postgres serve_concall, windowed poll ~10×/day + derived event tagsnew
backfillFull-market historical backfillnot built

The pages

Calendar live
Who reports when — the merged NSE+BSE results calendar. Month grid + agenda (collapsible days, scoped to the selected month). This same calendar seeds the nightly pipeline.
Results live
The screener: one row per company — latest reported quarter with QoQ/YoY on revenue, EBITDA and PAT, margin deltas in bps, EPS. EBITDA = PBT + Depreciation + Finance Costs − Other Income (PBT mandatory); banks and insurers use their normalized headline operating profit. MCap fills as price + balance-sheet coverage grows.
Filings live
The ingestion ledger: every filing the scrapers discovered, with per-document state (XBRL / PDF fetched or not) and links to the source documents.
Companies live
All listed companies, searched server-side and paginated. Click through for identity, listings, history, latest price, filing counts and the as-filed statements — a tab per (exchange × scope), every number exactly as filed, with result-PDF links per column where available.
Facts parked
Parked. It browsed raw L1 facts, which no longer ship to the UI. Full provenance is preserved per canonical fact (candidates_json); this page returns as a provenance browser.
Concalls partial
Earnings-call intelligence: each call's one-line highlight, derived event tags (M&A / Capex / Buyback / …), transcript PDF and audio, plus Tijori's AI summary as collapsible sections. Server-side search + status filter, keyset-paginated. Consumed from Tijori — no LLM of ours.
Formula MCP live
Human UI for the AI formula engine: pick a company and period, write arithmetic over reconciled keys (net_profit/total_assets*100) and get the value plus the exact inputs bound. The same engine serves AI clients over MCP.

APIs & docs

SurfaceBase URLDocs
quant-api — the data APIhttps://api.13-234-73-65.sslip.io/docs (Scalar, testable)
formula MCP — RESThttps://mcp.13-234-73-65.sslip.io/docs (Swagger) · /llm.txt
formula MCP — MCP protocolhttps://mcp.13-234-73-65.sslip.io/mcpfor Claude / Gemini MCP clients

How to fetch — in three steps

  1. Open https://api.13-234-73-65.sslip.io/docs — every endpoint is listed and testable right in the browser (no key, no setup).
  2. Pick an endpoint below and just call it: GET https://api.13-234-73-65.sslip.io/… → plain JSON back.
  3. Name a company by ISIN, NSE symbol, or BSE scrip — all three work wherever you see {id}. Money is in rupees; year is the fiscal-year end (2025 = FY 2024-25).

What to fetch

EndpointWhat you get
/companies?q=tataSearch / list companies (paginated)
/company/{id}One company — identity, listings, latest price, filing counts
/company/{id}/as-filed?kind=quarterEvery filed fact for that company — one tab per (exchange, scope)
/company/{id}/as-filed/metricsComputed per quarter — EBITDA, margins, growth, valuation
/company/{id}/prices?range=1yDaily close series + result-date markers
/company/{id}/concallsThat company's earnings-call transcripts + AI summaries
/screens/asfiled/metricsThe screener — latest quarter per company (revenue, EBITDA, PAT, margins); ?date= for "who reported today"
/concalls?status=recordedConference-call feed (newest first, keyset-paginated)
/calendarWho reports when (merged NSE+BSE results calendar)
/filingsEvery scraped filing + links to the source documents
/statsCorpus counts + snapshot dates

Example

curl "https://api.13-234-73-65.sslip.io/company/INE002A01018/as-filed/metrics?scope=Consolidated"

→ Reliance Industries' consolidated numbers per quarter — EBITDA, margins, growth, valuation — as JSON.

Data-quality guarantees

  • Two-source verification. NSE↔BSE equality verified exhaustively: 99.47% of 136k comparable fact-pairs bitwise identical, zero disagreement on any core line — residual differences are exchange-template artifacts, recorded as conflicts, never silently merged.
  • NSE cumulative-context correction. NSE routinely tags year-to-date values onto quarter contexts; the reconciler detects and corrects it.
  • Restatements never overwrite. A changed value becomes a new version; as-first-reported history is preserved.
  • “Not reported” is not zero. An all-zero filed context (a company that filed consolidated only annually) is dropped and flagged, not shown as a row of zeros.
  • Validated against Screener.in line-by-line: ~80% of line×year checks match, net profit ~92%; the rest is Screener’s editorial re-presentation — a documented ceiling, not reconciler error.

What does not work yet — honestly

  1. PDF cross-verification (“the PDF thing”). PDF extraction works (with confidence tiers), but the PDF↔XBRL verify-state machine is not wired into the reconciler yet — canonical facts today come from XBRL only. Parked by decision; it is the next reconciler rung. Until then PDF-only fields (notes, detail schedules) are not served.
  2. Coverage: 4,524 companies with as-filed results of 9,029 listed — grows every drain. The reconciled layer above it is parked: frozen at 264 companies since 2026-07-21, so the UI and API read as-filed only.
  3. Facts page parked (returns as a provenance browser over candidates_json).
  4. Prices are v1 — 1,247 trading sessions ingested; no corp-action adjustment or event windows yet, so MCap in Results is mostly empty.
  5. Result-PDF links are sparse — only filings scraped recently carry PDF URLs; backfills as drains continue.
  6. Gen-insurance quarterly net profit can be null (sector mapping gap).
  7. A few pre-fix rows in Postgres (phantom zero-quarters for 15 companies) — already re-queued; they clean up on the next drain and the UI hides them meanwhile.
  8. No authentication on any surface — deliberate (“security scan once everything is built”); IAP is the planned fix.

Where things run

AWS Mumbai (ap-south-1), two servers. The app server runs the UI, the API, the formula MCP and the filing upload (Docker, behind Caddy), plus the pipeline jobs on cron — the nightly pipeline, tg-fast, the watcher, the calendar and concalls. The DB server runs Postgres 17. S3 bucket quant-raw-868967336660 holds raw + the nightly state backup. Engineering source of truth: GUIDE.md, PLAN.md and TODO.md in the repo.