How a verdict is reached
engine 0.128.0 · data 2026.09.18-r1-PROVISIONAL
This page describes what actually happens when you run a check — the code paths, the data behind them, and the places where the answer is deliberately weaker than you might want it to be. It is written to be checked rather than believed. Where the tool cannot establish something, that is said here in the same words it is said on a result.
01
Where a verdict comes from
Every verdict, classification, limit comparison and citation in this product is computed by a deterministic rules engine — a self-contained TypeScript package that takes two things and returns one:
(your formulation, a snapshot of the schedules) → verdict JSON
The engine performs no network calls, reads no clock, uses no randomness and has no access to a language model. It cannot reach the database, so it cannot be affected by anything other than the snapshot it was handed. Given the same formulation and the same snapshot it returns the same bytes — which is what makes a verdict from six months ago reproducible rather than merely archived.
No language model is in the verdict path, at any point, in either module. Not to decide a colour, not to pick a schedule entry, not to compare a dose against a limit, and not to write the sentence a finding is explained in — those sentences are produced by the engine itself.
The two rule packs are separate because the questions are. The FSSAI pack answers whether a substance, at a stated daily dose, in an orally-consumed format, sits inside Schedules I–IV of the 2022 nutraceutical regulations. The CDSCO pack answers which drug schedule — G, H, H1, X or K — an active falls under, under the Drugs Rules, 1945. The drug pack has no “compliant” answer at all: an active is scheduled, restricted, or unresolved, and a tool that printed a clearance no instrument grants would be inventing one.
02
What a language model touches
A model is used in exactly four places, none of which is a verdict. Each is listed here with what it is not allowed to do, because the boundary is the interesting part.
- Reading a carton. Print-ready artwork often carries no text layer, so a model transcribes pixels into text. It returns one flat block of transcribed text and nothing else. It does not decide which label element a line belongs to — that split is done deterministically, using only phrases the rules themselves declare — and it is never asked whether anything complies.
- Pointing at a place on the page. When artwork has its type outlined as vector curves, there is no text object whose coordinates could be read. A model estimates a rectangle so a finding can be circled on the printout. The arrow carries a coordinate, not a judgement: which findings exist, what colour each one is, what it cites and how it is numbered were all settled by the engine before the model was called. Remove this step and every finding still appears, cited and numbered — in the margin instead of on the artwork.
- Drafting a proposed data amendment, in the admin tools. When the daily sweep detects a change on an official source, a model reads the excerpt and proposes values for at most five fields. It is instructed that an empty answer is a correct and common one. Nothing it proposes reaches live data: a person reviews every field against the same excerpt, and only their sign-off applies it.
- Purpose tags used for searching. The therapeutic purpose tags that let you search the catalogue by intent were proposed by a model and are unreviewed. The screen that uses them says so, and says the corollary out loud: an ingredient absent from a purpose search may simply be untagged rather than unsuitable. They order a search. They do not enter a verdict.
Your formulation is a trade secret and is treated as one: it is never written to logs or to analytics, and nothing about it is attached to an error report. Three things are worth stating precisely rather than reassuringly. A check run while you are signed in is saved to your own account, so you can reopen the report — an anonymous check is not saved in our database, apart from any ingredient name looked up in the registries, which is kept as described below. Artwork you upload, or a formulation document you upload on /invent or /check, is stored against your own account under a random name for the length of the read and deleted afterwards, and a picture of each page is sent to Anthropic’s API to be transcribed and to locate the words — nowhere else. And a name is looked up outside this product only after the engine has run, and only for a line the engine left unresolved whose name is not, and is not close to, the name of any schedule entry we hold — so not a name for which we suggested schedule entries that you did not confirm, and not a substance we matched to an entry but could not measure against its limit. What goes is only ever a name — the one on that line, or a component of a trade name we hold — never your doses and never the formulation, and it goes to the public chemical registries PubChem and the FDA’s GSRS to work out what the substance is. The name as you typed it is stored in our database with the answer, including an answer that found nothing, and kept with no expiry, so it is not sent again; a lookup cut short — a registry not answering, or the check reaching its time limit — is not kept, so that name can be sent again next time. That lookup never changes a verdict’s colour and writes nothing to any schedule, but it can change what the line says: when the registries name the substance, and neither the name on the line nor any name they give it matches a schedule entry we hold, the line reports it as absent from every Indian schedule checked — a conclusion drawn from our own entries, not from anything the registries know about Indian law.
One thing does leave by email, and it is worth being exact about because email is the only channel here that is addressed. When a regulation you have already been judged against changes, we send a notification — and that notification says nothing about your work. Not what changed, not which products, not how many. It is the same sentence for every recipient, because a mail naming a substance would tell anyone who saw it that your formulation contains it, which is the thing this page promises does not happen. What changed, and what it touched, stays behind your sign-in.
03
Where the data comes from
The FSSAI side is digitised from the official 2022 gazette PDF — Food Safety and Standards (Health Supplements, Nutraceuticals, Food for Special Dietary Use, Food for Special Medical Purpose and Prebiotic and Probiotic Food) Regulations, 2022 — one row per schedule entry, each carrying the page it was read from. The CDSCO side is digitised the same way from the Drugs Rules, 1945 schedules.
- 832
- FSSAI ingredient rows, Schedules I–IV
- every one page-cited to the gazette
- 646
- CDSCO drug actives, Schedules G · H · H1 · X
- plus Schedule K, which exempts rather than classifies
- 831
- rows carrying a human sign-off
- most accepted the machine extraction in bulk rather than opening the page
- 54
- rows quarantined as uncertain parses
- they carry no numeric limit and cannot answer green
The extraction is machine reading, and a sign-off on a row is not the same as somebody opening the page it cites. The CDSCO drug table has had no review pass at all. That is what the preview banner above every screen means, and what the PROVISIONAL suffix on the data version records. Closing it is the largest outstanding task in this project.
What is established mechanically, on every row: that the stored name appears on the page the row cites; that the stored dose appears on that page in a dose context; that the cited page exists in the source PDF; and four structural invariants covering the status field — a permitted row must cite a schedule, a quarantined row must not carry a limit, a not-permitted row must not cite one, and no row may lack a citation clause. At the last recorded run, every one of those checks passed on every row it applied to. That is a recorded result rather than a live figure, which is why it is not shown as a count above.
What no machine can establish is whether the row is the right row. A name appearing on the cited page does not prove the entry number is not off by one, that the species is not a neighbouring line, or that a limit does not belong to the row above. That is human work, and until it lands the honest reading of any result here is “this is what the machine read on that page”, not “this is what the gazette says”.
The structure those rows follow is the 2022 gazette’s own, which renumbered the 2016 schedules. As the instrument prints it: Schedule I nutrients — vitamins, minerals, amino acids and nucleotides, capped at one times the ICMR-NIN recommended daily allowance; Schedule II plants and botanicals, 439 entries carrying per-day ranges; Schedule III Part A molecules, isolates and extracts with stated numeric limits, 41 entries; Part B 192 further entries permitted without any stated numeric limit; Schedule IV prebiotics and probiotics. Melatonin, to name the one most often assumed otherwise, is permitted here — Schedule III Part A, entry 25, at 2–10 mg per day — and is not drug-classified.
These schedules are central and uniform across India. No state maintains a different permitted list. What varies by state is licensing, enforcement posture, and any No Objection Certificate for export-oriented manufacture — which is a different question, with its own directory of state authorities.
Sources are also filed with their legal status, because a draft or superseded instrument must never quietly carry a binding conclusion. An instrument is recorded as in force, as not able to carry a conclusion on its own — a draft, a proposal, something superseded or repealed, or guidance that describes how a regulator reads the law rather than imposing an obligation — or as status not recorded, which is neither of the other two and is today the common case. An unrecorded status answers normally and discloses itself, rather than turning the whole product grey on the strength of a gap in our own bookkeeping.
Of 16 registered instruments, 4 are recorded as in force, 0 as unable to carry a conclusion alone, and 12 have no status recorded.
Those unrecorded ones include the FSSAI 2022 nutraceutical regulations, which every nutraceutical verdict rests on. It is a gap in our filing rather than a doubt about the instrument. A verdict now carries this on its own face too — the same note you see below appears beside the citation itself on /report/[id] and /check — so a nutraceutical report will currently show it against its central citation, for the same reason this page does.
04
Why every verdict pins two versions
A verdict is stamped with both a db_version and an engine_version, and the pair is the whole traceability claim.
- db_version · 2026.09.18-r1-PROVISIONAL
- Which snapshot of the schedules answered. Formatted
YYYY.MM.DD-rN, and it moves on every data change a person has signed off. A live check today runs against the version shown here. - engine_version · 0.128.0
- Which rules ran. Semantic versioning, bumped on any change to rule logic — including changes that move no verdict at all, because “surely this one is harmless” is precisely the judgement that once let twelve commits of rule changes ship under a version that did not produce them. The engine’s own version file records, for each bump, which verdicts change and why.
An old verdict is never silently recomputed. It replays under the stamps it was issued with, so a report you filed months ago still says what it said, and the two numbers on it tell you exactly what would have to be re-run to see whether it would say the same thing today.
05
Absence is never permission
The schedules are positive lists. A substance that is not on one has not been permitted by omission — and it has not been prohibited by omission either. Every default in this system is built on that asymmetry, and all of them fail toward “a person must look at this”.
| FSSAI | CDSCO | What it asserts |
|---|---|---|
| Compliant | not issued | Every line was resolved and every line passed. The drug module has no equivalent — it cannot issue a clearance. |
| Conditional | Scheduled | A real finding that does not forbid the formulation: a dose over a stated cap, a condition unmet, or an active that is scheduled and carries obligations. |
| Not permitted | Restricted | The substance may not be used as formulated, or the active is restricted in who may sell it and on what prescription. |
| Verify manually | Verify manually | The tool did not reach a conclusion. Nothing about the substance is being asserted in either direction. |
On the drug side, no match is grey — never “over the counter”. An active that matches no schedule entry after every spelling, salt form, ester form, dosage-form and grade expansion the resolver knows returns “verify manually”. A wrong unscheduled call on a prescription drug is a licensing offence for the manufacturer, so absence of evidence is reported as absence of evidence. A grey line also states which schedules were searched and how many entries were in them, so a well-founded blank can be told apart from a lazy one.
On the FSSAI side, one unchecked line stops the summary saying “compliant”. Grey does not outrank a real finding — not knowing is not a breach, and letting unresolved names drown out an over-limit dose would bury the thing you needed to see. But a formulation with nothing found wrong and one line the engine could not resolve reports grey, not green, and states how many lines could not be checked. Compliance was not established; it was merely not contradicted.
A substance genuinely absent from every schedule is a finding, not an empty result. The tool distinguishes “we know exactly what this is and it is on no Indian schedule” from “we could not work out what this is”, because you can act on the first and cannot act on the second. The first names the route — a novel food ingredient application to FSSAI, and its outcome, before use. Silence is never offered as an answer.
One caveat about that exhaustiveness claim, stated because it is easy to over-read: when you run an FSSAI check, the schedules searched are the FSSAI ones. The drug schedules are consulted by the drug module, on its own screen. A “not present in the schedules” line on an FSSAI result is not a statement that the substance is absent from Schedule H or X.
06
How a dose limit is actually read
Schedule I, Schedule III Part A and Schedule IV state a single numeric ceiling per entry, and the comparison is arithmetic. Schedule II is different, and it is where most of the care went.
A Schedule II entry states its allowance per plant part and per preparation. Allium cepa permits 20–40 g of leaf as fresh, 10–20 ml of bulb juice, and 1–3 g of seed powder. Holding one number for “onion” is wrong in both directions — it tells someone using the leaf they may have 3 g where the schedule allows 40, and it would pass 20 g of seed powder if the stored number happened to be the leaf’s. So every printed range line is held separately, each with its own page citation.
When you do not say which part or preparation you are using:
- Inside the strictest printed range → permitted. Safe on every reading — there is no part you could have meant that would make it a breach.
- Above the strictest but within the most permissive → a finding, naming both. The answer turns on a fact the tool was not told, and picking the permissive bound silently is exactly the false green this design exists to prevent.
- Above every printed range → a finding against the most permissive one.
- Entered in a unit no printed range shares a dimension with → grey. An entry stating “5–10 drops” cannot be answered in milligrams, and the schedule gives no basis for converting. It says so rather than passing at any dose.
Naming the part narrows the comparison — but only to a line the entry actually prints. A part the schedule does not print for that plant falls back to the strictest range rather than matching nothing, so a vague or mistaken answer can never widen a limit. If the entry states a limit in a unit that could not be compared at all, that is reported too: quietly ignoring a printed limit and answering from the rest would be the most dangerous thing here.
A dose under a printed range is reported as below it, and grades exactly as being inside it does — green, no obligation. A schedule range is an allowance, not a minimum, and a tool that graded under-dosing as a breach would be telling a formulator to increase a botanical dose on invented authority.
Schedule II Notes 4 and 5 scale the permitted range for children — half the adult range for ages 5–16, a quarter for ages 2–5 — and both bounds are scaled, not just the ceiling. Duplicate lines for the same substance are summed before comparison rather than compared one at a time.
07
When the engine refuses to answer
Some failures cannot be degraded gracefully, and the design says so rather than answering anyway.
If the regulation database cannot be reached, the check does not run. There is a small bundled test fixture the app falls back to so that pages still render, and that fixture disagrees with the real data — among other things it treats melatonin at 10 mg/day as drug-classified, which the 2022 gazette does not. So every path that asserts a verdict checks whether it is holding fixture data and refuses. A regulatory answer computed on test rows is not a degraded answer; it is the wrong answer.
A geometric requirement is never judged from text. On artwork review, requirements about how large something must be printed come back grey with an instruction to measure the print-ready file — unless a real measurement of that file has been supplied, in which case a separate measured line is produced. A measured height is carried as an interval, not a point, because a height taken off rendered pixels is quantised and blurred; the rule asserts only when the whole interval lands on one side of the threshold. Straddling it returns grey and says measure it.
No finding is issued without a citation. Every non-grey line carries the instrument, the clause and the source page it rests on. Where a useful check exists but no Indian instrument states the requirement — barcode check-digit validation is the clearest example — the check runs and reports, and is deliberately not wired into any verdict, because a finding with nothing to cite is an opinion wearing a colour.
There is also a stricter setting than the one that ships. A food/drug dose corridor gate can be switched on, under which a substance appearing in both the FSSAI schedules and the drug schedules cannot appear in a green verdict without a verified corridor. It is off by default, and honestly so: no primary instrument in this dataset states such a corridor, and the cross-linked substance set holds a single row, so turning it on today would degrade one formulation and nothing else — noise without coverage. Either way the missing corridor is reported as a gap on the result rather than hidden.
08
What this deliberately does not do
- Almost no fuzzy matching, and never on a synonym. Names reach a schedule entry through the forms the entry itself prints, through salt and ester forms, through dosage-form, grade and strength wording, and across British and INN spellings — mostly by canonicalising both sides onto one spelling and comparing for equality. The one exception: a typo in a single word of a multi-word name is tolerated, capped at one edit, and only against a row's own canonical or botanical name — never a synonym, and never proposed as an automatic match, only as a candidate a person confirms. A name one letter short of a real INN's synonym still resolves to nothing, on purpose.
- No class judgement dressed as a determination. Some schedule entries are a single word covering hundreds of substances — “Antibiotics” is one. Where a published INN stem or the instrument’s own filing suggests membership, that appears as a separate, quoted, page-cited proposal, never as a classification, and never where a named entry is the right citation instead.
- No billing. There is no payment code in this product. Nothing here is behind a paywall, which is part of why the provisional-data disclosure matters as much as it does.
- No claim to be a regulatory decision. For Schedule X and narcotics-adjacent substances the tool gives the classification, the citation, and a note that specialist counsel is required — and stops there, naming licence forms at most.
One item used to be on this list and no longer is. Market-similarity scoring is built and live on every nutraceutical report’s “Closest products in market” panel (drug reports do not carry it — the two check types key ingredients to different id spaces this scorer does not reconcile). It ranks your formulation against a corpus of real marketed products by a fixed formula — weighted ingredient overlap (60%), dose closeness on shared ingredients (25%), pack-format match (10%) and product-category match (5%) — never by a model, and a model never ranks or narrates a match. The corpus holds 6 verified products today, drawn so far from a single storefront rather than a market survey, and every match names its source and carries its own line: presence in the market is not evidence of compliance, and the engine’s verdict for your combination may differ from what a listed product implies.
09
Why “pre-screening” is the honest word
Every screen and every report in this product carries the same sentence: this is a pre-screening tool, results are informational only, and it does not grant, replace or predict any decision by FSSAI, CDSCO or a state licensing authority.
That word is chosen carefully. “Pre-screening” describes what actually happened: your formulation was compared, deterministically, against a digitised copy of a published schedule, and the comparison was cited to a page. It says nothing about what an authority will decide, because nothing in this system has any way to know that — and what the schedules permit is only one input into that decision.
The strength of any result here is bounded by the weakest thing behind it, and today that is the review status of the underlying rows: 831 of 832 carry a human sign-off, most of them accepting the machine’s extraction in bulk rather than opening the page. A green line means the engine compared your dose against what the machine read on the page it cites, and found no breach. It does not mean anyone has cleared your product. Confirm anything you rely on against the gazette itself, and take licensing questions to your state authority.
Browse the ingredient data to see the rows, their citations and their review state directly, or run a check and read the result against this page.