← Back to blog
May 12, 2026

Why we don't let an LLM invent hazard classification data

It would be easy to build a version of ChemDocs where you type in a chemical name and an AI model writes you a Safety Data Sheet on the spot. It would also be dangerous, and we think it would be dishonest to ship.

Large language models are trained to produce text that reads as plausible and fluent. For most writing tasks that is exactly what you want. For hazard classification — a flash point, an LD50, a GHS category, a CAS number — plausible is not the bar. Correct is the bar. A model can produce an H-statement that looks completely reasonable and is simply wrong, with nothing in the output to signal the difference.

The consequence of a wrong number on a marketing page is embarrassment. The consequence of a wrong number on a Safety Data Sheet is that someone handling your product doesn't know it can cause serious eye damage, or doesn't know it's an oxidizer that shouldn't be stored near flammables. That is a real safety and legal risk, not a demo inconvenience.

So ChemDocs draws a hard line: hazard classification data comes from a structured database, seeded from public regulatory sources — ECHA's harmonized CLP classifications and PubChem's aggregated GHS data. Every record in that database carries its source URL and the date it was retrieved. Document assembly then takes that verified data and inserts it into the standard 16-section SDS template using fixed regulatory phrasing, not freeform generated text.

Where it gets interesting is mixtures — most real products aren't a single pure compound. The UN GHS "Purple Book" defines exact mathematical rules for classifying a mixture from the classification and concentration of its components (ATE additivity for acute toxicity, concentration-cutoff summation for skin and eye hazards, a summation method for aquatic hazard). We implemented those rules as actual calculation logic, tested against known reference cases — not asked a model to estimate the answer.

And when a chemical in our database doesn't have a confirmed harmonized classification — which happens, ECHA doesn't have Annex VI entries for everything — we flag it explicitly instead of generating a plausible-looking placeholder. You see "needs manual sourcing," not a confident-looking number that happens to be invented.

AI has a role in this product eventually — helping draft company-specific handling notes, for instance, with a human reviewing before anything ships. It does not have a role in deciding what H-code goes on your label. That part stays boring, sourced, and auditable, on purpose.