MAI-Alchemy
A series

Precisely Stated Problems

A problem precisely stated… is a problem you can finally work on. These are hard analytical problems — stated precisely, with an honest look at where a new class of tools might help find a way through. This is early, exploratory work, openly under development — potential directions, not finished answers. No secrets given away — just the thinking.

① The problem, precisely stated② How it's traditionally handled③ How we're exploring it④ Where it points
Problem 01

The Noise Problem

"A LUMA spectral library isn't possible — the noise in the signal corrupts the 12-band fingerprint, so you can never match a compound reliably."
Why it's true

A LUMA fingerprint is a shape — how a compound absorbs across 12 VUV bands. But at real, working concentrations, every band carries electronic noise. Noise nudges each band a little, the shape wobbles, and two runs of the same compound stop looking alike. Build a library on wobbly fingerprints and nothing matches. That's the wall that kept a LUMA library from ever existing.

How it's traditionally handled

The standard toolkit: smoothing (Savitzky–Golay, moving average, Whittaker), baseline correction, and signal averaging. These are necessary — everyone uses them to strip electronic fuzz and instrument drift. But they hit a ceiling. Smooth too hard and you erase real shoulders or merge two peaks into one. They work one channel at a time, in the time dimension only. And they can't tell a faint real signal from noise — a height threshold either keeps a trace peak (noise and all) or throws it out.

REAL MOLECULE — COHERENT NOISE — INCOHERENT vs the test isn't height — it's shape
A real absorption holds a consistent, on-template shape across all 12 bands. Noise is twelve independent random wiggles — no molecular shape at all.
How we're exploring it — conceptually

The key move is to stop treating VUV as a 2D trace. It's 3D: at every instant the detector records a 12-band spectrum — a shape. And here's the physical fact we lean on: a real molecule's absorption is coherent across those 12 bands — one consistent, on-template shape — while noise is incoherent, twelve random wiggles that form no molecular shape.

So instead of asking "is this wiggle tall enough to be real?" we ask "is this 12-band shape coherent — does it look like a real molecular absorption, or like noise?" Coherence is a stricter test than height: it recovers real peaks that height missed and rejects noise — and it can never invent a false one. That's what lets you pull a clean, trustworthy fingerprint out of a noisy run, even near the floor. (The concept is the point; the implementation stays under the hood.)

Where it points

A potential foundation. In the runs we've done so far, the fingerprints appear to hold — and a LUMA library, long assumed impossible, starts to look feasible. If it holds up under wider testing, it becomes the base the rest of this work explores from. Early, and openly under development.

Problem 02

The σ-Wall — When Spectra Lie

"A VUV spectrum can't identify saturated sulfur species — sec-butyl and isobutyl mercaptan give fingerprints a spectral match calls over 99% identical. So spectral identification of sulfur is impossible."
Why it's true

This one is physics, and it deserves respect. A VUV fingerprint comes from a molecule's electronic transitions. In saturated molecules, the deep-VUV absorption comes from the σ-bond electrons — the ordinary single bonds that make up the skeleton. Move one methyl group from the end of the chain to a branch and you've changed the molecule's name, its boiling point, its behavior — but you've barely touched its σ electrons. The spectrum hardly moves. We call it the σ-wall: a hard physical ceiling on how much a 12-band spectrum can distinguish saturated look-alikes. Below is the wall in our own measured data — not a simulation.

sec-butyl mercaptan isobutyl mercaptan cosine: “over 99% identical” 125130135140145150155160170180200220 VUV band (nm) — normalized absorbance, measured apex spectra
Two real fingerprints from an ASTM D5623 calibration standard, measured on one instrument. Same atoms, same bonds — one branch apart. Band for band, nearly twins.
How it's traditionally handled

Two ways — both dodge the physics. The classic way is retention time: run a standard, note when each compound elutes, and identify by the clock. It works until it doesn't — a peak drifts, two compounds co-elute, a column ages, a method changes, and the clock quietly points at the wrong name. The newer way is to trust the match score anyway: a cosine similarity of 0.999 feels authoritative. But when two spectra are twins, a high score isn't evidence — it's the σ-wall telling you the question can't be answered this way, dressed up with decimal places. The score is confident precisely where it deserves to be humble.

How we're exploring it — conceptually

We stopped asking the spectrum to do what physics won't let it do. Restate the problem precisely and it changes shape: the claim isn't "the spectrum identifies every sulfur" — it's "the spectrum identifies the family, and something else must pick the twin." Stated that way, the answer is sitting in the run: the twins may share a spectrum, but they don't share a boiling point — and boiling point is what decides elution order on the column. Structure sets the property; the property sets the order; the order is measured physics, not a lookup.

So identification becomes convergence: the fingerprint narrows the field, elution order and boiling point discriminate the look-alikes, and the certified contents of the standard anchor the truth. No single axis is trusted — the ID stands only where the axes agree. And when they don't agree, the system says so, out loud, instead of hiding behind a 0.999. A system that knows when it can't know is worth more than one that always answers.

Where it points

On the D5623 standard we've run, all eight certified sulfurs — the twins included — appear to land on the right names once the axes are made to agree, each ID auditable axis by axis. That's an early result on one standard, not a closed case. The σ-wall doesn't move; the potential is in walking around it rather than through it — and in a system that says so out loud when it can't be sure.

Problem 03

The Property Problem — Reading What a Product Can Do

"A detector can only tell you how much is there. Physical properties — refractive index, boiling point, density — come from bench tests, not from a spectrum."
Why it's true

The whole calibration paradigm is built on Beer's law: absorbance is proportional to amount. For the detectors chromatography grew up on, that really is the whole story — a single-channel detector hands you one number per instant. Amount, nothing else. Properties live in structure, and structure simply wasn't in the signal. So the lab evolved around the gap: the chromatograph tells you how much, and a bench full of other instruments tells you what it can do.

How it's traditionally handled

Three ways. Run the bench tests — density meters, distillation, vapor pressure, engine tests. Accurate, and slow: each property is its own instrument, its own sample volume, its own queue. Calculate from composition — a full DHA names every peak and property tables do the rest. It works, but it demands a 100-meter column, hours per run, and every identification to be right. Or black-box chemometrics — regress spectra against measured properties and hope. Inside its training set it's fine; drift outside — a new blend chemistry, an oxygenate that wasn't in the model — and it fails silently, because it never knew any physics to violate.

12 VUV bands electronic structure RI · BP · density from the same electrons that did the absorbing physical properties what the product can DO
The causal chain. Bands are electronic structure; structure sets properties; properties decide what a product is worth.
How we're exploring it — conceptually

Restate the problem precisely and the assumption shows itself: "a detector only measures amount" was true of single-channel detectors. A LUMA isn't one. Twelve VUV bands are a readout of electronic structure — which transitions are available, which bonds are present. And structure is where properties come from. The cleanest case is refractive index: refraction and VUV absorption are the same electrons responding to light — the physics practically insists they correlate. Within a chemical family, they do, strongly.

So the approach isn't "regress spectrum against property and hope." It's physics first: read the family from the shape of the fingerprint, then apply that family's structure–property law — local models anchored to known physics instead of one global black box. Anchoring buys two things chemometrics alone can't give: reach (out on unfamiliar blends, the physics-anchored local models cut prediction error nearly in half versus generic regression) and honesty — a family-first model knows when a sample doesn't belong to any family it's seen, and says so instead of guessing. Some properties follow rigorously; others hold within a family. We say which is which — that's the difference between a claim and a con.

Where it points

The potential: one injection that reads composition and a property profile — refractive index from the physics, boiling point and density family by family. Early results are encouraging within families; broader validation is ongoing. If it holds, the question could shift from "what's in the sample?" toward the one that always mattered: "what might this product do?"

Problem 04

Co-elution — Two Peaks, One Lump

"When two compounds elute together, the only fix is better separation — a longer column, a slower ramp, another method. The detector can't un-mix them."
Why it's true

A single-channel detector records one summed number per instant. Two overlapped peaks arrive as one lump, and there is literally no information in the signal to split them — the sum has erased the parts. That's why separation science carried the entire burden for decades: it's why DHA columns are 100 meters long and the runs take hours. If the detector can't tell two compounds apart, the column has to keep them apart.

How it's traditionally handled

More separation: longer columns, slower ovens, or a second dimension entirely (GC×GC) — real fixes, paid for in run time, hardware, and complexity. Or shape guessing: fit the lump with two idealized peak shapes and split the area by geometry. That holds only as long as peaks behave ideally — which trace-level, tailing, real-world peaks don't.

what one channel sees what twelve bands see
Same moment in the run. One channel sums the overlap into a single lump; in 12-band space the two compounds never stopped being two.
How we're exploring it — conceptually

State it precisely and the assumption is visible: "the detector can't un-mix them" assumes the detector records a number. A LUMA records a spectrum — every instant of the lump is a 12-band shape, and the lump's shape is just the two compounds' shapes added together in known proportions that change across the peak. Walk through the overlap and the blend ratio shifts point by point — the spectrum visibly migrates from the leading compound's shape toward the trailing one's.

That migration is the tell. It says there are two compounds, it shows what each one's spectrum looks like (the leading edge is nearly pure one, the trailing edge nearly pure the other), and it makes splitting the amounts a piece of algebra instead of geometry — un-mixing shapes, not guessing curves. The separation happens in spectral space, not on the column.

Where it points

The potential: where two spectra differ enough, co-eluters may be separated in spectral space rather than on the column — letting the column do less because the data does more. When two co-eluters are themselves spectral twins, the split is flagged uncertain rather than guessed. Early and exploratory — you can watch the idea in the Il-LUMA-nate demo, co-elution view on.

Problem 05

The Calibration Treadmill

"Quantitation demands calibration — every compound, every level, on your instrument, re-run on schedule. No standards, no numbers."
Why it's true

Detectors respond differently to different compounds, so an area means nothing until a response factor turns it into an amount. And response factors have always been treated as instrument-specific — your flame, your detector, your drift — so every lab measures its own, for every compound, at multiple levels, over and over. Calibration isn't a step; it's a treadmill that eats the lab's calendar.

How it's traditionally handled

Multi-level calibration curves per compound, per instrument, refreshed on schedule. Internal standards to patch drift between calibrations. Surrogates when the compound can't be bought. All of it necessary under the assumption — and none of it questioning the assumption: does the response factor really belong to the instrument?

How we're exploring it — conceptually

Beer's law: A = ε · l · c. Absorbance equals absorptivity × path × concentration. Look at where the response lives: ε is a property of the molecule — how strongly its electrons couple to light at each wavelength. It isn't your instrument's opinion; it's the compound's physics. And VUV absorbance is famously linear over a wide range.

Stated precisely, the problem inverts: calibrate the molecule once, not the instrument forever. Response factors live in the library, alongside the fingerprints — measured once, or carried across a family — and each run is tied down with an anchor instead of a fresh curve for every compound. Honesty matters here: this is pseudo-absolute, not magic — anchors, checks, and validation against certified standards stay in the loop. What leaves is the treadmill.

Where it points

The potential: a response factor that lives in the molecule (its ε) rather than the instrument — so a run might be anchored and validated against certified standards instead of rebuilt as a fresh curve per compound per week. Pseudo-absolute by design: the physics proposes the number, certified standards keep it honest. Still under development — but if it holds, the calendar goes back to the chemistry.

Problem 06

Trace — Below the Height Floor

"Below the noise floor a peak is undetectable. If it doesn't rise above the baseline, it isn't there — and no software can change that."
Why it's true

Height thresholds exist for a reason. With one channel, height is the only evidence there is — noise and signal look identical except for size. Lower the threshold and false positives flood the peak table; raise it and trace compounds silently vanish. Every lab lives on that knife edge, because the knife edge is all a height test offers.

How it's traditionally handled

Get more signal: concentrate the sample, inject more, or move to a more sensitive — and destructive — detector. All legitimate, all costly, and none of them changes the logic: detection still means tall enough.

12-band coherence score — the cliff real trace peaks ≈ 1.0 the cliff noise ≈ 0.08 · 1/12 floor
Not a gray zone — a cliff. Real molecular absorptions stay coherent across all 12 bands even at trace level; noise never does.
How we're exploring it — conceptually

Same move as the noise problem, taken to the floor. A trace compound is small — but it's still coherent: its shape shows up faintly and consistently in every band. Noise is loud but incoherent. Score that, and there's no knife edge at all: real absorptions score near 1.0; noise sits at the mathematical floor — 1 in 12, about 0.08, fixed by the number of bands. A cliff, not a slope. Detection stops meaning "tall enough" and starts meaning "shaped like a molecule" — and a shape test can reach below the height floor, with the library telling it exactly what shape to look for.

And the honest boundary: this does not turn a LUMA into an ultra-trace detector — raw sensitivity is set by the optics, and dedicated sulfur detectors still win on pure detection limit. The claim is different and, for real decisions, often worth more: down at trace levels, when something is there, you don't just learn that sulfur exists — you learn which one. Detection with identity attached.

Where it points

The potential: a shape test that may reach below the height floor — recovering trace peaks with a name attached, and rejecting noise by the same test. Not ultra-trace sensitivity (dedicated detectors still win on pure detection limit); the possible gain is identity at low levels, where the old logic saw only noise. Early findings, still being validated.

The list keeps growing

Every hard problem gets the same treatment — stated precisely, then explored with the new tools. Queued up:

Reading a whole blend at once — PIONA from one run Carrying a library across instruments Column health from data you already have

A problem precisely stated… is where the work begins.

← Back to What I Learned