Ad Astra ResearchTechnical Validation

Measurement confidence · error propagation · independent verification

Most breakthrough results are measurement artifacts.

Ad Astra Research

I do technical validation for deep tech — the unglamorous work of determining whether a claimed effect survives scrutiny, or whether it was instrumentation all along. Independent, adversarial, and documented. Ad Astra Research is the name I take freelance validation work under.

Why validation fails

Wrong results rarely announce themselves

An experiment producing a false positive looks exactly like one producing a real effect. The difference only shows up when someone goes looking for the ways the measurement could be lying — and most teams have neither the time nor the incentive to do that to their own work.

01

The instrument is part of the experiment

Sensors drift. Calibration made under one condition does not transfer to another. A thermocouple reading five degrees high produces a beautiful, reproducible, entirely fictional result — and it reproduces because the error is systematic, not random.

02

Reproducibility is not validity

A systematic error reproduces perfectly. Teams often treat repeatability as proof of a real effect when it is equally consistent with a stable flaw in the measurement chain. The right question is not whether the result repeats, but whether it survives a deliberate attempt to break it.

03

Nobody runs enough null experiments

Without a properly matched control — same apparatus, same conditions, inactive material — there is no baseline against which an anomaly means anything. Null experiments are the least interesting thing to run and the most important thing to have.

04

Error bars are estimated, not propagated

Uncertainty is frequently assigned by intuition rather than computed through the full instrument chain. When every sensor, shunt, flow meter, and conversion step is accounted for, claimed effects often fall inside the margin they were supposed to exceed.

05

Motivated reasoning is structural, not personal

Founders under fundraising pressure and researchers under publication pressure are not dishonest — they are subject to an incentive gradient that makes a positive interpretation easier to reach than a negative one. Independent validation exists to remove that gradient.

Methodology

A staged protocol, run in order

Each stage is designed to kill the claim. If the claim survives all seven, that survival means something. The sequence matters: expensive stages are never run before cheap ones have had their chance to end the investigation.

1

Claim decomposition

Reduce the technology's headline claim to a set of specific, falsifiable measurements. "Thirty percent more efficient" becomes a defined set of input and output quantities, under defined conditions, with a defined comparison baseline. Ambiguity here invalidates everything downstream.

Output: a testable claim specification
2

Measurement chain audit

Map every instrument between the physical phenomenon and the reported number. Establish each one's calibration history, traceability, stated accuracy, and environmental sensitivity. Identify the weakest link — usually one sensor or one conversion step carrying disproportionate uncertainty.

Output: annotated instrument chain with per-stage uncertainty
3

Error propagation analysis

Compute total system uncertainty through the full chain rather than estimating it. Determine the minimum effect size the apparatus can resolve with confidence. Frequently this stage alone resolves the question: the claimed effect is smaller than what the instrument can honestly detect.

Output: resolvable-effect threshold and confidence interval
4

Adversarial hypothesis generation

Enumerate every mundane explanation that would produce the same observation. Chemical pathways, parasitic energy sources, interference, thermal transport differences, circuit artifacts. Each conventional explanation becomes a hypothesis to be tested and eliminated by design — not dismissed by argument.

Output: ranked alternative-explanation register
5

Null and control experiment design

Build matched controls that isolate the claimed mechanism as the only variable. The control must share the apparatus, the conditions, and the measurement chain. Where calibration baselines differ physically from active conditions, that mismatch is characterized rather than assumed away.

Output: control protocol with matched conditions
6

Independent instrumentation

Measure the critical quantity a second way, using a physically different method. If a thermal claim rests on flow calorimetry, verify with an independent thermal method. Agreement between independent methods is far stronger evidence than precision within a single one.

Output: cross-method agreement or discrepancy analysis
7

Documented determination

Deliver a written finding with the data, the reasoning, the residual uncertainty, and the explicit limits of what was tested. The report states what would change the conclusion. A validation that cannot be audited by a third party is not a validation.

Output: investor- or board-ready technical determination

What a determination looks like

Three honest endings

A validation engagement resolves to one of three outcomes. None of them is "it works." Overstating certainty in either direction is the failure mode of validation itself.

Flaw identified

A specific error in the measurement chain or experimental design accounts for the observed effect. The finding names the mechanism, demonstrates it, and shows what the corrected measurement yields.

Within measurement error

The effect, if present, is not resolvable above the uncertainty of the apparatus. This is not a negative result — it is a statement about the instrument. The report specifies what measurement capability would be needed to resolve it.

Survives scrutiny, inconclusive

The claim withstood every alternative explanation tested and the effect exceeds propagated uncertainty. The honest conclusion is a defined next stage of investigation, not a verdict.

A validation report that concludes "the technology works" is usually a report that did not look hard enough. The value of independent validation is the specificity of its uncertainty, not the confidence of its conclusion.

Failure mode catalog

The errors I check for first

These recur across energy, thermal, and detection systems regardless of the underlying technology. Each one has, somewhere, produced a result that a competent team believed was real. I check the catalog explicitly rather than relying on memory.

Concentration measured where mass flow was the claim
An instrument reporting a species as a fraction of sample volume says nothing about total quantity emitted unless flow rate is measured alongside it. Any process that adds diluent to the stream will read as a dramatic reduction on a concentration-only instrument while emitting exactly as much as before. The check is simple and frequently skipped: confirm that the quantity being measured is the quantity the claim is about.
Parasitic chemical or stored energy pathways
Ordinary chemistry — oxidation, recombination, phase change, stored reactants — producing energy that is attributed to the novel mechanism. Quantifying the maximum possible conventional contribution is a standard first check on any excess-energy claim.
Flow measurement error in thermal systems
Calorimetry that depends on mass or volume flow inherits every error in the flow measurement. Flow meters drift, respond nonlinearly across their range, and are sensitive to temperature and entrained gas. Independent verification — gravimetric cross-check against the flow reading — catches what calibration certificates do not.
Temperature sensor defects and calibration drift
Defective, uncalibrated, or improperly mounted temperature sensors produce systematic offsets that survive repetition. Thermal contact quality, self-heating, and lead resistance all introduce errors that look like signal. Sensor swaps and in-situ recalibration are cheap and frequently decisive.
Radio-frequency interference generating artifacts
High-frequency modes in a system couple into measurement electronics and produce readings unrelated to the physical quantity. A reliable diagnostic: interrupt the input and observe whether the measurement changes on a timescale consistent with real physics or with an electrical artifact.
Calibration baseline mismatch
A calibration performed under conditions that differ physically from active operation — different gas, different loading state, different heat transfer regime — silently misassigns the baseline. The apparent anomaly is then the difference between two incomparable conditions rather than a real effect.
Phase-shift artifacts in power measurement
Reactive components, including stray inductance introduced by the measurement setup itself, produce phase shifts that corrupt computed power — occasionally yielding physically impossible values such as negative input power. Where power measurement matters, the circuit topology is characterized rather than assumed.
Unintended flow paths in the apparatus itself
Leaks, cracks, and unsealed joints create routes the experimental design never accounted for — recirculating a product back into an input, bypassing a measurement point, or admitting ambient air into a supposedly closed stream. These are usually found after the campaign rather than before, which is an argument for pressure-testing and tracing the flow path as a precondition for running rather than as diagnosis afterward.
Sampling position and instrument condition drifting mid-campaign
Probe placement that varies between runs, sensors degraded by the conditions they are measuring in, and instruments repaired partway through a series all introduce differences between datasets that look like differences between experimental conditions. Position and instrument state need to be recorded per run, not assumed constant across a campaign.
Surface damage mistaken for detection events
In track-based and surface-sensitive detectors, chemical or physical damage to the detector medium can be visually indistinguishable from genuine particle tracks. Exposure controls and independent detection methods are required before attributing surface features to the claimed phenomenon.
Detection thresholds asserted rather than demonstrated
Claims of low-energy or low-flux detection frequently rest on a threshold that was never independently established for that specific setup. Shielding geometry, background characterization, and discrimination logic must be validated against known sources before any detection claim is credible.

Engagement templates

How I would approach three typical ventures

These are illustrative templates, not client work — worked examples of how the protocol applies to common deep tech claim structures. Each shows the questions I would ask, the tests I would design, and the conditions under which I would tell an investor to walk.

Template A — Thermal / energy claim

A venture claiming anomalous excess heat from a solid-state reactor

"Our reactor produces 1.4× thermal output relative to electrical input, sustained over multi-hour runs, reproducible across dozens of cells."

First questions

  • How is input power measured, and does the method account for reactive components?
  • What calorimetry type, and what is its demonstrated accuracy at this power level?
  • What does a null cell — identical apparatus, inactive material — produce under the same conditions?
  • What is the total chemical energy available in the cell over the run duration?

Tests I would design

  • Gravimetric cross-check of flow measurement against the flow meter's own reading
  • Joule-heating calibration across the full operating range with a known resistive load
  • Input-power measurement by two independent methods: precision shunt with oscilloscope, and a high-integrity power meter
  • Input interruption test to distinguish thermal mass response from measurement artifact
  • Sealed-cell mass balance to bound chemical contribution

Kill conditions

  • Claimed excess falls within propagated uncertainty of the calorimeter
  • Null cells show comparable excess
  • Available chemical energy accounts for the integrated output
  • Effect disappears under independent power measurement

What the investor gets: a stated resolvable-effect threshold for the current apparatus, a determination of whether 1.4× is above or below it, and — if the effect survives — a specification of the instrumentation upgrade required to make the claim defensible to a skeptical third party.

Template B — Detection / sensing claim

A venture claiming a novel detector with order-of-magnitude sensitivity improvement

"Our detector resolves particle events an order of magnitude below the noise floor of conventional systems, with sub-percent false positive rate."

First questions

  • What known-source calibration establishes the claimed threshold, and was it done on this specific apparatus?
  • How is background characterized, and over what duration?
  • What is the discrimination logic, and was it developed on the same data used to validate it?
  • Can detector medium damage or environmental artifacts mimic a true event?

Tests I would design

  • Blind calibration against traceable known sources at and below the claimed threshold
  • Extended background runs with the source removed, analyzed with identical discrimination logic
  • Held-out validation: re-derive discrimination parameters on one dataset, test on another
  • Deliberate artifact injection — environmental, chemical, electrical — to test discrimination robustness
  • Shielding geometry verification against modeled attenuation

Kill conditions

  • Threshold not reproducible against traceable sources
  • Background runs produce events at comparable rates
  • Discrimination logic fails on held-out data
  • Injected artifacts pass as genuine events

What the investor gets: an independently established detection threshold with a measured false-positive rate, and a clear statement of whether the sensitivity claim is a property of the detector or of the analysis pipeline applied to it.

Template C — Efficiency / power electronics claim

A venture claiming a step-change in conversion efficiency

"Our topology achieves 99.2% conversion efficiency across the operating envelope, versus 96–97% for conventional designs."

First questions

  • At 99% efficiency, the measurement must resolve a 1% loss — what is the instrument's accuracy relative to that?
  • Is efficiency measured directly, or inferred from a thermal or modeled proxy?
  • What is the comparison baseline, and was it measured on the same rig?
  • Does the claim hold across load, temperature, and duty cycle, or only at a favorable operating point?

Tests I would design

  • Direct input/output power measurement with instruments whose accuracy is demonstrably better than the loss being measured
  • Calorimetric loss measurement as an independent cross-check on the electrical method
  • Full-envelope sweep across load, ambient temperature, and duty cycle
  • Head-to-head against a conventional unit on identical instrumentation
  • Phase and topology characterization to rule out measurement-induced reactance

Kill conditions

  • Measurement accuracy insufficient to resolve the claimed difference
  • Advantage disappears outside a narrow operating point
  • Baseline was measured on different instrumentation
  • Electrical and calorimetric loss measurements disagree

What the investor gets: a cross-verified efficiency figure across the real operating envelope, an apples-to-apples comparison against the incumbent on identical instrumentation, and an explicit statement of where the claimed advantage does and does not hold.

Capability

What I can measure

Validation is only as good as the instrumentation behind it. I build and operate the measurement systems myself rather than subcontracting the part that matters most.

Thermal and calorimetric

  • Flow, air-flow, Seebeck, and isoperibolic calorimetry
  • Flow calorimetry resolving ±0.05°C differential — roughly ±1.5 W on a 100 W system
  • Independent mass-flow verification against gravimetric reference
  • Baseline drift correction and uncertainty quantification
  • Custom four-wire PT1000 RTD probes, built and calibrated in-house

Electrical and power

  • Precision shunt and oscilloscope-based power measurement
  • High-integrity power metering for cross-verification
  • Phase and topology characterization
  • Custom interface electronics for pulsed and mixed-mode drive

Vacuum, pressure, and gas

  • High-vacuum through high-pressure system design
  • Mass spectrometry for leak checking and gas analysis
  • Mass flow control and custom valve systems
  • Low-cost precision dosing and regulation hardware

High-voltage and plasma

  • AC plasma and pulsed discharge experiments on UHV platforms
  • Capacitor-bank pulsed power systems
  • Experimental safety procedure design and failure mode analysis
  • RF systems and interference characterization

Radiation and particle detection

  • Solid-state charged particle detection below 100 keV
  • Pulse-shape discrimination with RF amplification
  • Track-based detection, diamond detectors, gamma spectroscopy
  • Custom peak discrimination software for low-energy events
  • Shielding geometry design and background characterization

Data acquisition

  • Multi-channel isolated DAQ hardware built in-house
  • Distributed multi-site acquisition with synchronized data streams
  • Structured experiment and session data schemas
  • Parallel experiment recording with independent control
  • Real-time monitoring and automated anomaly flagging

Materials characterization

  • Electron microscopy with elemental analysis (SEM/EDS)
  • Time-of-flight secondary ion mass spectrometry
  • X-ray diffraction and Raman spectroscopy
  • Surface profilometry and morphology analysis
  • Sample preparation: plating, annealing, etching, rolling

Analysis and tooling

  • Custom analysis pipelines in Python, SQL, and MATLAB
  • Uncertainty quantification and error propagation modeling
  • Retrieval systems for querying large technical document sets
  • Version-controlled release practice for measurement software

Working together

Three ways an engagement usually starts

Scope depends on what you already have. A team with three years of data needs a different engagement than one with a prototype and a hypothesis.

Pre-investment due diligence

  • For investors evaluating a technical claim before committing capital
  • Review of existing data, measurement methodology, and instrument chain
  • Independent assessment of whether claimed effects exceed defensible uncertainty
  • Written determination suitable for an investment committee

Adversarial pre-raise review

  • For founders who want the hard questions asked privately first
  • I attack the result the way a skeptical technical diligence partner would
  • Findings delivered to you, not to the market
  • Output is a defensible claim, or a clear picture of what would make it one

Measurement system design

  • For teams whose instrumentation cannot resolve what they need to prove
  • Instrument chain specification, calibration protocol, and uncertainty budget
  • Custom DAQ and control hardware where commercial options do not fit
  • Built to be auditable by a third party from the start

Principal

Where the method came from

Ad Astra Research is led by Shriji Barot, an engineering leader with roughly a decade running experimental programs across plasma and gas-discharge systems, electrolytic cells, and gas-phase solid-state experiments. It is a domain where anomalous results are common, almost all of them turn out to be instrumentation, and the entire job is telling the difference. Experience gained from working at Industrial Heat LLC.

The role spanned senior technical oversight and program execution: directing experimental design, validation, and failure mode analysis across multiple laboratory sites; architecting the distributed data acquisition infrastructure those experiments ran on; and reporting technical findings and investment recommendations directly to the CEO and board. The measurement systems were built in-house — calorimeters of several types, vacuum and pressure systems, radiation detection chains, custom control electronics — because commercial equivalents either did not exist or could not resolve what needed resolving.

It also involved finding, documenting, and cataloging internally a long list of ways the measurements had been wrong. The failure mode catalog above is a generalized version of that list. The discipline that produced it is the product: build the measurement, attack the result, report the uncertainty honestly.

Education

  • M.S. Aerospace Engineering, University of Illinois at Urbana-Champaign
  • B.S. Aerospace Engineering, University of Illinois at Urbana-Champaign
  • Graduate research on deuterium and hydrogen interactions with chemically reactive metal hydrides — reaction kinetics, thermal systems, surface chemistry

Domains worked in

  • Plasma discharge and high-voltage systems
  • Electrolytic and gas-phase solid-state experiments
  • Calorimetry and thermal measurement
  • Radiation and charged particle detection
  • Materials characterization and preparation
  • Distributed instrumentation and control systems

Technical practice

  • Instrumentation and control software in Python, LabView, SQL, and C
  • Custom hardware design, CAD, and electronics integration
  • Retrieval and language-model tooling for private technical document sets
  • Automated analysis pipelines and anomaly detection
  • Configuration control and version-managed release practice

Industrial Heat LLC maintains a limited public presence as a private research organization. References, including direct contact with the CEO, are available on request to verify the scope and nature of that experience. Full professional background is at shriji-barot.com.

Why I build my own instruments

Cost is a validation problem

The reason bad measurements persist in early-stage research is usually not carelessness. It is that the right instrument costs more than the program can justify, so the team makes do with the wrong one and hopes.

Two representative examples from prior work. A precision gas dosing and pressure regulation system for electrolytic and metal hydride experiments, built for under three hundred dollars using orifice gaskets, low-cost solenoid actuators, and a custom control layer. And custom four-wire PT1000 temperature probes using top-tier sensing elements, built and calibrated in-house for under twenty dollars each — against catalog prices many times that, for programs needing dozens of them.

Neither was a compromise. Each was specified against what the experiment actually needed to resolve, and met that specification.

The harder half is the software

Cheap sensors are the easy part. The thing that actually determines whether a lab's data is trustworthy is whether there is one place it all lands. Most research groups accumulate instruments over years, each with its own vendor software, its own file format, and its own clock — and then reconcile them by hand in spreadsheets afterward. That reconciliation is where errors enter, and it is invisible in the final result.

The alternative is centralized acquisition and control: one interface that speaks to every instrument on the bench regardless of vendor or protocol, streams all of it into a single timebase, and lets an operator configure and run an experiment without touching six different applications. It needs to be simple enough that a technician uses it correctly under pressure, and structured enough that a dataset recorded eighteen months ago can still be traced to the exact configuration that produced it.

Across six laboratory sites, that software was what made the data defensible — synchronized multi-sensor streams, parallel experiments running under independent control, structured session and channel schemas, and a browser-based interface for monitoring runs remotely.

A recorded walkthrough of a centralized DAQ web interface built along these lines:

Watch the interface walkthrough

* This video is an example of what centralized DAQ streaming and control software can look like, based on a system I have worked on in the past. It is shown to illustrate the approach, not as a product on offer. A comparable system can be built from scratch for a client's specific instrument set.

The same body of work — hardware and software together — sits behind Ad Astra DAQ.

This matters for validation work because it changes what can be recommended. When a claim cannot be resolved with the instrumentation a team already has, the answer is often a measurement system that can be built rather than a budget line they cannot approve.

Where this is headed

A shared lab, eventually

Validation work pays for itself, but it is not the end goal. The thing I actually want to build is a lab that more than one person can use.

The picture is a centralized, well-instrumented space where people come to build what they genuinely want to build. Not a rented bench in a facility that treats them as tenants, and not an incubator where the work is shaped by whatever a program manager thinks is fundable. A room with real measurement capability in it, open to people with something they are serious about making.

The part that makes it worth doing is the cross-engagement. Someone deep in power electronics sits next to someone doing materials work, and the questions they ask each other are the ones neither would have thought to ask alone. Most hard problems in one discipline have already been solved in a neighboring one, and the only reason that knowledge does not transfer is that the people holding it are in different buildings. A lab where projects overlap by default turns that from an occasional lucky conversation into the normal condition.

It is also how skill actually propagates. Hands-on technical capability is learned by standing next to someone doing it, asking why they made a particular choice, and then making that choice yourself with someone watching. That does not survive well in environments where everyone is siloed on their own deliverable. It does survive in a room where helping with someone else's build is the ordinary thing to do, and where the person you help today knows something you will need next month.

And people build better when the thing they are building matters to them. Motivation is not a fixed quantity to be managed — it comes from working on something you find genuinely interesting and believe will do some good. Give capable people real tools and a reason to care, let them keep the upside of what they make, and the output is not comparable to the same people executing someone else's roadmap.

That is a long way off and it is not what I am selling today. But it is the direction the consulting work, the instrumentation, and the DAQ platform all point in, and it seemed worth stating plainly rather than leaving implied.

Contact

Bring me the claim you cannot afford to be wrong about

I work with venture investors conducting technical due diligence, deep tech teams who want an adversarial review before raising, and research organizations that need independent verification of an internal result.

Ad Astra Research is a one-person practice — the umbrella I take freelance validation and due diligence work under, alongside full-time engineering work. Based in Raleigh, North Carolina.