METER.

A reusable engine for human measurement.

METER turns supported response data into latent scores, diagnostics and auditable results—without fitting a new model to every dataset.

Pretrained scoring — no per-dataset fittingVersioned capability contractExplicit support checks and refusals
Demo dataset
Respondents
Items
Per-dataset METER fittingNone
Capability decision

Watch a real run

Choose an example to see METER check the data, determine what it can support and return the result.

Waiting to start…

  1. Read
  2. Check
  3. Decide
  4. Measure

Open the full demo →

How METER works

1 · Upload

Upload item-response data.

2 · Define

Choose items, coding and factor structure.

3 · Compare

Compare METER with conventional scoring.

4 · Assess agreement

See how respondent rankings differ between methods.

5 · Download

Download scores, comparison report and Methods text.

Run a guided example →

Built for researchers who repeatedly measure latent characteristics

Have a dataset? See if METER supports it →

No per-dataset fitting

METER is trained once across simulated measurement problems and then left unchanged. For supported datasets, scoring is one inference pass, with no fitting or recalibration on this dataset.

Refusals, not extrapolation

Before scoring, METER checks the request against a versioned capability contract, and the submitted responses against executable data checks: sample size, scale length, item variance, whether the items cohere, and per-factor adequacy. Outside the supported region METER limits or refuses rather than extrapolating.

Provenance on every result

Each run records the model, capability contract, data schema, item mapping, runtime and run ID, so every result can be audited backward to exactly what was asked and what answered.

Evaluated against conventional psychometrics

Known-truth recovery

On 40 unseen synthetic worlds with known latent truth, the pretrained core recovered person scores at median r = 0.939 (a per-dataset specialist reached 0.944).

Synthetic ground truth; real data has no truth to compare against.

External transfer without refitting

Prospectively locked before the data were opened: pooled agreement r = 0.985 with country-fitted models across 28 countries (66,812 respondents), replicated at 0.985 in an earlier wave.

Agreement with fitted models — convergence, not latent-truth accuracy.

Five-factor supplied structure

Factor-wise agreement 0.92–0.98 in an independent cohort, and median 0.979 when the same model weights scored a different 50-item instrument.

Scores transferred; factor-correlation recovery failed its prespecified gates, and METER never discovers structure.

See all benchmarks, including failed evaluations →

Starting with questionnaire and assessment data, METER is being built as a common measurement layer across supported instruments and populations. The current research release focuses on scoring, diagnostics, capability checks and auditable provenance.