# widgetspec

**Answers that are interfaces, with receipts.** A small JSON format for answer screens, a safe formula language,
an auditor that checks the polish against the numbers, a rule-based chooser that says when plain text is enough,
a plain-text fallback and an HTML renderer. Zero dependencies, one file, Node 18+ and browsers.

On 7 October 2026 ChatGPT started answering with screens instead of paragraphs: comparisons side by side,
charts, calculators with inputs you can change, timelines and checklists, picked per question from a library of
components. A screen looks finished whether or not its numbers are right. widgetspec is for anyone who builds,
generates or reviews answers like that: describe the screen as data, check it, then show it, or print it as text.

```
npm test          # 14 tests
npm run demo      # a short tour
```

## The format

A spec is a question and a list of blocks. Six kinds:

| kind | what it shows | required fields |
| --- | --- | --- |
| `text` | paragraphs | `text` |
| `compare` | items side by side, one row per attribute | `items`, `rows[{label, values, unit}]` |
| `chart` | bar or line chart | `type`, `x`, `series[{name, values}]`, `unit` |
| `calc` | inputs you can move and outputs that follow | `inputs[{id, label, value, min, max}]`, `outputs[{id, label, formula}]` |
| `timeline` | steps in order | `steps[{when, label, detail}]` |
| `checklist` | things to tick | `items[{label, done}]` |

Every block may carry `title`, `source` and `note`. The source and every formula travel with the block and are
printed under it as its receipt.

```json
{ "v": 1, "question": "Split a $186 bill between 5 with an 18% tip",
  "answer": [
    { "kind": "text", "text": "With an 18% tip the total is $219.48, so each of the 5 pays $43.90." },
    { "kind": "calc", "title": "Split the bill",
      "inputs": [ { "id": "bill", "label": "Bill", "value": 186, "min": 0, "max": 1000, "unit": "$" },
                  { "id": "tip_pct", "label": "Tip", "value": 18, "min": 0, "max": 30, "unit": "%" },
                  { "id": "people", "label": "People", "value": 5, "min": 1, "max": 20 } ],
      "outputs": [ { "id": "tip", "label": "Tip", "formula": "bill * tip_pct / 100", "unit": "$", "format": "money" },
                   { "id": "total", "label": "Total", "formula": "bill + tip", "unit": "$", "format": "money" },
                   { "id": "each", "label": "Each person pays", "formula": "total / people", "unit": "$", "format": "money" } ] } ] }
```

## Formulas

Numbers (`1_000`, `1.5e3`, `18%` = 0.18), names, `+ - * / ^`, comparisons, `&&`, `||`, `!` and these functions:
`min max round floor ceil abs sqrt ln log10 pow if pmt fv`. `pmt(rate, periods, principal)` is a loan payment;
`fv(rate, periods, payment, start)` is what regular saving grows to. The parser is written out by hand: nothing
outside this language can run, formulas are limited to 500 characters and 64 levels of nesting, and an output can
use the inputs and the outputs above it.

## Audit: polish is not proof

`audit(spec)` returns findings with a path, a level and a code, and a grade: `pass`, `check` (warnings) or `fail`.

| code | level | what it catches |
| --- | --- | --- |
| `INVALID` | error | the spec is malformed (every problem, with its path) |
| `TOTAL_MISMATCH` | error | a "Total" row that is not the sum of the rows above it with the same unit |
| `UNKNOWN_VAR` | error | a formula uses a name that is not an input or an earlier output |
| `NOT_FINITE` | error | an output has no value at the start, or breaks when an input is moved to either end |
| `OUT_OF_RANGE` | error | an input starts outside its own min and max |
| `UNGROUNDED_NUMBER` | warn | a number in the text that appears in no widget's data and not in the question |
| `NO_SOURCE` | warn | a chart or comparison with numbers and no source |
| `NO_UNIT` | warn | numbers with no unit |
| `MISSING_VALUE` | warn | an empty cell or a gap in a series |
| `UNUSED_INPUT` | warn | a slider that changes nothing |
| `THIN_LINE` | warn | a line chart with fewer than three points |

`mutants(spec)` plants one known fault at a time (a text number nudged by 7%, a source or unit removed, a total
off by 10%, a typo in a formula name, a divisor allowed to reach zero, an input outside its range, a slider that
does nothing, a gap in the data, `*` swapped for `+`) and `trial(spec)` reports which ones the audit notices. On
the five clean example specs it catches 24 of 24. Two more mutations are there to show the limit: a chart point or
a table cell changed by 7% that nothing else in the answer refers to passes unnoticed (0 of 2). The audit checks
that an answer agrees with itself, not that it agrees with the world.

`data/specs/broken.json` is a polished answer with seven kinds of fault planted in it; the CLI exits with code 4 on
it, so `widgetspec audit` can sit in a test script.

## Text or interface?

`decide(question)` picks a kind with weighted rules and returns its reasons. It answers with plain text unless a
signal is strong enough. On the 120 labelled questions in `data/questions.json` (written by us, 20 per kind) the
rules were tuned on the 60 even-numbered ones:

| questions | right | note |
| --- | --- | --- |
| tuning half (even ids) | 60 / 60 | what tuning on the same questions looks like |
| held-out half (odd ids) | 47 / 60 (78.3%) | the number to quote |
| held-out, chose an interface | 37 of 37 correct | it never added a screen to a text question |
| held-out, interface questions | 37 of 50 found | the 13 misses all fell back to plain text |

Run `widgetspec bench --split test` for the matrix and the misses; the largest group (6 of 13) are "A or B?"
questions with no "vs" in them.

## Command line

```
widgetspec decide "<question>"           text or interface? which kind, and why
widgetspec validate <spec.json>          is it a well-formed spec (exit 1 if not)
widgetspec audit <spec.json> [--json]    check the polish against the numbers (exit 4 on errors)
widgetspec text <spec.json>              the plain-text answer (every number and formula kept)
widgetspec html <spec.json> [--out f]    a standalone page with working sliders
widgetspec calc <spec.json> [id=value]   run the calculators in a spec
widgetspec eval "<formula>" [id=value]   evaluate one formula
widgetspec bench [--split test|tune|all] how often decide() agrees with the labelled questions
```

## API

```js
const W = require("./src/widgetspec.js");      // or <script src="widgetspec.js"> → window.WidgetSpec
W.validate(spec)            // { ok, errors: [{ path, msg }] }
W.audit(spec)               // { grade, errors, warnings, findings: [{ path, level, code, msg }] }
W.decide(question)          // { kind, interface, reasons, scores }
W.compute(calcBlock, set)   // { values, errors }
W.toText(spec)              // plain text
W.toHTML(spec)              // static HTML (class names start with ws-; style with src/widgetspec.css)
W.render(spec, element)     // browser: HTML plus working sliders and checklists
W.expr.evaluate("pmt(6.5% / 12, 360, 350000)")   // 2212.24
W.bench(questions, filter)  // accuracy, confusion matrix, misses
W.mutants(spec) / W.trial(spec)   // plant known faults; which does the audit catch?
```

## Examples

- `examples/audit-an-answer.js` audits a spec and prints the text version when it fails.
- `examples/text-or-interface.js` runs `decide()` on a few questions or on yours.
- `examples/build-a-calculator.js` builds a car-loan answer in code, audits it and writes a working page.

## Limits

`decide()` reads words, not meaning; it is a baseline and a fallback, not a model. The audit checks arithmetic,
grounding and robustness, not whether a source is true. Charts are bar and line only. The questions and specs are
ours; they are small and in English.

MIT licence.
