widgetspecDS-WS-01 · Product datasheetRev 1.0 · October 2026

Open-source library · MIT · v0.1.0 · zero dependencies

widgetspec

Answer components with receipts

1Features

  • Six block kinds: text, compare, chart, calc, timeline, checklist
  • Formula language with a hand-written parser: nothing outside it can run
  • Audit with eleven finding codes; exit code 4 for build pipelines
  • Plain-text fallback that keeps every number and every formula
  • decide() answers in text unless an interface is called for: 37 of 37 held-out interface choices right
  • One file, Node 18+ and browsers, 14 tests

2Applications

  • Reviewing generated answer screens before they reach a reader
  • Calculators and comparisons in documentation and help centres
  • Test suites for chat products that answer with interfaces
  • A plain-text view for readers who turn the visuals down

3Description

On 7 October 2026 ChatGPT began answering with interfaces: comparisons, charts, calculators, timelines and forms, picked per question from a library of components. A screen looks finished whether or not its numbers are right.

widgetspec describes an answer screen as data, checks that data against itself (totals, units, sources, formulas, every number quoted in the text, every slider at both ends) and renders it as an interface or as plain text.

Typical application. A calculator answer, live, with its receipt and its audit result underneath.
widgetspecDS-WS-01 · Product datasheetRev 1.0 · October 2026

4Functional block diagram

questionfrom the reader decide()text or which kind textwhen nothing calls spec (JSON)blocks + receipts validate()shape, paths audit()11 finding codes render()interface toText()plain text expr: formulas, compute()used by validate, audit and the sliders
Figure 4-1. A question goes to decide(); when an interface is called for, the answer is a spec that is validated and audited before it is shown or printed.

5Block kinds

KindShowsRequired fieldsdecide() listens forThe audit checks
textParagraphstextwhy, what is, explain, writing tasks; and anything with no strong signalevery number quoted is in some widget's data or in the question
compareItems side by side, a row per attributeitems, rows[label, values, unit]vs, versus, compare, difference between, which is bettertotals per column, units, source, empty cells
chartBar or line charttype, x, series[name, values]trend, over time, since a year, by year or region, chartunit, source, gaps, a line with fewer than three points
calcInputs you can move, outputs that followinputs[id, label, value, min, max], outputs[id, label, formula]how much … if, split, loan, per month, % off, calculateunknown names, unused sliders, inputs out of range, outputs that break at either end
timelineSteps in ordersteps[when, label, detail]plan a…, itinerary, schedule, roadmap, milestones, in ordershape only
checklistThings to tickitems[label, done]checklist, packing, what do I need, kit, before I leaveshape only

Every block may also carry title, source and note. The source and every formula travel with the block and are printed under it as its receipt.

widgetspecDS-WS-01 · Product datasheetRev 1.0 · October 2026

6Detailed description: try it

Everything on this sheet is the library itself, running in your browser. Ask a question to see which form decide() picks, or edit a spec and watch the audit.

6.1 Ask

6.2 Spec

6.3 Output

6.4 Audit

    widgetspecDS-WS-01 · Product datasheetRev 1.0 · October 2026

    7Typical characteristics

    Measured in your browser on the release's own data: 120 labelled questions (data/questions.json) and six example specs.

    Figure 7-1. What decide() picked for the 60 held-out questions. Rows: the form we labelled. Columns: the form it chose.
    Figure 7-2. Share of each kind found, on the questions the rules were tuned on and on the held-out ones.
    Table 7-1. Text or interface, and which kind
    MeasureTuning halfHeld-out half
    Table 7-2. Audit of the example specs
    SpecBlocksErrorsWarningsGrade
    widgetspecDS-WS-01 · Product datasheetRev 1.0 · October 2026

    8Formula language

    Table 8-1. Operators, tightest first
    OperatorMeaning
    ( ) f(a, b)grouping, function call
    ^power, right to left: 2 ^ 3 ^ 2 is 512
    - + !sign, not: -2 ^ 2 is −4
    * /times, divide
    + -plus, minus
    < > <= >= == !=comparison, gives 1 or 0
    && then ||and, or
    18% 1_000 1.5e3numbers: 18% is 0.18
    Table 8-2. Functions, each example evaluated live
    FunctionExampleResult

    8.1 Evaluate a formula

    9Absolute maximum ratings

    ParameterMaxBeyond it
    widgetspecApplication note AN-10110 October 2026

    Application note

    Polish is not proof: checking generated answer screens against themselves

    The widgetspec maintainers

    Abstract. Chat products now answer some questions with interfaces assembled from components [1, 2]. We describe such an answer as data and check it against itself: totals against their rows, numbers in the text against the widgets, sources and units, and every calculator at both ends of every slider. A rule-based chooser picks text or one of five interface kinds; tuned on half of 120 labelled questions, it gets 47 / 60 of the other half right, never adds an interface to a question we labelled as text, and falls back to text for the rest. Planting one fault at a time into five clean answers, the audit catches 24 of 24 faults of the kinds it checks and neither of two wrong numbers that nothing else in the answer refers to.

    1. Introduction

    On 7 October 2026 ChatGPT's Chat tab moved to GPT-6 with what OpenAI calls Intelligent UI: the model decides per question whether text or an interface answers it better, and builds the interface from a library of common elements, among them comparisons, diagrams, maps, calculators and forms [1, 2, 3]. A reader can move an input and watch a result change. One review put the risk plainly: polish looks like proof, and a calculator is only as accurate as its formula [4].

    A finished-looking screen invites trust it has not earned. We want the opposite default: an answer screen should arrive with its receipts (where each number came from, which formula made it) and should be checked before anyone sees it.

    2. What an answer must agree with

    We check four kinds of agreement, none of which needs to know whether a source is true.

    1. Arithmetic. A row labelled total equals the sum of the rows above it that share its unit.
    2. Grounding. Each number in the text (other than small counts and years) appears in a widget's data, in a calculator's result at its starting inputs, or in the question, allowing for rounding.
    3. Robustness. Each calculator output has a value at the start and stays finite when any one input is moved to its minimum or maximum, the places a reader is most likely to drag it.
    4. Provenance. Charts and comparisons with numbers carry a unit and a source; no slider is decorative.
    widgetspecApplication note AN-10110 October 2026

    3. Text or interface

    decide() scores a question for each kind with weighted patterns and answers in text unless an interface kind scores at least 3 and beats the text score. We wrote 120 questions, 20 for each form, and labelled the form we think fits best. Rules were tuned on the even-numbered half only. Table AN-1 shows what tuning on the same questions looks like next to the held-out result.

    Table AN-1. decide() on the labelled questions
    RightInterface chosen and rightInterface questions found

    Every held-out miss went to plain text: 13 of 13. The largest group are "A or B?" questions that never say "vs" or "compare". Falling back to text is the cheaper error: a reader loses a table, not the answer.

    4. Planting faults

    mutants() copies a clean spec and plants one fault, of a kind chosen from a fixed list, wherever the spec has a place for it; trial() counts it as caught when the expected finding appears more often than in the original. Two extra mutations change a number that nothing else refers to; they measure what the audit cannot see.

    Table AN-2. Planted faults across the five clean example specs

    5. Limits

    The rules read words, not meaning, and the questions are ours: few, in English and written by the people who wrote the rules. The audit checks that an answer agrees with itself, not with the world; a wrong interest rate used consistently passes. Only bar and line charts are checked. Choosing between interface kinds is measured only by agreement with our labels.

    References

    1. OpenAI, "GPT-6 for everyone," October 2026. openai.com
    2. "ChatGPT, Intelligent UI that creates screens based on questions," Design Compass, 8 October 2026. designcompass.org
    3. "OpenAI rolls out GPT-6 and Intelligent UI in ChatGPT," Let's Data Science, October 2026. letsdatascience.com
    4. "What is ChatGPT's Intelligent UI, and can you turn it off?" The AI Career Lab, October 2026. theaicareerlab.com
    widgetspecDS-WS-01 · Product datasheetRev 1.0 · October 2026

    10Software

    10.1 API

    CallReturns
    validate(spec){ ok, errors: [{ path, msg }] }
    audit(spec){ grade, errors, warnings, findings }
    decide(question){ kind, interface, reasons, scores }
    compute(calc, set){ values, errors }
    toText(spec)plain text
    toHTML(spec) / render(spec, el)HTML / a live answer in el
    expr.evaluate(src, names)a number
    mutants(spec) / trial(spec)planted faults / which were caught
    bench(questions, filter)accuracy, matrix, misses

    10.2 Command line

    $ widgetspec decide "compare tea and coffee"
    compare (an interface)
    $ widgetspec audit data/specs/broken.json
    FAIL: 4 errors, 5 warnings
    $ echo $?
    4
    $ widgetspec calc data/specs/dinner-split.json people=4
      Each person pays: $54.87
    $ widgetspec html data/specs/savings.json --out answer.html

    10.3 Run the tests

    unzip widgetspec-0.1.0.zip && cd widgetspec && npm test

    widgetspec-0.1.0.zipwidgetspec.jsREADME.md

    10.4 Source files

    FileLinesBytes

    Select a file to read it. The README is shown formatted.

    widgetspecDS-WS-01 · Product datasheetRev 1.0 · October 2026

    11Ordering information

    Orderable partPackagePrice24 hMarket capVolume 24 hLiquiditySupplyTax
    $WIDGETSPL token, Solana—————1,000,000,0000%

    Loading the chart…
    Figure 11-1. Package drawing: live price, 1-second candles. Open the full chart

    Device marking

    Where to get it