Analysis & measurement

Instruments and their validity: model judges, platform panels, simulated agent societies. The longitudinal measurement campaign here is the weekly design's methodological precedent.

A4 · 1 of 5 campaigns have entries ingested here · 13 of 91 entries

Entries are repositories — libraries, benchmarks, specifications — admitted as evidence, not as products.

Tasks in this category

Measure platforms longitudinally

Campaigns placed on this task

  • 2026-08-03__platform-ecosystem-longitudinal-measurement

    run type (the campaign's own designation) high-recall-map 13 of 91 entries ingested

    Ledger rows
    42
    Screening rows
    96
    Repos in ledger
    16

Entries — 13 of 91 on this hub

Tasks with no entry today

6 of 7 tasks in this category carry no entry today. Each is listed with its state: ingested with no repository evidence in its ledger, on the record and not yet ingested, or planned and not yet run. Opening a row shows the campaigns placed on it, with their ledger and screening rows as recorded.

  • Calibrate model judges Not yet ingested 1 campaign on the record; one modern ledger has not been brought into the registry.

    Campaigns placed on this task

    • 2026-07-30__multi-model-contract-calibration

      run type (the campaign's own designation) high-recall-map Not yet ingested

      Ledger rows
      36
      Screening rows
      56
      Repos in ledger
      18
  • Simulate agent societies & markets Not yet ingested 2 campaigns on the record; one modern ledger has not been brought into the registry. The other predates the current standard; it carries flags, not counts.

    Campaigns placed on this task

    Multiple runs of this task exist; each is its own sealed record — no supersession is implied.

    • 2026-07-27__agent-markets-societies-simulations

      run type (the campaign's own designation) pre-standard Not yet ingested

      Ledger rows
      Screening rows
      Repos in ledger
      0

      Coverage ceiling, quoted from the sealed record:

      SCOPE.md frozen before external discovery with comprehensive-continuation addendum; coverage claim: effort-bounded, v2 mechanism-class stop rule satisfied
    • 2026-07-30__interactive-gamified-agent-worlds

      run type (the campaign's own designation) high-recall-map Not yet ingested

      Ledger rows
      28
      Screening rows
      100
      Repos in ledger
      31
  • Validate LLMs as instruments Not yet ingested 1 campaign on the record; one modern ledger has not been brought into the registry.

    Campaigns placed on this task

    • 2026-08-03__llm-as-measurement-instrument-validity-drift

      run type (the campaign's own designation) high-recall-map Not yet ingested

      Ledger rows
      48
      Screening rows
      106
      Repos in ledger
      4
  • Preregistration & power Planned No campaign has run yet. Grounded in crosswalk row C28 (internal crosswalk id).

  • Psychometrics Planned No campaign has run yet. Grounded in crosswalk row N03 (internal crosswalk id).

  • Missing data Planned No campaign has run yet. Grounded in crosswalk row N04 (internal crosswalk id).