DUMMYEVIDENCE INTELLIGENCE
SHADOW · READ ONLY
The evidence-gated intelligence loop

DUMMY THINKS

Forty-five independent loops turn point-in-time public evidence into competing probabilities, force disagreement into the open, learn only from settled outcomes, and keep research intelligence structurally separate from authority over capital.

The Dummy loopback-only operator board
The Organism visualizes persisted evidence. The DOM truth ribbon remains authoritative; the page is loopback-only, GET-only, and incapable of placing an order.
46
competing sources
45
autonomous loops
8
cycle phases
11
sports challengers
7
sports leagues
3
crypto scopes
2
dedicated crypto loops
4
LLM roles
0
automatic authority

Capability counts describe registered repository components, not profitability, deployment readiness, or permission to trade. Current launch status remains NO-GO until every external evidence gate is satisfied.

Visible by design

Two system maps, loaded as standalone assets

These diagrams sit near the top of the release and open at full resolution. They are ordinary SVG image assets—not script-rendered diagrams—so GitHub Pages, reduced-motion browsers, screen readers, and static mirrors all receive the same architecture.

Complete public capability catalog

Every major ability—and the surface that proves it exists

This catalog is intentionally broader than a feature list. Each ability names its owning system and its proof boundary, so research tooling, operational visibility, and execution authority cannot blur together.

Observe

Market discovery & normalization

Kalshi and public context feeds become timestamped, deduplicated market identities with provenance and freshness.

proof: observation ledger + source health
Crypto · Observe

Multi-timeframe chart visualization

BTC, ETH, and SOL candlesticks across 15m, 1h, 4h, 1d, and 1w with deterministic indicators and pattern markers.

proof: immutable Market Observer chart bundles
Crypto · Loop

Crypto paper twin & horizon evidence

Dedicated 5-minute paper-twin and 10-minute horizon-evidence tasks grade asset × timeframe × strategy lanes.

proof: forward paper and horizon artifacts
Sports · Observe

History & play-by-play lakes

Seven league surfaces combine point-in-time games, box scores, EPA, period state, comeback matrices, and player context.

proof: source manifests + event-purged lakes
Reason

Forecasting & simulation

Market anchors, statistical kernels, power ratings, scoring distributions, live models, and scenario simulators emit attributed probabilities.

proof: preregistered source forecasts
Reason

Specialist council & Model Arsenal

Vertical specialists and four schema-bound LLM roles preserve dissent under exact routing and paid-call gates.

proof: redacted stored witnesses, zero auto authority
Calibrate

Trust, uncertainty & fusion

Brier, log loss, ECE/MCE, debiasing, contested-market scoring, and scope-specific uncertainty determine earned weight.

proof: settled contested forecasts
Evaluate

Walk-forward & backtest evidence

Temporal folds, event-cluster bootstrap, fees, liquidity, CLV, partial fills, and negative controls challenge every claim.

proof: predict-before-update reports
Allocate

Candidate portfolio allocation

Evidence-adjusted edge and settlement velocity divide one bounded pot across holdable, correlation-aware candidates.

proof: Σ grants ≤ pot and grant ≤ ask
Constrain

Risk governor & execution firewall

Stage, drawdown, clusters, liquidity, TTL, sealed caps, session, credential, and LIMIT-only checks can only shrink intent.

proof: typed blocks + transport witnesses
Truth

Reconciliation, settlement & audit

Orders, fills, cancellations, outcomes, corrections, and account snapshots remain separate evidence layers.

proof: append-only identities + corrections
Learn

Autoresearch, tuning & evolution

Strategy mining, parameter tuning, quality-diversity search, true crossover, ablation, chaos, and fragility create challengers.

proof: observational until promoted by evidence
Question

Self-scout & film room

Bias tendencies, worst-call reconstruction, recruiting, matchup lens, top threat, and development tracking expose blind spots.

proof: report failures surface as failures
Operate

Watchdog, healing & durability

Independent task health, self-heal, snapshots, ledger retention, pruning, vacuum, and allowlisted rotation contain failures.

proof: cadence-bound artifacts + bounded storage
Explain

Operator Board & desktop notifier

Overview, scopes, charts, Arsenal, glossary, readiness, health, and freshness render persisted truth without mutation controls.

proof: loopback-only GET surface
Expose

Market Observer MCP

Allowlisted public candles, indicators, patterns, chart bundles, network status, and source health through read-only tools.

proof: false execution / order / allocation authority
Dummy operator Board showing all sixteen public-release capability families
All abilities on the Board — the same sixteen capability families are now visible inside the read-only operator UI, including market perception, both named crypto loops, chart visualization, settlement memory, metacognition, the firewall, fleet reliability, and Observer MCP.
Intelligence is a system property

Six capabilities. One inspectable memory.

Dummy does not call one model “the intelligence.” Its intelligence comes from preserving provenance, making probabilistic claims before outcomes are known, recording dissent, grading every eligible claim, and constraining what learned confidence is allowed to change.

Perception

Timestamped market, venue, macro, weather, roster, schedule, and game-state observations enter with source, freshness, and market identity attached.

Stale, missing, or ambiguous evidence becomes an abstention—not a synthetic fact.

Probabilistic reasoning

Market anchors, statistical kernels, simulations, specialists, and quarantined language-model roles emit schema-bound probabilities at explicit horizons.

Model confidence and fluent prose are not treated as calibration.

Dissent

Every source keeps its identity. The market, specialist panels, power ratings, and challengers can disagree without their evidence being flattened away.

Consensus is not manufactured by hiding the minority forecast.

Memory

Observations, forecasts, settlements, corrections, calibration, promotion dossiers, and execution witnesses are retained as separate evidence layers.

A narrative report cannot overwrite a transport witness or settled outcome.

Metacognition

Self-scouting, film-room reconstruction, ablation, drift checks, walk-forward replay, and fragility tests ask where the organism is wrong and why.

Thin samples, hindsight, and correlated events cannot masquerade as durable skill.

Constrained action

Allocation, risk, caps, freshness, session, and firewall gates form a monotone chain: each stage can shrink a proposal; none can expand upstream authority.

Research success, paper results, and model votes cannot authorize capital.
Capability map — intelligence with inspectable feedback Each capability reads or writes the same auditable evidence memory while keeping its own contract, uncertainty, and refusal boundary.
solid = forward evidence · dashed = learning
Dummy intelligence capability diagram Six capabilities form a feedback loop around an auditable evidence ledger: perception, probabilistic reasoning, dissent, constrained action, metacognition, and memory. None grants live authority. 01 · OBSERVE PERCEPTION provenance · freshness · identity missing → abstain 02 · PREDICT PROBABILISTIC REASONING models · simulations · horizons schema-bound forecasts 03 · CHALLENGE DISSENT market · specialists · challengers identity stays visible 06 · BOUND CONSTRAINED ACTION allocation · risk · firewall cannot mint authority 05 · QUESTION METACOGNITION self-scout · ablation · fragility thin evidence stays thin 04 · REMEMBER MEMORY claims · outcomes · corrections append-only evidence layers SHARED TRUTH PLANE AUDITABLE EVIDENCE LEDGER point-in-time · settled · correctable INTELLIGENCE MAY RECOMMEND · ONLY OPERATOR AUTHORITY MAY PERMIT
What it prices

The arsenal

Every market is priced by a purpose-built model, cross-checked by challengers, and gated on settled evidence before it can earn trust.

Crypto

BTC · ETH · SOL — 15m / hourly / daily / weekly
  • Realized & implied volatility, market and macro risk regime (S&P / DXY / VIX / 10y / gold / oil)
  • Momentum, technical, order-book and cross-venue evidence
  • Multi-timeframe structure, candlestick & divergence chartist, patience/confirmation entries

Sports

MLB · NFL · NBA · NHL · WNBA · NCAAF · NCAAMB
  • Winners, totals, spreads, team totals, YRFI/NRFI, first-five, halves, quarters, periods
  • Platoon splits, bullpen quality, rivalry awareness, plate-appearance simulation
  • Power-ratings ensemble (FPI/BPI, Elo, Massey, tie-aware Colley) + licensed multi-book de-vig
  • 162,915-game point-in-time lake → 11 walk-forward-tuned challengers: Glicko-2, MOV-Elo, Pythagenpat, Four Factors, EPA/play, scoring sigmas, rest shift, live win-prob, referee totals, player props
  • 32,298-game play-by-play knowledge base — empirical comeback matrices for live re-pricing
  • Player props via metered licensed consensus; correlation-aware parlay pricing

Autonomous improvement

the machine tunes itself
  • Two-stage promotion ladder — challengers earn their place from settled evidence
  • Evolution lab: quality-diversity archive, true crossover, jitter-fragility verdicts
  • Self-scout, film room, recruiting board, matchup lens, top-threat, development tracker
  • Fusion ablation, property-invariant sweeps, chaos drills, and sharp/sheep diagnostics retained as evidence

Council & debate

specialists + LLM voices
  • Per-vertical specialist subagents with coherence, CLV, and trust surfaces
  • Exact 4-model panel — Gemini 3.6 Flash, GPT-5.6 Luna, Claude Sonnet 5, GLM-5.2 — 7-call atomic, double-locked, graded per voice
  • Every specialist fails closed — missing data means abstain, never a degraded guess
Crypto is a first-class intelligence lane

Observe, chart, price, grade, and learn across five horizons

Dummy does not hide crypto inside a generic market loop. BTC, ETH, and SOL have dedicated observation, charting, paper-twin, and horizon-evidence paths, while shared forecast, calibration, risk, and research loops retain the asset and timeframe identity.

Dummy Crypto Research Charts showing BTC candlesticks, deterministic indicators, pattern markers, provenance, and no production authority
Crypto Research Charts — the actual artifact-only chart UI, shown with a clearly labeled synthetic release fixture. The production observer accepts rights-reviewed closed public candles; both paths compute deterministic RSI / EMA / ATR / MACD / Bollinger / stochastic / OBV facts, final-bar markers, age, provenance, and explicit false execution authority.
Market Observer READ-ONLY SIDECAR

Coinbase Exchange public candles → immutable content-addressed chart bundles. The dashboard never contacts the provider itself.

DummyCryptoPaperTwin EVERY 5 MIN

Runs BTC / ETH / SOL paper lanes for 15m, 1h, 1d, and 1w cohorts with isolated asset × timeframe × strategy evidence and no capital authority.

DummyCryptoHorizonEvidence EVERY 10 MIN

Builds forward horizon matrices and evaluates whether each crypto cohort has enough time-consistent evidence to remain a challenger.

Shared crypto intelligence ATTRIBUTED

ShadowPredator, MispricingMonitor's crypto-fast lane, recalibration, backtests, autoresearch, allocation, and risk preserve the crypto scope instead of merging it into sports evidence.

Research boundary: charts display facts, not price targets; paper twins record hypothetical decisions, not fills; horizon evidence can challenge a model, not authorize an order. The chart page is GET-only and all execution, order, cancel, amend, allocation, and promotion authority fields remain false.
The cognition cycle

Eight phases, each with a contract

Public evidence flows in one direction. Each phase produces a bounded artifact for the next one, so a failure can be located, replayed, and explained without pretending the entire organism is one indivisible model.

ScanSignalFuseAllocateRiskExecuteReconcileLearn
01 / OBSERVE

Scan

Discover supported markets and attach point-in-time public context, venue state, freshness, and provenance.

output → normalized observations
02 / PREDICT

Signal

Run every eligible source at the correct asset, league, market type, and horizon; preserve abstentions and reasons.

output → attributed probabilities
03 / CALIBRATE

Fuse

Debias each opinion and combine it by earned, scope-specific trust while retaining the market as a standing anchor.

output → contested consensus
04 / PRIORITIZE

Allocate

Rank evidence-adjusted edge by settlement velocity and divide one bounded pot across candidates that can actually be held.

output → capped candidate grants
05 / CONSTRAIN

Risk

Apply stage, drawdown, correlation-cluster, bankroll, liquidity, price, and time-to-live gates.

output → reduced order intent
06 / FIREWALL

Execute

Shadow by default. Any controlled live proof still requires explicit authority, an active session, sealed caps, fresh state, and LIMIT-only transport.

output → paper receipt or transport witness
07 / TRUTH

Reconcile

Match orders, fills, cancellations, market settlement, and account evidence without promoting local gate blocks into broker claims.

output → outcome-linked evidence
08 / ADAPT

Learn

Brier-score all eligible forecasts, refresh calibration, inspect failures, tune challengers, and append the next promotion evidence.

output → forward-only trust

The monotone rule: fusion cannot outvote calibration, allocation cannot exceed the candidate pot, risk cannot exceed the allocation, and the firewall cannot exceed sealed caps. Intelligence may recommend; only explicit operator authority may permit.

One pot, many candidates

Capital allocation

Sizing one order and dividing a budget across many are different problems. The risk brain solved the first — quarter-Kelly under a SHADOW → CANARY → RAMP → CRUISE ladder. Nothing solved the second: the top-ranked candidate took the entire remaining budget its caps allowed, a near-equal rival got nothing, and forty qualifying candidates sized identically to three.

Three split policies

explicit configuration workflow
  • kelly_prorata (default) — each ask scaled by weight; apportioned to fit only when the pot is oversubscribed. Thin nights deliberately deploy less than the cap
  • proportional — share of the pot by weight, clamped to the ask
  • top_k — fund the highest-weighted K in full, starve the tail

Weighted by proof, not promise

contested Brier advantage
  • Weight comes from how much better the model's Brier is than the market's own price on the same rows
  • Uses the lower 95% bound — the same quantity the promotion gate tests. Twelve lucky rows cannot size up on evidence the sample won't support
  • The weight floor is never zero: a scope that can never be allocated can never settle, never accrue evidence, and never earn its way up

Two rules learned the hard way

now property-tested
  • Weights are a contention tiebreak, not a discount. An early version scaled every ask by its weight unconditionally — a lone 100¢ ask from an unproven scope became 25¢, below one contract, and trading stopped silently. A change that sizes everything to zero looks conservative and is not safety
  • Divide across holdable slots, not evaluated candidates. A hundred are evaluated per cycle; a stage may permit five open markets. A hundred-way split puts every grant under one contract

Invariants

proven over generated inputs
  • Σ granted ≤ pot, always
  • granted ≤ ask, always — the allocator can only reduce, and every existing cap still binds behind it
  • Adding a candidate never raises an existing grant
  • The operator throttle can only shrink the pot; enlarging it past the sealed ceiling is a registration ceremony, not a slider
The autonomous fleet

45 loops, isolated by responsibility

Dummy is not one immortal process. Its intelligence is distributed across independent Windows scheduled tasks that survive reboots, pick up merged code on their next fire, and are individually observed. One slow report, broken feed, or league backfill cannot hold the cognition cycle hostage.

Cadence — the heartbeat

8 loops
  • The brain cycle, live poller, sports board refresh, mispricing monitor, DummyCryptoPaperTwin every 5 minutes, the vNext shadow organism, the USE sidecar, and account snapshots

Predator · LivePoller · SportsBoard · Mispricing · CryptoTwin · vNext · USE · AccountSnapshot

Learning & self-improvement

10 loops
  • Weight recalibration split into a fast weight-only core and heavy diagnostics, so a slow report can never stall a cycle — plus the tuner, strategy miner, autoresearch lab, simulation trainers, and DummyCryptoHorizonEvidence every 10 minutes

Weights · Backtest · Tune · SelfImprove · StrategyMiner · Autoresearch · Simulation

Sports data & grading

18 loops · per league
  • History backfill and event-purged walk-forward grading for all seven leagues, box-score ingest for three, and NFL EPA — each league on its own schedule so one outage cannot stall the rest

Lake × 7 · WalkForward × 7 · BoxScore × 3 · NFL EPA

Health & durability

5 loops
  • A fleet-wide watchdog that grades the other tasks rather than its own exit code, a five-minute self-heal loop, readiness reporting, and a snapshot writer so the board never holds a ledger lock

Watchdog · Healer · Readiness · Dashboard · DashboardSnapshot

Bounded footprint

4 loops
  • Ledger retention, pruning, a self-skipping vacuum, and log rotation on an explicit allowlist — state, audit and full-history tapes are never blind-truncated

Retention · Prune · Vacuum · LogRotation

Why independent loops matter

fault containment is intelligence infrastructure
  • Each cadence has its own timeout, evidence artifact, and next retry
  • The watchdog evaluates the fleet instead of trusting its own exit status
  • Read-heavy diagnostics cannot stall price formation or hold the ledger open
  • Recovery resumes from persisted artifacts rather than hidden process memory
Where the crypto loops live: DummyCryptoPaperTwin is one of the 8 heartbeat tasks; DummyCryptoHorizonEvidence is one of the 10 learning tasks. ShadowPredator, MispricingMonitor's crypto-fast pass, recalibration, backtests, and autoresearch also process crypto under preserved asset × timeframe identities. The two dedicated tasks are included in the total of 45, not added on top of it.
Loop fleet map — 45 isolated jobs, one evidence plane Five responsibility bands run on independent cadences and publish bounded artifacts. The board observes those artifacts; it is not a control path.
8 + 10 + 18 + 5 + 4 = 45
Dummy 45-loop fleet diagram Forty-five independent scheduled loops are divided into heartbeat, learning, sports data and grading, health and durability, and bounded footprint groups. They exchange persisted artifacts through an evidence plane observed by a read-only dashboard. 8 HEARTBEAT price · poll · refresh · snapshot 10 LEARNING recalibrate · tune · research · simulate 18 SPORTS DATA + GRADING lake × 7 · walk-forward × 7 · box · EPA 5 HEALTH + DURABILITY watch · heal · report · serve · snapshot 4 BOUNDED FOOTPRINT retain · prune · vacuum · rotate PERSISTED COORDINATION EVIDENCE PLANE observations attributed forecasts cycle receipts settlements + calibration promotion dossiers health + freshness audit + corrections 45 INDEPENDENT RETRY BOUNDARIES MARKETS + BOARD fresh candidates · ranked scopes TRUST + CHALLENGERS forward evidence · ablation · tuning POINT-IN-TIME LAKES league-isolated history + replay READINESS TRUTH fleet state · blockers · snapshots BOUNDED STORAGE hot path shrinks · audit stays intact READ-ONLY OBSERVER THE ORGANISM GET-only · no control edge SEPARATE AUTHORITY PLANE EXECUTION FIREWALL no scheduler can self-authorize NO DASHBOARD OR LOOP CONTROL PATH

The fleet can improve models, weights, calibration, and candidate selection. It cannot rewrite truth rules, lower promotion evidence, widen sealed caps, mint an operator signature, or grant itself a live session.

The command board

One hardened view over persisted runtime truth

The Dummy operator board is a loopback-only, read-only web board at 127.0.0.1:8787. It reads persisted artifacts and exposes no scheduler, configuration, authority, capital, model-provider, or broker mutation endpoint. Split-flap values change only when evidence changes, and a ⌘K palette jumps to any coin, league, model role, or saved theme.

The Organism

data-bound neural field · progressive enhancement
  • The four-model arsenal forms the cortex; sources and scopes orbit by tier
  • Pulses appear only when persisted evidence changes, never as fake activity
  • WebGL2, WebGL1, Canvas 2D, and static fallbacks preserve the same truth
  • Reduced motion stops animation without hiding evidence

Truth before spectacle

accessible DOM remains authoritative
  • Sticky status ribbon separates account freshness, authority, board, and system health
  • Forecast diagnostics are labelled as probability quality—not trading profit
  • Stored model witnesses are redacted and never refreshed by opening the page
  • Responsive scope docks retain keyboard and touch navigation
Updated Dummy Overview showing the Organism, truth ribbon, account state, and evidence diagnostics
Overview · current Organism buildUpdated UI capture — sticky execution truth, cached account state, system health, forecast diagnostics, and the data-bound neural field in one loopback-only operator surface.
Sports scope
Per-league scope — graded forecast quality, model vs book, live picks by edge, accuracy by bet type, and a day-by-day games breakdown
Crypto scope
Per-coin scope — the model beating the closing line, accuracy by bet type, every priced market by category, and today's settled calls
The evidence ladder

What the organism knows—and what that knowledge cannot do

Dummy keeps four layers deliberately separate. Moving upward requires new evidence; no lower layer is allowed to imply the layer above it.

LAYER 01

Forecast

A timestamped probability from an identified source, scope, market type, and horizon—made before the outcome.

Cannot prove calibration, edge, profitability, or authority.
LAYER 02

Settled evidence

Outcome-linked Brier, log-loss, closing-line, sample-size, and event-cluster evidence produced by forward grading.

Cannot bypass uncertainty bounds, replication, or promotion review.
LAYER 03

Promotion

A bounded dossier showing a challenger met preregistered evidence gates for a defined role and scope.

Cannot create broker credentials, sealed caps, or a live session.
LAYER 04

Authority

An explicit, operator-held, time-bounded control contract checked again at the central execution firewall.

Cannot be manufactured by code, UI state, paper results, or model consensus.
Non-negotiable

Safety invariants

  • Paper / shadow by default — no real capital at risk
  • Promotion to capital is human-only, never autonomous
  • Execution is fail-closed and LIMIT-only behind a hardened firewall
  • Challengers are evidence-gated on settled, contested outcomes
  • Risk caps are byte-sealed — code may register a new baseline but is structurally forbidden from manufacturing the operator approval needed to use it. Self-authorization is unrepresentable, not merely discouraged
  • Markets are deny-by-default — an order cannot form unless the market is positively authorized by exact ticker or by series. Series matching is boundary-aware identifier matching, never category inference; category compliance comes only from fetched venue metadata
  • Claims carry timestamps — a validated candidate's tradability expires, and a market never observed reports null rather than false. "Not tradable" and "we never looked" are different facts
  • Broker contact is claimed only on transport evidence. Two early proofs recorded as BROKER_REJECTED never reached a broker — no submit call, no endpoint, no payload, submit disabled throughout. Both were local gate blocks wearing a broker's label; the rejection classifier and truth layer exist because of them
  • Research may evolve; truth rules, promotion standards & the firewall may not
  • Every paper decision is explained and recorded in an auditable ledger