The Dokimion · proof before praise
Evaluation evidence
No blended vanity score. Results are separated by evaluation class, dataset, sample and limitation.
Credibility gate: HOLD
The historical suite is only a structural smoke test. Representative agent quality and product outcomes have not been established.
Legacy structural suite
HOLDClass
Synthetic smoke test
Sample
4 historical single runs
Observed result
3 of 4 pass under v1 rescoring
Limitation
Structure only; not representative, repeated, outcome-based, or market-comparable.
Representative agent evaluation
NOT_EXECUTEDClass
Real-task evaluation
Sample
0 reviewed task sets
Observed result
No approved result
Limitation
Gold tasks, comparators, repetitions, uncertainty, and human calibration are pending.
Agentic security suite
NOT_EXECUTEDClass
Adversarial evaluation
Sample
0 attack suites
Observed result
Not executed
Limitation
Prompt injection, privilege, memory, leakage, supply-chain, and denial-of-wallet tests remain to be built.
PythiaGrid product benchmark
HOLDClass
Observable product outcome
Sample
0 consented design partners
Observed result
Not established
Limitation
No representative external user outcome or comparative product evidence yet.