Internal · Editorial operations
AI Content Operations Dashboard
Evaluations
Demo mode: Demo scores come from a deterministic heuristic harness (real overlap and structure analysis, not random numbers). They are advisory inputs to review, not validated production benchmarks.
Scenarios:
No result yet.
Try the “Unsupported claims” scenario — safety and tone dimensions should flag it.
Dimension methodology
- Grounding: lexical overlap between draft language and supplied context.
- Relevance: share of sentences addressing the queried terms.
- Completeness: length and structural development vs task norm.
- Tone: penalises unhedged absolutes and intensifiers.
- Safety: injection-pattern screen — hard FAIL on detection.
- Accessibility: heading structure and length heuristics.
- Source coverage: citation markers across supplied sources.