Manual evaluation
These diagnostics do not score factual correctness. Each Run case button starts a full live research request using configured credentials and may incur API charges.
medical/scientific — What does research suggest about sleep and memory consolidation in healthy adults?
Expected topics: study design, population, limitations
Important sources: Systematic reviews and controlled studies
Review evidence quality manually; not medical advice.
technology — How do rooftop solar and small wind turbines compare in urban settings?
Expected topics: capacity factor, location, cost
Important sources: Field performance studies
Check geography and publication dates.
economics — What evidence links minimum wage increases to employment changes?
Expected topics: causality, effect sizes, regional differences
Important sources: Quasi-experimental studies and reviews
Look for conflicting findings and scope limitations.
general factual research — Why do seasons occur on Earth?
Expected topics: axial tilt, day length, hemispheres
Important sources: Authoritative astronomy sources
Academic search relevance may be low despite a straightforward question.
controversial/conflicting evidence — What are the benefits and limitations of intermittent fasting compared with continuous calorie restriction?
Expected topics: adherence, trial duration, confounding
Important sources: Randomized trials and systematic reviews
Check whether disagreements and uncertainty are disclosed.
little academic evidence — Does using a purple notebook improve productivity for freelance illustrators?
Expected topics: lack of direct evidence, assumptions
Important sources:
An explicit lack of evidence is appropriate; do not reward invented support.
Compare research runs
Latest 100 matching runs. Ratings are manual developer observations, never input to synthesis. Costs and latency include all attempts; models lists include historical attempts.
| Question / run | Models used | Latest judge | Cost USD | Latency ms | Active papers | Latest manual ratings |
|---|---|---|---|---|---|---|
| #1 hello COMPLETED | openai/gpt-6-astra, google/gemini-3.1-pro-preview, qwen/qwen3.8-max-0902 | x-ai/grok-4.20 | 0.0326941000 | 35399 | 20 | Not rated |
Run diagnostics
· Sept. 22, 2026, 6:21 p.m. · completed
hello- total_duration_ms
- 35399
- interrupted_execution_metrics_incomplete
- False
- successful_providers
- 4
- failed_providers
- 0
- successful_agents
- 3
- consensus_success
- True
- synthesis_success
- True
- paper_count
- 20
- total_estimated_openrouter_cost
- 0.0326941000
- known_cost_subtotal
- 0.0326941000
- attempts_missing_cost
- 0
- valid_source_percent
- 100.0
- removed_source_ids
- 0
- agreements_count
- 2
- disagreements_count
- 1
- evidence_strength_distribution
- {'limited': 2, 'moderate': 2}
- judge_confidence
- 80.0
- answer_length
- 42