Personal build
Rotten Tom-AI-toes
- The question:
- How do we compare AI models on trust, beyond a single score?
- What I built:
- Built an evaluation wizard and comparison dashboards covering 6 trust dimensions and 17 foundation models. This is an exploration of evaluation workflows, not a validated clinical decision tool.
- Evaluation workflows
- Model comparisons
Evaluation workspace
Make the basis of trust visible.
Use case → Model selection → Evaluation review
- AccuracyReview evidence
- SafetyReview evidence
- FairnessReview evidence
- ExplainabilityReview evidence
- ComplianceReview evidence
- EfficiencyReview evidence