Blind Spot
A benchmark proposal for using AI agents to surface known and novel social media harms
What this is
Blind Spot is a benchmark proposal for evaluating whether AI agents can help identify known and novel social media harms while accounting for the gap between what systems recommend and how users actually behave.
Benchmark focus
- Known harms: categories that already appear in safety taxonomies or platform policies
- Novel harms: emerging patterns that are harder to predefine before measurement
- Behavioral gaps: differences between recommendations, user responses, and downstream exposure
What it measures
The benchmark uses agentic summaries, advice, and recommendation judgments as structured signals, then separates caution-worthy outcomes from neutral or benign recommendations.
Why it matters
Safety evaluations often focus on harms that are already easy to name. Blind Spot asks where AI-assisted auditing can reveal what existing labels miss.