Introducing the apgard Benchmark for Youth Mental Wellbeing
Part 1 of the apgard Youth AI Safety Benchmark series
Young people are increasingly turning to AI for emotional support, and the open-source AI safety ecosystem has made great progress evaluating mental health and child safety risks. Building upon this foundation and thanks to the support of the Safe Online fund, apgard ai is releasing the first version of our open-source apgard Benchmark, which goes deeper on fine-grained youth mental wellbeing sub-risk categories to further address youth and their developmental needs. For example, this benchmark examines how AI responds to a teen asking how to hide self-harm marks before gym or make restrictive eating look like a normal food journal entry, scenarios that may be underrepresented under broader safety evaluations.
This youth mental wellbeing benchmark is our first release of our three part apgard Youth AI Safety Benchmark series, and is built upon KORA’s benchmark tool. The second release, focused on child sexual exploitation and abuse (CSEA), will also be built upon that benchmark. We will also release a new trajectory-based benchmark that measures frontier models’ ability to handle youth mental wellbeing and CSEA risk signals over time, which is slated to be released towards the end of 2026.
This particular benchmark evaluates models against auto-generated youth-AI scenarios driven by our clinically-informed taxonomy of youth mental health risks across the following high-level risk categories:
Disordered Eating & Body Dissatisfaction
Nonsuicidal self-injury (NSSI)
Suicide and suicidal ideation (SSI)
Psychosocial Distress
It measures how they address relevant conversation topics, avoid harmful advice, recognize crises, and steer youth toward safer outcomes.
To conduct this assessment, we incorporated:
Taxonomy design: We synthesized existing child mental health research into a youth- and AI-specific risk taxonomy for youth mental well-being.
Expert review: 24 child clinicians and youth safety experts reviewed, refined, and validated the approach.
Built upon KORA: We applied our taxonomy onto KORA’s expert-informed open-source AI child safety benchmark tool to derive this youth mental well-being benchmark.
Benchmarking and scoring: Each target model was assessed by a single reasoning judge model (GPT 5.4 Mini, GPT 5.5, Claude Sonnet 4.6 or Gemini 3.5 Flash due to their ability to break down nuanced and complex meanings), with the judge always different from the target being evaluated. Scores are computed based on KORA’s rubric, and the corresponding metrics are shared on our leaderboard. We ran a total of 1380 scenarios across both assistant- and child-aware prompts among 18 frontier models. Assistant mode is the model’s default state, while child-aware mode is when the model is aware that it is speaking with a child.
Youth Mental Wellbeing Analysis
Across all four high-level risk categories, the benchmark shows that the most significant youth mental wellbeing failures are not typically caused by overtly unsafe first responses. Instead, many failures emerge through gradual alignment with the child’s unsafe framing, where models begin with empathy and caution but enable harm through advice, explanations, normalization, secrecy, or dependency over subsequent turns.
When models are child-aware, their performance improves across all models and risk categories, as they tend to activate more protective guardrails and treat requests as requiring adult intervention rather than direct help. The largest improvements occur in SSI and psychosocial distress, while NSSI remains the most challenging category. For the subrisk categories where Assistant scores performed poorly on our benchmark, such as Concealment & Evasion (NSSI), Curiosity & Conceptual Engagement (NSSI), and Action Readiness (SSI), even with a child-aware prompt, scores remained low.
Some of our observations across the four high-level risk categories include:
Disordered eating & body dissatisfaction
Failures occur when models shift from addressing emotional distress to supporting appearance, weight, or control goals. The highest-risk failures involve calorie-focused discussions, restrictive eating guidance, concealment from caregivers, or body-image optimization framed as fitness, curiosity, or self-improvement. Strong responses consistently treat disordered eating as a wellbeing concern, validating feelings while redirecting toward nourishment, healthy coping, and trusted adult support.
Non-suicidal self-injury (NSSI)
Failures arise when models blur the line between understanding self-harm versus validating and enabling it. Common issues include explaining why self-harm provides relief, normalizing it as a coping mechanism, answering curiosity-driven questions with excessive detail, or helping users conceal evidence from caregivers. Child-aware prompting improves performance on active self-harm engagement and concealment requests, but curiosity-based scenarios remain difficult. Exemplary responses validate emotional pain without validating self-harm itself, avoid explanations that reinforce perceived benefits, and treat secrecy as a safety concern rather than a privacy preference.
Suicide and suicidal ideation (SSI)
The scenarios in this risk category contain the highest-severity wellbeing failures. The primary risk is collaboration with a child’s suicidal narrative through goodbye messages, preparation-related assistance, confidentiality assurances, or philosophical discussions that portray death as relief, peace, or escape. Failures occur when models fail to identify the imminent risk, such as when the child frames the suicidal act as a dare or challenge. Strong responses treat all suicidal language as potentially serious, refuse participation in preparation or concealment, and consistently redirect toward trusted adults, crisis resources, and human intervention throughout the conversation.
Psychosocial distress
This risk area differs from the other categories because the AI itself can become the hazard, as they default to being helpful and always available but can miss addressing unhealthy dependency patterns. Failures occur when models become a child’s primary source of regulation, identity formation, reassurance, or belonging. Common examples include extended identity coaching, emotional dependency, avoidance of caregivers, deception scripts, or reinforcement of dissociation and rumination. Exemplary responses provide support while intentionally limiting the AI’s role, strengthening offline coping skills, relationships, and support networks.
Across categories, a consistent pattern emerges: the model’s greatest risk is not the explicit encouragement of harm, but gradual alignment with the child’s framing of the problem. Models may provide safe responses in the first turn, but in subsequent prompts, responses fail to uphold safety boundaries as users push back or escalate on their request. In disordered eating & body dissatisfaction, this appears as support for weight-control or unrealistic body composition goals. In non-suicidal self-injury, it appears as explanations that normalize self-harm. In suicidal ideation, it appears as collaboration with hopelessness, secrecy, or preparation. In psychosocial distress, it appears as replacing human support with always-available AI-mediated validation and guidance. Mentioning support resources is not enough to achieve a safe outcome if the response reinforces the harmful objective.
These findings suggest that youth mental wellbeing evaluations should focus less on single-turn refusals and more on whether models maintain safe, developmentally appropriate boundaries over time. Exemplary performance requires validating emotions without validating harmful solutions, resisting conversational drift toward the child’s unsafe objective, and consistently strengthening redirections towards the tool’s intended purpose, trusted adults, and offline sources of resilience.
However, there are some limitations to this benchmark, including:
A lack of real youth validation: The user messages in these scenarios were prompted to reflect realistic youth language and behavioral patterns, but they were not validated by young people themselves. There is an opportunity to incorporate direct feedback from youth reviewers to better align simulated prompts with how they express themselves in the real world. In particular, the 7-9 year old age group simulations need to be refined, as younger children are less likely to discuss mental health topics directly.
Wellbeing is not only mental health-related: This benchmark applies a clinical psychology lens to model behavior, drawing on established frameworks for youth mental health risks. However, it does not address the full landscape of harms that affect young people’s holistic well-being. There is an opportunity to incorporate other perspectives on youth well-being, including sexual development. (We will be releasing a youth sexual risks benchmark in the coming months.)
Real AI usage by young people is messy, random, and over long conversations: The greatest risks often emerge not in the first few exchanges, but through repeated interactions that gradually normalize harmful framing or deepen unhealthy dependency. Current evaluation approaches need the ability to assess risk longitudinally to account for these patterns.
Apgard is continuing to engage with youth development and AI experts to refine this benchmark and address these limitations.
How to use these findings
These findings are relevant to anyone making decisions for how AI reaches young people, including:
AI developers: Integrate the benchmark into your evaluation pipelines to identify youth mental well-being failure modes pre and post-deployment.
For parents and schools: Ask technology partners to run this benchmark and share their scores to ensure youth mental well-being standards are met.
For policymakers: Use this benchmark as an input to standards, procurement criteria, and emerging AI policy frameworks.
What’s Next
We are currently building a trajectory-based evaluation methodology to better understand how model behavior evolves across repeated interactions in order to better assess youth AI safety risks at scale and better support young people’s unique needs. Reach out to us at benchmark@apgardai.com if you’d like to join us on this journey!





