Evaluating Fairness in AI-Generated Emergency Alert Translations: A Quantitative Case Study of California IPAWS Messages in Spanish and Hindi
Public agencies increasingly rely on artificial intelligence (AI) based language technologies to disseminate emergency information to linguistically diverse populations. While automated translation systems enable rapid multilingual communication, concerns remain regarding whether such systems preserve fairness-critical communicative features across languages, particularly in high-stakes crisis contexts. This study evaluates fairness in AI-generated Spanish and Hindi translations of official emergency alerts issued through California’s Integrated Public Alert and Warning System (IPAWS). Using publicly available IPAWS alerts as the primary corpus, AI-generated translations were segmented and quantitatively coded using a fairness rubric encompassing procedural and interactional dimensions. A derived dataset was analyzed using descriptive statistics, Welch’s independent-samples t-tests, effect size estimation where appropriate, and one-way ANOVA to assess cross-language differences. Results indicate statistically significant disparities between Spanish and Hindi translations in procedural fairness and overall fairness, with large effect sizes. The findings demonstrate that lexical accuracy alone is insufficient to ensure equitable emergency communication and underscore the need for domain-specific fairness evaluation frameworks in public-sector AI deployments.
