HomeCensorshipAutomated Moderation Is Here to Stay—Accountability Must Keep Pace

Automated Moderation Is Here to Stay—Accountability Must Keep Pace

When whistleblower Frances Haugen leaked a set of documents from Meta in 2020, among the revelations was a jarring statistic: The company’s algorithms designed to detect terrorist content incorrectly deleted nonviolent Arabic-language content 77 percent of the time, while failing to detect hate speech under the company’s own policies in many instances. Meta’s own transparency report released later that year demonstrated similar findings. Five years later, researchers in the region report that overzealous moderation remains a problem, while paths to remedy have all but collapsed.

Where these systems are faltering in Arabic, they’re positively failing in less-resourced languages. As a 2025 report from the Center for Democracy and Technology found, labeled datasets in certain languages and dialects such as Maghrebi Arabic and Kiswahili contain inconsistencies, bias, and inaccuracies due to the limited hiring of annotators who actually speak the languages as well as shifts in the languages themselves.

But language disparities are just one of several concerns as automated moderation becomes more widespread. From the systemic suppression of content from Palestine to the repeated misclassification of LGBTQ+ content as adult or explicit material, these varied examples demonstrate the risks of overreliance on automated moderation—and the need for stronger safeguards.

Transparency, Cultural Competence, Appeals

Automated systems can process content at a scale that humans never could, potentially enabling better moderation at scale and alleviating the psychological load on ill-paid moderators whose jobs require them to view incredibly disturbing content. But automated systems also reproduce existing biases, struggle to understand context, and often make mistakes that disproportionately affect journalists, activists, artists, and other vulnerable and marginalized communities.

Despite those intrinsic flaws, there is a great deal companies, policymakers, and civil society can do to help ensure that highly-automated systems operate in ways that respect human rights, minimize predictable harms, and provide meaningful accountability when they fail. That evolution can start with committing to the Santa Clara Principles 2.0, which reflect the needs and expectations of the global community and specifically address automation.

Drawing on the Santa Clara Principles 2.0, international human rights standards, and years of research documenting the shortcomings of automated moderation, EFF proposes recommendations for policymakers and companies: automated technologies should help, not replace, human moderators; companies must be transparent about when and how automation is used in content decisions; companies must regularly audit their automated systems for bias, with particular attention to low-resource languages and vulnerable communities; users must have the ability to appeal decisions, with appeals promptly evaluated by human moderators; and lawmakers should avoid promoting legislation that mandates automated moderation systems or dictates platforms’ technical and design choices.

Because content moderation shapes public discourse and fundamental rights, its design and oversight must respond to the concerns of policymakers, civil society, independent researchers, and the communities most affected by these systems.

This article was originally published by the Electronic Frontier Foundation (EFF). Republished under Creative Commons CC BY 4.0. Read the original article: https://eff.org/deeplinks/2026/07/part-2-automated-moderation-here-stay-accountability-must-keep-pace.

Get the weekly briefing

Five things worth your attention — censorship, privacy, algorithms and digital rights. One email a week, no noise.

We’ll send you a confirmation email first. No tracking, no sharing, unsubscribe in one click. See our Privacy Policy.

Must Read

spot_img