Content moderation determines what billions of people can say and see online. For years, platforms relied on human moderators to enforce their rules, and increasingly, on automated moderation. With the rise of LLMs, powerful generative AI systems are now used for tasks ranging from automated content removal to behaviour detection and user notification. Yet this shift has been happening largely without scrutiny, transparency, and without any known systematic assessment of what it means for human rights.
For civil society, activists, journalists, and marginalised communities, the promises of LLM moderation are appealing, but the risks are also alarming: content related to protests and dissident discourse is often underrepresented in the datasets LLMs are trained on. At the same time, the moderation power is concentrated in a small number of foundational AI models. Most platforms use pre-built LLMs rather than building their own. As a result, a handful of foundational models shape the public discourse at scale, with almost no democratic accountability.
ECNL responded to this challenge with a comprehensive, interdisciplinary research effort. We reviewed over 200 computer science and related academic papers and conducted in-depth legal analysis under international human rights law.
Our Algorithmic Gatekeepers research is structured around the specific rights at stake and the real-world harms LLM moderation poses to them. ECNL's analysis finds that the most promising uses of LLMs in moderation are in procedural safeguards: assisting human moderators, triaging complex cases, generating clearer explanations of moderation decisions, powering faster appeals, and identifying systemic errors. Used well, these systems could make content governance more transparent and fair; but used as blunt removal tools, they pose serious threats to civic freedoms.
The research has since been used by industry leaders and civil society, thereby informing both practice and policy. On the policy side, ECNL continues to advance regulation of advanced AI systems through its work on the EU AI Act implementation, the General-Purpose AI Code of Practice, and broader international standard-setting processes at the Council of Europe, the United Nations, and OECD working groups. The research directly informs these advocacy efforts, helping policymakers with rights-grounded understanding of how LLMs work.
The difference this work made:
- Established the first comprehensive human rights analysis of LLM-based content moderation, helping define how these systems should be understood from a rights perspective across policy, industry, and civil society.
- Shifted the debate on AI-driven content moderation away from efficiency narratives towards rights-based concerns.
- Strengthened the case for using these systems to support human decision-making rather than replace it.
- Informed ongoing policy and standard-setting discussions at EU, Council of Europe, UN, and OECD level, ensuring that emerging regulation reflects the specific risks of LLM-driven moderation systems on civic freedoms.
Image: Hanna Barakat & Archival Images of AI + AIxDESIGN / https://betterimagesofai.org /