Conspiracy Watch | The Conspiracy Observatory
"Repetition does not transform a lie into a truth"
Franklin D. Roosevelt
Conspiracy Watch | The Conspiracy Observatory

Conspiracy Watch Joins Global Experts to Stress Test AI Systems for Antisemitism and Online Hate

Participants in the Artificial Intelligence systems ‘Red-Teaming’ event found it relatively easy to prompt several chatbots into generating dubious and even false and dangerous answers

Through its Technology and Human Rights Institute (TecHRI), the World Jewish Congress convened more than 45 experts for a two-day meeting.

Experts in antisemitism, online hate and related conspiracy theories who gathered in New York to stress test leading Artificial Intelligence systems and platforms found it did not take much to break their safety guardrails. By posing as an "objective" researcher or adopting a different identity, testers were able to induce chatbots to produce inconsistent, questionable and even false answers around Holocaust denial, Hamas terrorist attacks, antisemitic conspiracy theorists, and narratives of so-called “Jewish control”. In some cases the systems produced instructions on how to violently attack Jews. The experiment revealed striking differences depending on which language or persona was used in prompts.

Conspiracy Watch joined dozens of international experts at the two-day AI Red-Teaming Event. The gathering was organized by Technology and Human Rights Institute (TecHRI) of the World Jewish Congress. The initiative aims to combat antisemitism, bias, and discrimination in artificial intelligence systems.

More than 45 experts in antisemitism, online hate, technology, and digital rights took part, drawn from academia, civil society, and the technology sector around the world. Their goal was to stress-test leading AI systems: evaluating whether they generate hateful or harmful outputs, whether their guardrails can be easily circumvented, and what processes could help curtail the antisemitic hate speech and misinformation that is prevalent online.

The gathering opened with a conversation between CNN political commentator Van Jones and Anne Neuberger, former U.S. deputy national security adviser for cyber and emerging technologies. The two discussed the promise and challenges of rapidly advancing AI technologies, and the importance of ensuring such systems reflect democratic values and human rights principles.

Over the following day, participants worked collaboratively to "red team" major AI systems — subjecting them to tests designed to identify blind spots, inconsistencies, and vulnerabilities related to antisemitism and online hate, comparing findings and discussing trends and recommendations. The exercise found that sometimes the systems carried ingrained biases capable of significantly altering the information presented to users as "truth."

In one example, a simple change to the user's stated identity caused a chatbot to revise its answer about Hamas's use of sexual violence on October 7 — shifting from a clear "yes" to casting doubt on whether there was sufficient reliable information to support the claim. Participants also found that the sources an AI system drew from were critical to its output: contrary to the assumption that these systems draw evenly on the entire internet, a small number of prioritized sources could account for large portions of a response and skew its framing.

Disturblingly, participants found it relatively easy to manipulate several chatbots into breaking their own guardrails — for instance, by framing a query as coming from an "objective" persona, researchers were able to induce systems to produce content instructing users on how to violently attack Jews, or arguing that the Holocaust did not occur.

"AI systems are rapidly becoming arbiters of information for billions of people, but they are only as effective as our ability to teach them to recognize the complexities of hate," said Yfat Barak-Cheney, executive director of the WJC Technology and Human Rights Institute. "As more social media companies also turn to content moderation using generative AI, this has broader effects if not done right. It wasn't about proving AI is biased — it was about finding where the bias shows up, so we can do something about it. That is why bringing together experts to rigorously test these systems is not only a technical exercise, but a human rights imperative."

Barak-Cheney said she thought it would be hard to "jailbreak" the AI chatbots, and that “we need experts to recognize subtle nuances and think of ways to trick the chatbots into biased responses”.

“Turns out manipulating AI safety guardrails is easier than we think,” she said after the event. “Participants bypassed content filters with minimal effort — framing requests as debate prep, legal research, or a grandmother's bedtime story. The models complied, even serving instructions on how to cause the most harm during a shooting attack.

“We're not getting objective information. We're getting personalized bias. Just like social media algorithms, chatbots appear to serve users the answers they expect to receive. In one experiment, the same question about Hamas's actions on October 7th yielded opposing answers depending on the assumed identity of the user, presented simply by name change.

“We've seen this movie before with social media platforms, and we know how it ends. Some chatbots did better than others, which brings me back to my mantra for 2026 - Everything is a policy decision. AI companies can, and should, make better policy choices.”

Conspiracy Watch's Contribution

Conspiracy Watch writer and academic Dr. Stephanie Share, a researcher at the Elizabeth and Tony Comper Center for the Study of Antisemitism, focused her red-teaming work on evaluating how several AI systems respond to Holocaust denial and distortion.

Using the LLM-monitoring tool developed specifically for the project, participants were able to compare the outputs of multiple AI systems simultaneously, assess their responses, and identify areas for improvement.

Conspiracy Watch's own experiments tested how various LLMs handled questions around conspiracy theories such as the Great Replacement — evaluating both whether the systems correctly identified the theory as antisemitic, and the quality of the sources they cited in their answers.

The team also tested how AI chatbots handled a widely-viewed YouTube conversation between notoriously antisemitic online demagogue Candace Owens and Young Turks podcast host Ana Kasparian. The discussion drew criticism from Jewish advocacy groups and media watchdogs for advancing antisemitic tropes and conspiracy theories, including claims about a "Zionist lobby" influencing Owens's career and Holocaust films functioning as "brainwashing." Released in late May and viewed roughly 1.8 million times since, the video was used as a test case for whether AI systems could independently recognize antisemitic conspiracy content when prompted about it.

The results were fairly poor. In the United States, Gemini automatically offered on the right hand side of the YouTube screen, conspiratorial follow-up questions generated from the video's transcript – without any disclaimer flagging the antisemitic content discussed in the video. Most but not all of the other systems tested were unable to identify the antisemitic conspiracy theories present in the conversation when asked directly.

Building Better AI

The project remains ongoing, but its goal is to provide analyses and recommendations to major platforms to help develop AI tools that are more reliable, more equitable, and more resilient against hate speech and the manipulation of history.

Participants also discussed tracking how LLM responses shift over time in the immediate aftermath of a sensitive event — such as a shooting, a conflict, or a terrorist attack. Such an experiment could show how conspiracy theories thrive as a story first emerges and existing information sources are manipulated in real time. Proposed methodologies include measuring the percentage of LLM responses that cite conspiracy websites as sources, and creating a benchmark to measure the extent to which LLMs promote conspiracy theories in their generated responses.

The full findings are expected to be published in the coming months, with specific analysis shared with AI companies to help them improve their systems.

For sixteen years, Conspiracy Watch has been diligently spreading awareness about the perils of conspiracy theories through real-time monitoring and insightful analyses. To keep our mission alive, we rely on the critical support of our readers.

DONATE!
ABOUT THE AUTHOR
Emma-Kate Symons
Emma-Kate Symons
Emma-Kate Symons is a Paris-based journalist and columnist who has been published in The Washington Post, The Wall Street Journal, Foreign Policy, The Atlantic, The Financial Review and Reuters. A contributing editor at The New World, she is a regular contributor to French weekly Franc-Tireur and France 24 . Educated at the University of Sydney and Columbia University, Emma-Kate has reported from Europe and the Middle East, as well as from New York, Washington, Manila, Bangkok and Canberra.
ALL ARTICLES BY Emma-Kate Symons
SHARING:
Conspiracy Watch | The Conspiracy Observatory
Blue Sky
© 2026 An initiative of the Observatoire du conspirationnisme (nonprofit organization) with the support of The Foundation for the Memory of the Shoah.
Fondation pour la Mémoire de la Shoah
cross