Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Are Safety Failures Model-Specific or Prompt-Universal? Comparing Hallucinations and Refusals Across LLMs
A comparative study of GPT‑4, Claude, and open‑source LLMs, analyzing hallucinations, refusals, and Chain‑of‑Thought prompting effects on safety and disagreement signals.
This project explores how different large language models—such as GPT-4, Claude, and open-source models like Mistral—respond to the same prompts, particularly in terms of hallucinations and refusals. By comparing these behaviors across architectures, we aim to uncover whether safety failures (like confidently wrong answers or unjustified refusals) are inherent to specific models or more prompt-dependent
The study also examines how Chain-of-Thought (CoT) prompting affects safety: Does asking models to “think step by step” reduce hallucinations or actually worsen overconfidence? Additionally, a test can be done to check whether disagreement between models can act as an alignment signal — a simple mechanism to flag risky outputs.
( I started working on this project a few months ago as part of a paper idea for a AI alignment conference but couldnt complete the study. The experiment is currently in progress, and clear results, visualizations, and takeaways will be ready the next 1–2 weeks.)
Compose Email
Loading recent emails...