Reading the sheet…
Reading the sheet…
Two live sources, each labelled with its own status. Telltale does not publish a score for any model it did not grade from a transcript a reviewer supplied, and it never presents fallback data as current.
| Model | Provider | Licence | Framework |
|---|---|---|---|
| google/gemma | Gemma | pyTorch | |
| google/gemma-2 | Gemma | pyTorch | |
| google/gemma-3 | Gemma | pyTorch | |
| google/gemma-3n | Gemma | transformers | |
| google/gemma-4 | Apache 2.0 | transformers | |
| keras/gemma | keras | Gemma | keras |
| keras/gemma2 | keras | Gemma | keras |
| keras/gemma3 | keras | Gemma | keras |
| cohereforai/aya-vision | cohereforai | Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) | transformers |
| deepseek-ai/deepseek-r1 | deepseek-ai | MIT | transformers |
| metaresearch/llama-3 | metaresearch | Llama 3 Community License | pyTorch |
| metaresearch/llama-3.1 | metaresearch | Llama 3.1 Community License | pyTorch |
| metaresearch/llama-3.2 | metaresearch | Llama 3.2 Community License | pyTorch |
| mistral-ai/mistral | mistral-ai | Apache 2.0 | pyTorch |
| mistral-ai/mistral-small-24b | mistral-ai | Apache 2.0 | transformers |
| tatsu-lab/alpaca | tatsu-lab | CC BY-NC-SA 4.0 | pyTorch |
| deepseek-ai/deepseek-r1-0528 | deepseek-ai | MIT | transformers |
| mistral-ai/devstral-small-2505 | mistral-ai | Apache 2.0 | transformers |
| qwen-lm/qwen-3 | qwen-lm | Apache 2.0 | transformers |
| qwen-lm/qwen-3-5 | qwen-lm | Apache 2.0 | transformers |
| qwen-lm/qwen-3-vl | qwen-lm | Apache 2.0 | transformers |
| qwen-lm/qwen3-coder | qwen-lm | Apache 2.0 | transformers |
| qwen-lm/qwen3-next-80b | qwen-lm | Apache 2.0 | transformers |
| qwen-lm/qwen3-tts | qwen-lm | Apache 2.0 | transformers |
| domdejonge/essay-detection | domdejonge | BSD-3-Clause | pyTorch |
| mistral-ai/devstral-small-2507 | mistral-ai | Apache 2.0 | transformers |
| mistral-ai/magistral-small-2506 | mistral-ai | Apache 2.0 | transformers |
| mistral-ai/magistral-small-2506-gguf | mistral-ai | Apache 2.0 | gguf |
| mistral-ai/magistral-small-2509 | mistral-ai | Apache 2.0 | transformers |
| mistral-ai/mistral-small-3.1 | mistral-ai | Apache 2.0 | transformers |
| alpie/alpie-core | alpie | Apache 2.0 | transformers |
| deepseek-ai/janus-pro | deepseek-ai | Other (specified in description) | pyTorch |
| huikang/deepseek-r1 | huikang | Apache 2.0 | transformers |
| qwen-lm/qwq-32b | qwen-lm | Apache 2.0 | transformers |
| shelterw/deepseek-r1 | shelterw | Apache 2.0 | transformers |
| ai21labs/ai21-jamba-reasoning-3b | ai21labs | Other (specified in description) | transformers |
| embeddedravi/mcq-generator-phi-3-4k-tuned-with-50k-lora-weights | embeddedravi | CC BY-NC-SA 4.0 | keras |
| eyppler/lbai-1 | eyppler | MIT | transformers |
| ibm-research/granite-speech | ibm-research | Apache 2.0 | transformers |
| jagatkiran/microsoft-phi-4 | jagatkiran | Other (specified in description) | transformers |
| sayedsalem/unslothphi-4-unsloth-bnb-4bit | sayedsalem | Apache 2.0 | transformers |
| shelterw/qwen2.5 | shelterw | Apache 2.0 | transformers |
| ashok205/gpt-oss-120b-uncensored | ashok205 | Apache 2.0 | transformers |
| danielhanchen/gpt-oss-120b | danielhanchen | Apache 2.0 | transformers |
| danielhanchen/gpt-oss-20b | danielhanchen | Apache 2.0 | transformers |
| huikang/gpt-oss-120b-aimo3 | huikang | Apache 2.0 | transformers |
| llkh0a/gpt-oss-20b-gguf | llkh0a | Apache 2.0 | pyTorch |
| reyvan14/gpt-oss-sft-aimo3 | reyvan14 | Apache 2.0 | transformers |
| shelterw/gpt | shelterw | Apache 2.0 | transformers |
| yeoyunsianggeremie/nvidia-gpt-oss-120b-eagle3 | yeoyunsianggeremie | Apache 2.0 | pyTorch |
https://github.com/aniruddhaadak80/telltale and read the task definition in benchmark/telltale-hold.npm run bench:prepare to write the case file, or pass the script id straight to the Kaggle runner.node scripts/run-kaggle.mjs --script queue-latency --models google/gemma-3,qwen-lm/qwen-3 with your Kaggle credentials.Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users' expressed views. In reality, sycophancy rarely happens in a single exchange; it may emerge organically as users repeatedly insist or subtly steer the dialogue over time. Current…
2609.39863v1 · 2026-09-30 · Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth et al.
Large language models (LLMs) can influence people's beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target's beliefs more extreme. We examine radicalization along two pathways:…
2609.38296v1 · 2026-09-29 · Ozgur Can Seckin, Shalmoli Ghosh, Alessandro Flammini et al.
Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap…
Kaggle. Model catalogue from the Kaggle public API (https://www.kaggle.com/api/v1/models/list), used under the Kaggle Terms of Use. Licences are reported as Kaggle lists them and are the reviewer's responsibility to check.
arXiv. Paper metadata from the arXiv Atom API (https://export.arxiv.org/api/query). Metadata is supplied by arXiv contributors under CC0-style terms; the papers themselves remain under their authors' licences. Abstracts are truncated.
2609.37616v1 · 2026-09-29 · Abhinav Rajeev Kumar, Paras Chopra
Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit…
2609.35544v1 · 2026-09-28 · Xu Wang, Difan Zou, Xuansheng Wu
Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocates a wrong answer and then withdraws that pressure. We introduce a recovery-after-pressure protocol for multiple-choice…
2609.33672v1 · 2026-09-27 · Adi Shnaidman
Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while principle-based methods specify high-level values without fully operationalizing them. To address this…
2609.33086v1 · 2026-09-27 · Johann D. Gaebler, Calvin Isley, Max Lamparth et al.
Sycophancy is a language model's tendency to cave when a user pushes back, abandoning a correct answer for the user's. Several benchmarks now measure it by scripting an objection and recording how often the model caves. Because that objection is a prompt template, whatever else the template varies is measured along with the property it claims to isolate. We audit SycEval, which reports that objections raised before…
2609.32867v1 · 2026-09-26 · Atharv Gupta, Akshat Jindal, Lavanya Nigam et al.
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses.…
2609.32720v1 · 2026-09-26 · Shuang Liu, Yongliang Miao, Yanguang Liu et al.