Bibliographie
Sources sur lesquelles repose le catalogue. Chaque référence est cotée A–F et liée aux fiches qui la citent.
Chaque source est cotée selon l'échelle de preuve A–F du manuel. Un DOI doit résoudre vers le titre correct ; un arXiv est indiqué comme préprint. Les sources sans identifiant résolvable sont signalées pour vérification.
- BGrade B
(2026). DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models. LREC.
https://arxiv.org/abs/2510.27543Cité dans :B72 - BGrade B
Braun,. (2025). Acquiescence Bias in Large Language Models. EMNLP Findings.
https://arxiv.org/abs/2509.08480Cité dans :B58 - BGrade B
Robinson et al.,. (2025). AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic. Findings ACL.
https://arxiv.org/abs/2412.04193Cité dans :B72 - BGrade B
Shao et al.,. (2025). Benford's Curse: Tracing Digit Bias to Numerical Hallucination. NeurIPS.
https://arxiv.org/abs/2506.01734 - BGrade B
Kim, Garg, Peng & Garg,. (2025). Correlated Errors in Large Language Models. ICML.
https://arxiv.org/abs/2506.07962 - AGrade A
Yan et al.,. (2025). Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic. EMNLP.
doi.org/10.18653/v1/2025.emnlp-main.681 - BGrade B
Zhi et al.. (2025). Exposing Product Bias in LLM Investment Recommendation.
https://arxiv.org/abs/2503.08750Cité dans :B37 - AGrade A
Peters & Chin-Yee. (2025). Generalization bias in LLM summarization. R. Soc. Open Sci..
doi.org/10.1098/rsos.241776Cité dans :B30 - AGrade A
Hu et al.,. (2025). Generative language models exhibit social identity biases. Nature Computational Science 5.
doi.org/10.1038/s43588-024-00741-1Cité dans :B34 - AGrade A
Zhu et al.,. (2025). Is Your LLM Outdated? A Deep Look at Temporal Generalization. NAACL.
doi.org/10.18653/v1/2025.naacl-Cité dans :B09 - BGrade B
Ibrahim et al.. (2025). MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers · arXiv:2508.14925. — position paper.
https://arxiv.org/abs/2508.14925 - BGrade B
Ibrahim et al.. (2025). Measuring and mitigating overreliance to build human-compatible AI. — position paper.
https://arxiv.org/abs/2509.08010 - BGrade B
Wu et al.. (2025). MultiRAG: Mitigating Hallucination in Multi-source RAG.
https://arxiv.org/abs/2508.03553Cité dans :B28 - BGrade B
Zou et al.,. (2025). Muse Spark Safety & Preparedness Report (Meta Superintelligence Labs, 2026). USENIX Security.
https://arxiv.org/abs/2402.07867Cité dans :B40 - BGrade B
Nasr et al.,. (2025). Scalable Extraction of Training Data from (Production) Language Models. ICLR.
https://arxiv.org/abs/2311.17035Cité dans :B76 - AGrade A
Kalai et al.,. (2025). |Why Language Models Hallucinate. OpenAI.
https://arxiv.org/abs/2509.04664 - AGrade A
Tao et al.,. (2024). Cultural bias and cultural alignment of LLMs. PNAS Nexus.
doi.org/10.1093/pnasnexus/pgae346 - AGrade A
Wyllie et al.,. (2024). Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias. FAccT.
doi.org/10.1145/3630106.3659029Cité dans :B38 - AGrade A
Zeng et al.,. (2024). How Johnny Can Persuade LLMs to Jailbreak Them. ACL.
doi.org/10.18653/v1/2024.acllong.773Cité dans :B11 - BGrade B
Zhan et al.,. (2024). — préprint retiré (déc. 2025) ; B69 re-sourcée vers Holtzman et al., « The Curious Case of Neural Text Degeneration » (ICLR 2020) InjecAgent. ACL Findings.
https://arxiv.org/abs/2403.02691 - BGrade B
Panickssery et al.,. (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS.
https://arxiv.org/abs/2404.13076Cité dans :B54 - BGrade B
(2024). LLMs Are Not Robust Multiple Choice Selectors (Zheng et al., ICLR 2024); Pezeshkpour & Hruschka. NAACL Findings.
https://arxiv.org/abs/2309.03882Cité dans :B50 - AGrade A
Liu et al.,. (2024). Lost in the Middle: How LMs Use Long Contexts. TACL.
doi.org/10.1162/tacl_a_00638 - BGrade B
Sclar et al.,. (2024). Quantifying LMs' Sensitivity to Spurious Features in Prompt Design. ICLR.
https://arxiv.org/abs/2310.11324Cité dans :B57 - AGrade A
Yuan et al.,. (2024). |Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations. EMNLP.
doi.org/10.18653/v1/2024.emnlp-main.155Cité dans :B05 - BGrade B
(2023). Chain-of-Thought Reasoning In The Wild Is Not Always Faithful (Arcuschin et al.); Turpin et al.. NeurIPS.
https://arxiv.org/abs/2503.08679Cité dans :B64 - BGrade B
Jin et al.,. (2023). CLadder: Assessing Causal Reasoning in Language Models. NeurIPS , vol. 36.
https://arxiv.org/abs/2312.04350 - BGrade B
Zheng et al.,. (2023). Judging the Judges: Position Bias in LLM-as-a-Judge (Shi et al., AACL-IJCNLP 2025); MT-Bench. NeurIPS.
https://arxiv.org/abs/2406.07791Cité dans :B51 - BGrade B
Turpin et al.,. (2023). Language Models Don’t Always Say What They Think. NeurIPS , vol. 36.
https://arxiv.org/abs/2305.04388Cité dans :B64 - BGrade B
Kandpal et al.,. (2023). Large Language Models Struggle to Learn Long-Tail Knowledge. ICML , PMLR 202.
https://arxiv.org/abs/2211.08411 - BGrade B
Mallen et al.,. (2023). |When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. ACL — PopQA.
https://arxiv.org/abs/2212.10511Cité dans :B56 - AGrade A
Parrish et al.,. (2022). Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care (Zack et al., The Lancet Digital Health 2024;6(1):e12-e22). Findings ACL , p. -2105.
doi.org/10.18653/v1/2022.findings-acl.165Cité dans :B33 - BGrade B
Carlini et al.,. (2021). Extracting Training Data from Large Language Models. USENIX Security.
https://arxiv.org/abs/2012.07805 - AGrade A
Fire & Guestrin. (2019). Goodhart's Law in action. GigaScience.
doi.org/10.1093/gigascience/giz053Cité dans :B14 - AGrade A
Head et al.,. (2015). The Extent and Consequences of P-Hacking. PLOS Biology.
doi.org/10.1371/journal.pbio.1002106 - BGrade B
AgentProp-Bench — Evaluating Tool-Using Language Agents .
https://arxiv.org/abs/2604.16706Cité dans :Annexe D - AGrade A
AI models collapse when trained on recursively generated data. Nature.
doi.org/10.1038/s41586-024-07566-yCité dans :B38 - BGrade B
- AGrade A
Dubois et al.. AraTrust, AraDiCE (S-21/22) — Verified · Peer-reviewed (COLING 2025); MENAValues, DialectalArabicMMLU (Reported).
https://arxiv.org/abs/2602.23971 - BGrade B
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge .
https://arxiv.org/abs/2510.18196Cité dans :B63 - BGrade B
2412.18626. Counting Ability of LLMs and Impact of Tokenization (Zhang et al.); Why LLMs Struggle to Count Letters? (Fu et al.).
https://arxiv.org/abs/2410.19730Cité dans :B52 - BGrade B
- BGrade B
Directional AI Advice: Experimental Evidence from Healthcare .
https://arxiv.org/abs/2607.08706Cité dans :B92 - BGrade B
Huang et al.. Emergent Social Intelligence Risks in Generative Multi-Agent Systems.
https://arxiv.org/abs/2603.27771 - BGrade B
- AGrade A
Advani. From Confident Closing to Silent Failure — False Success in LLM Agents.
doi.org/10.1145/3582269.3615599 - AGrade A
Kotek et al.,. Gender bias and stereotypes in Large Language Models. ACM CI ’23.
doi.org/10.1145/3582269.3615599Cité dans :B33 - BGrade B
- AGrade A
Lin. Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review.
doi.org/10.18653/v1/2024.acllong.773 - BGrade B
- BGrade B
- BGrade B
Induction Head Toxicity Explains Repetition Curse in LLMs .
https://arxiv.org/abs/2505.13514 - BGrade B
Judging the Judges — Bias Mitigation in LLM-as-a-Judge Pipelines .
https://arxiv.org/abs/2604.23178Cité dans :B55 - BGrade B
- BGrade B
- BGrade B
Zhao et al.. LLM hallucinations in the wild: Large-scale evidence from non-existent citations.
https://arxiv.org/abs/2605.07723 - BGrade B
DELEGATE-52. LLMs Corrupt Your Documents When You Delegate.
https://arxiv.org/abs/2604.15597 - BGrade B
Salecha et al.. LLMs Show Human-like Social Desirability Biases in Survey Responses.
https://arxiv.org/abs/2405.06058 - BGrade B
Malberg et al. — Comprehensive Eval of Cognitive Biases in LLMs (NLP4DH/ACL 2025; arXiv:2410.15413): availability heuristic (B06) & survivorship bias (B18) among 30 biases × 20 LLMs .
https://arxiv.org/abs/2410.15413 - BGrade B
- BGrade B
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples. Anthropic, UK AI Security Institute, Alan Turing Institute.
https://arxiv.org/abs/2510.07192 - AGrade A
- BGrade B
Alvarez et al.. ProvenanceGuard — Source-Aware Factuality Verification.
https://arxiv.org/abs/2606.18037 - BGrade B
Ding et al.. Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges.
https://arxiv.org/abs/2602.13576 - AGrade A
Wu et al.. Self-Preference / familiarity (Wataoka et al.) + LLM-as-a-judge audits · arXiv:2410.21819. Nat. Comm..
doi.org/10.1038/s41467-025-58551-6Cité dans :B55 - AGrade A
- BGrade B
- AGrade A
Leng et al.. Taming Overconfidence in LLMs: Reward Calibration in RLHF.
doi.org/10.1371/journal.pbio.1002106Cité dans :B66 - BGrade B
OIT. The Refusal–Compliance Tradeoff — A Large-Scale Safety Behavior Audit · arXiv:2605.05427.
https://arxiv.org/abs/2605.05427Cité dans :B70 - AGrade A
Lundin et al.. |The Token Tax: Systematic Bias in Multilingual Tokenization.
https://arxiv.org/abs/2509.05486 - BGrade B
Zhang et al.. |Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
https://arxiv.org/abs/2510.01171Cité dans :B61 - BGrade B
Singhal et al.. |Verbosity Bias in Preference Labeling (Saito et al.); Length Correlations in RLHF.
https://arxiv.org/abs/2310.10076Cité dans :B53 - BGrade B
- BGrade B
Li & Shi. |Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents.
https://arxiv.org/abs/2606.30383Cité dans :B85