LLM-NewsBench: Assessing Factuality, Hallucination, Headline Consistency, and Journalistic Quality in LLM-Generated News
DOI:
https://doi.org/10.47941/ijce.3967Keywords:
LLM NewsBench, factual evaluation, news generation, source-grounding, artificial intelligenceAbstract
Purpose: Large language models (LLMs) are increasing capabilities for producing fluent, cohesive, and publishable news stories but language quality by itself is not an indication of factual correctness. This paper champions LLM-NewsBench, a source-grounded benchmark for factual evaluation of LLM-generated news articles against a human-authored reference corpus.
Methodology: LLM-NewsBench comprises 50 news reports published by Dawn (Pakistan), each independently reconstructed three times from ChatGPT, Gemini, and DeepSeek. Both headlines and full news articles were evaluated. Headlines were scored for fidelity with the reference headline, while articles were scored for source-grounding completeness hallucination, journalistic style and neutrality, news article structure readability source fidelity, compression, and coherence. The proposed weighted compositional metric weighted factual (30%), completeness (20%), hallucination (20%), journalistic style (10%), structure (10%), and readability (10%) score yielded optimized aggregate benchmark scores of 0.87 for ChatGPT, 0.83 for Gemini, and 0.79 for DeepSeek (sets 150).
Findings: ChatGPT also achieved the highest factual grounding neutrality source fidelity, and compression in the current evaluation corpus. Gemini produced more detailed and comprehensive reconstructions but proved more aggressive in editing expansions than the more conversative DeepSeek which produced more analytical and feature-rich reconstructions but posed a greater threat of unsupported add-ons and narrative drift. Headline analysis revealed reference matching headlines of 78% for ChatGPT, 70% for DeepSeek, and 48% for Gemini. Error analysis revealed made-up entities, fabricated events, unsupported data, time-disarticulation, location hijacking, and overcommitted causal expansion.
Unique contribution to theory, practice and policy: These results demonstrate that headline similarity, semantic presentation, and stylistic quality are insufficient indicators of source-based fact complexity, and LLM-NewsBench establishes an operationalized reproducible multi-faceted computer-based evaluation infrastructure for news generation benchmarking, automated scoring, human validation, and visualization.
Downloads
References
Achiam, Josh, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida et al. "Gpt-4 technical report." arXiv preprint arXiv:2303.08774 (2023).
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
Chen, S., Gao, S., & He, J. (2023). Evaluating factual consistency of summaries with large language models. arXiv preprint arXiv:2305.14069.
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171-4186).
Durmus, E., He, H., & Diab, M. (2020, July). FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th annual meeting of the association for computational linguistics (pp. 5055-5070).
Jafari, N., Allan, J., & Iyyer, M. (2026). Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation. arXiv preprint arXiv:2604.03141.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., ... & Fung, P. (2023). Survey of hallucination in natural language generation. ACM computing surveys, 55(12), 1-38.
Kim, S., & Kim, S. (2024). Can language models evaluate human written text? case study on Korean student writing for education. arXiv preprint arXiv:2407.17022.
Kryściński, W., McCann, B., Xiong, C., & Socher, R. (2020, November). Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) (pp. 9332-9346).
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33, 9459-9474.
Lin, C. Y. (2004, July). Rouge: A package for automatic evaluation of summaries. In Text summarization branches out (pp. 74-81).
Lopez-Lira, A., & Tang, Y. (2026). Can ChatGPT forecast stock price movements? return predictability and large language models. Journal of Financial Economics, 184, 104335.
Manakul, P., Liusie, A., & Gales, M. (2023, December). Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 conference on empirical methods in natural language processing (pp. 9004-9017).
Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020, July). On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th annual meeting of the association for computational linguistics (pp. 1906-1919).
McCutcheon, A., & Brogly, C. (2025, September). Do small language models generate realistic variable-quality fake news headlines?. In 2025 3rd International Conference on Artificial Intelligence, Blockchain, and Internet of Things (AIBThings) (pp. 1-5). IEEE.
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W. T., Koh, P., ... & Hajishirzi, H. (2023, December). FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 conference on empirical methods in natural language processing (pp. 12076-12100).
Ogawa, S. (2024). Linearization of transition functions along a certain class of Levi-flat hypersurfaces. Hokkaido Mathematical Journal, 53(3), 485-510.
Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002, July). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics (pp. 311-318).
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., ... & Batsaikhan, B. O. (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... & Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
Wang, A., Cho, K., & Lewis, M. (2020, July). Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th annual meeting of the association for computational linguistics (pp. 5008-5020).
Xu, W., Napoles, C., Pavlick, E., Chen, Q., & Callison-Burch, C. (2016). Optimizing statistical machine translation for text simplification. Transactions of the Association for Computational Linguistics, 4, 401-415.
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., & Choi, Y. (2019). Defending against neural fake news. Advances in neural information processing systems, 32.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2019). Bertscore: Evaluating text generation with Bert. arXiv preprint arXiv:1904.09675.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Ayesha Muiz Mir, Abbas Rashid Butt

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution (CC-BY) 4.0 License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.