FolioStart free

Computer Science · Literature

Research papers on Retrieval-augmented generation in language models

Recent and highly-cited academic work on retrieval-augmented generation in language models, gathered from Semantic Scholar, CrossRef and OpenAlex.

Search all 200M+ papers on this topic, free →Or track new retrieval-augmented generation in language models papers automatically as they publish
  1. Retrieval-Augmented Generation for Large Language Models: A Survey

    Yun-Fan Gao, Yun Xiong, Xin-Yu Gao, et al. · 2023 · ArXiv · 4,225 citations

    Large Language Models (LLMs) showcase impressive capabilities but encounter challenges like hallucination, outdated knowledge, and non-transparent, untraceable reasoning processes. Retrieval-Augmented Generation (RAG) has emerged as a promising solution by incorporating knowledge from external databases. This enhances the accuracy and credibility of the generation, particularly for knowledge-intensive tasks, and allows for continuous knowledge updates and integration of domain-specific information. RAG synergistically merges LLMs' intrinsic knowledge with the vast, dynamic repositories of external databases. This comprehensive review paper offers a detailed examination of the progression of

    Save this paper
  2. Benchmarking Large Language Models in Retrieval-Augmented Generation

    Jiawei Chen, Hongyu Lin, Xian-Pei Han, et al. · 2023 · 695 citations

    Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different large language models, which make it challenging to identify the potential bottlenecks in the capabilities of RAG for different LLMs. In this paper, we systematically investigate the impact of Retrieval-Augmented Generation on large language models. We analyze the performance of different large language models in 4 fundamental abilities required for RAG, including noise robustness, negative rejection, information integration, and counterfactual robustness

    Save this paper
  3. PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models

    Wei Zou, Runpeng Geng, Binghui Wang, et al. · 2024 · 354 citations

    Large language models (LLMs) have achieved remarkable success due to their exceptional generative capabilities. Despite their success, they also have inherent limitations such as a lack of up-to-date knowledge and hallucination. Retrieval-Augmented Generation (RAG) is a state-of-the-art technique to mitigate these limitations. The key idea of RAG is to ground the answer generation of an LLM on external knowledge retrieved from a knowledge database. Existing studies mainly focus on improving the accuracy or efficiency of RAG, leaving its security largely unexplored. We aim to bridge the gap in this work. We find that the knowledge database in a RAG system introduces a new and practical attack

    Save this paper
  4. Retrieval augmented generation for large language models in healthcare: A systematic review

    L. M. Amugongo, Pietro Mascheroni, Steve Brooks, et al. · 2025 · PLOS Digital Health · 282 citations

    Large Language Models (LLMs) have demonstrated promising capabilities to solve complex tasks in critical sectors such as healthcare. However, LLMs are limited by their training data which is often outdated, the tendency to generate inaccurate (“hallucinated”) content and a lack of transparency in the content they generate. To address these limitations, retrieval augmented generation (RAG) grounds the responses of LLMs by exposing them to external knowledge sources. However, in the healthcare domain there is currently a lack of systematic understanding of which datasets, RAG methodologies and evaluation frameworks are available. This review aims to bridge this gap by assessing RAG-based appro

    Save this paper
  5. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework

    Simone Kresevic, M. Giuffré, Miloš Ajčević, et al. · 2024 · NPJ Digital Medicine · 196 citations

    Large language models (LLMs) can potentially transform healthcare, particularly in providing the right information to the right provider at the right time in the hospital workflow. This study investigates the integration of LLMs into healthcare, specifically focusing on improving clinical decision support systems (CDSSs) through accurate interpretation of medical guidelines for chronic Hepatitis C Virus infection management. Utilizing OpenAI’s GPT-4 Turbo model, we developed a customized LLM framework that incorporates retrieval augmented generation (RAG) and prompt engineering. Our framework involved guideline conversion into the best-structured format that can be efficiently processed by L

    Save this paper
  6. Integrating Retrieval-Augmented Generation with Large Language Models in Nephrology: Advancing Practical Applications

    Jing Miao, C. Thongprayoon, S. Suppadungsuk, et al. · 2024 · Medicina · 177 citations

    The integration of large language models (LLMs) into healthcare, particularly in nephrology, represents a significant advancement in applying advanced technology to patient care, medical research, and education. These advanced models have progressed from simple text processors to tools capable of deep language understanding, offering innovative ways to handle health-related data, thus improving medical practice efficiency and effectiveness. A significant challenge in medical applications of LLMs is their imperfect accuracy and/or tendency to produce hallucinations—outputs that are factually incorrect or irrelevant. This issue is particularly critical in healthcare, where precision is essenti

    Save this paper
  7. A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models

    Qing-Gang Zhang, Shengyuan Chen, Yuanchen Bei, et al. · 2025 · ArXiv · 156 citations

    Large language models (LLMs) have demonstrated remarkable capabilities in a wide range of tasks, yet their application to specialized domains remains challenging due to the need for deep expertise. Retrieval-Augmented generation (RAG) has emerged as a promising solution to customize LLMs for professional fields by seamlessly integrating external knowledge bases, enabling real-time access to domain-specific expertise during inference. Despite its potential, traditional RAG systems, based on flat text retrieval, face three critical challenges: (i) complex query understanding in professional contexts, (ii) difficulties in knowledge integration across distributed sources, and (iii) system effici

    Save this paper
  8. Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness

    Yuhe Ke, Li-Yuan Jin, K. Elangovan, et al. · 2025 · NPJ Digital Medicine · 152 citations

    Large Language Models (LLMs) hold promise for medical applications but often lack domain-specific expertise. Retrieval Augmented Generation (RAG) enables customization by integrating specialized knowledge. This study assessed the accuracy, consistency, and safety of LLM-RAG models in determining surgical fitness and delivering preoperative instructions using 35 local and 23 international guidelines. Ten LLMs (e.g., GPT3.5, GPT4, GPT4o, Gemini, Llama2, and Llama3, Claude) were tested across 14 clinical scenarios. A total of 3234 responses were generated and compared to 448 human-generated answers. The GPT4 LLM-RAG model with international guidelines generated answers within 20 s and achieved

    Save this paper
  9. CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models

    Yuanjie Lyu, Zhi-Yu Li, Si-Min Niu, et al. · 2024 · ACM Transactions on Information Systems · 147 citations

    Retrieval-augmented generation (RAG) is a technique that enhances the capabilities of large language models (LLMs) by incorporating external knowledge sources. This method addresses common LLM limitations, including outdated information and the tendency to produce inaccurate “hallucinated” content. However, evaluating RAG systems is a challenge. Most benchmarks focus primarily on question-answering applications, neglecting other potential scenarios where RAG could be beneficial. Accordingly, in the experiments, these benchmarks often assess only the LLM components of the RAG pipeline or the retriever in knowledge-intensive scenarios, overlooking the impact of external knowledge base construc

    Save this paper
  10. Simple is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation

    Mufei Li, Siqi Miao, Pan Li · 2024 · ArXiv · 118 citations

    Large Language Models (LLMs) demonstrate strong reasoning abilities but face limitations such as hallucinations and outdated knowledge. Knowledge Graph (KG)-based Retrieval-Augmented Generation (RAG) addresses these issues by grounding LLM outputs in structured external knowledge from KGs. However, current KG-based RAG frameworks still struggle to optimize the trade-off between retrieval effectiveness and efficiency in identifying a suitable amount of relevant graph information for the LLM to digest. We introduce SubgraphRAG, extending the KG-based RAG framework that retrieves subgraphs and leverages LLMs for reasoning and answer prediction. Our approach innovatively integrates a lightweight

    Save this paper
  11. BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

    Jiaqi Xue, Meng Zheng, Yebowen Hu, et al. · 2024 · ArXiv · 109 citations

    Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However, this reliance on external knowledge introduces significant security vulnerabilities, as many RAG systems (e.g., Google Search) rely on large and unsanitized data repositories (e.g., Reddit). In this paper, we unveil a novel threat in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base. When a user's query contains attacker-specified trigger words, the RAG retrieves and refers to these malicious passages, enabling the att

    Save this paper
  12. DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models

    Wei-Hang Su, Yi-Chen Tang, Qing-Yao Ai, et al. · 2024 · 101 citations

    Dynamic retrieval augmented generation (RAG) paradigm actively decides when and what to retrieve during the text generation process of Large Language Models (LLMs). There are two key elements of this paradigm: identifying the optimal moment to activate the retrieval module (deciding when to retrieve) and crafting the appropriate query once retrieval is triggered (determining what to retrieve). However, current dynamic RAG methods fall short in both aspects. Firstly, the strategies for deciding when to retrieve often rely on static rules. Moreover, the strategies for deciding what to retrieve typically limit themselves to the LLM's most recent sentence or the last few tokens, while the LLM's

    Save this paper
  13. Graph Retrieval-Augmented Generation for Large Language Models: A Survey

    T. Procko, Omar Ochoa · 2024 · 2024 Conference on AI, Science, Engineering, and Technology (AIxSET) · 99 citations

    Large Language Models (LLMs) demonstrate general knowledge, but they suffer when specifically needed knowledge is not present in their training set. Two approaches to ameliorating this, without retraining, are 1) prompt engineering and 2) Retrieval-Augmented Generation (RAG). RAG is a form of prompt engineering, insofar as relevant lexical snippets retrieved from RAG corpora are vectorized and aggregated with prompts. However, RAG documents are often noisy, i.e., while relevant to a given prompt, they can contain much other information that obfuscates the desired snippet. If the purpose of pretraining a LLM on massive and general corpora is to engender a generally applicable model, RAG is no

    Save this paper
  14. Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey

    B. Ni, Zhe-Yuan Liu, Yong-Jia Lei, et al. · 2025 · ACM Computing Surveys · 66 citations

    Retrieval-Augmented Generation (RAG) enhances AI-generated content by integrating external knowledge, improving relevance, and reducing hallucinations. However, RAG also introduces risks related to reliability, safety, privacy, fairness, explainability, and accountability, which impact trustworthiness. While various methods aim to address these concerns, a unified framework is lacking. This survey bridges that gap by presenting a comprehensive roadmap for trustworthy RAG systems. We provide a structured analysis of key challenges, existing solutions, and future directions across these aspects. We further organize the field around how the retrieval and generation stages and the six trustworth

    Save this paper
  15. Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey

    Ao-Ran Gan, Hao Yu, Kai Zhang, et al. · 2025 · ArXiv · 55 citations

    Recent advancements in Retrieval-Augmented Generation (RAG) have revolutionized natural language processing by integrating Large Language Models (LLMs) with external information retrieval, enabling accurate, up-to-date, and verifiable text generation across diverse applications. However, evaluating RAG systems presents unique challenges due to their hybrid architecture that combines retrieval and generation components, as well as their dependence on dynamic knowledge sources in the LLM era. In response, this paper provides a comprehensive survey of RAG evaluation methods and frameworks, systematically reviewing traditional and emerging evaluation approaches, for system performance, factual a

    Save this paper
  16. SurgeryLLM: a retrieval-augmented generation large language model framework for surgical decision support and workflow enhancement

    Chin Siang Ong, Nicholas T. Obey, Yanan Zheng, et al. · 2024 · npj Digital Medicine · 47 citations

    SurgeryLLM, a large language model framework using Retrieval Augmented Generation demonstrably incorporated domain-specific knowledge from current evidence-based surgical guidelines when presented with patient-specific data. The successful incorporation of guideline-based information represents a substantial step toward enabling greater surgeon efficiency, improving patient safety, and optimizing surgical outcomes.

    Save this paper
  17. RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

    Bang An, Shiyue Zhang, M. Dredze · 2025 · ArXiv · 41 citations

    Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. However, despite the widespread use of the Retrieval-Augmented Generation (RAG) framework, AI safety work focuses on standard LLMs, which means we know little about how RAG use cases change a model's safety profile. We conduct a detailed comparative analysis of RAG and non-RAG frameworks with eleven LLMs. We find that RAG can make models less safe and change their safety profile. We explore the causes of this change and find that even combinations of safe models with safe documents can cause unsafe generations. In addition, we evaluate some existing red teaming methods for RA

    Save this paper
  18. PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization

    Yang Jiao, Xiao-Dong Wang, Kai Yang · 2025 · Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval · 39 citations

    Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of applications, e.g., medical question-answering, mathematical sciences, and code generation. However, they also exhibit inherent limitations, such as outdated knowledge and susceptibility to hallucinations. Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm to address these issues, but it also introduces new vulnerabilities. Recent efforts have focused on the security of RAG-based LLMs, yet existing attack methods face three critical challenges: (1) their effectiveness declines sharply when only a limited number of poisoned texts can be injected into the knowledge database (2) th

    Save this paper
  19. Adaptive Control of Retrieval-Augmented Generation for Large Language Models Through Reflective Tags

    Chengyuan Yao, Satoshi Fujita · 2024 · Electronics · 15 citations

    While retrieval-augmented generation (RAG) enhances large language models (LLMs), it also introduces challenges that can impact accuracy and performance. In practice, RAG can obscure the intrinsic strengths of LLMs. Firstly, LLMs may become too reliant on external retrieval, underutilizing their own knowledge and reasoning, which can diminish responsiveness. Secondly, RAG may introduce irrelevant or low-quality data, adding noise that disrupts generation, especially with complex tasks. This paper proposes an RAG framework that uses reflective tags to manage retrieval, evaluating documents in parallel and applying the chain-of-thought (CoT) technique for step-by-step generation. The model sel

    Save this paper
  20. Information retrieval from textual data: Harnessing large language models, retrieval augmented generation and prompt engineering

    Asen Hikov, Laura Murphy · 2024 · Journal of AI, Robotics & Workplace Automation · 15 citations

    This paper describes how recent advancements in the field of Generative AI (GenAI), and more specifically large language models (LLMs), are incorporated into a practical application that solves the widespread and relevant business problem of information retrieval from textual data in PDF format: searching through legal texts, financial reports, research articles and so on. Marketing research, for example, often requires reading through hundreds of pages of financial reports to extract key information for research on competitors, partners, markets and prospective clients. It is a manual, error-prone and time-consuming task for marketers, where until recently there was little scope for automat

    Save this paper
  21. The Development and Evaluation of a Retrieval-Augmented Generation Large Language Model Virtual Assistant for Postoperative Instructions

    Syed Ali Haider, Srinivasagam Prabha, Cesar Abraham Gomez Cabello, et al. · 2025 · Bioengineering · 9 citations

    BACKGROUND: During postoperative recovery, patients and their caregivers often lack crucial information, leading to numerous repetitive inquiries that burden healthcare providers. Traditional discharge materials, including paper handouts and patient portals, are often static, overwhelming, or underutilized, leading to patient overwhelm and contributing to unnecessary ER visits and overall healthcare overutilization. Conversational chatbots offer a solution, but Natural Language Processing (NLP) systems are often inflexible and limited in understanding, while powerful Large Language Models (LLMs) are prone to generating "hallucinations". OBJECTIVE: To combine the deterministic framework of tr

    Save this paper

Write your paper with these sources

Folio is the integrity-first research workspace: search 200M+ papers, save sources, and write with citations that format themselves. Free for students and researchers.

Start writing free →