FolioStart free

Computer Science · Literature

Research papers on Retrieval-augmented generation in language models

Recent and highly-cited academic work on retrieval-augmented generation in language models, gathered from Semantic Scholar, CrossRef and OpenAlex.

Search all 200M+ papers on this topic, free →Or track new retrieval-augmented generation in language models papers automatically as they publish
  1. Large language models encode clinical knowledge

    Karan Singhal, Shekoofeh Azizi, Tao Tu, et al. · 2023 · Nature · 3,762 citations

    Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias. In addition, we evaluate Pathways Lan

    Save this paper
  2. Retrieval-Augmented Generation for Large Language Models: A Survey

    Yunfan Gao, Yun Xiong, Xinyu Gao, et al. · 2023 · arXiv (Cornell University) · 732 citations

    Large Language Models (LLMs) showcase impressive capabilities but encounter challenges like hallucination, outdated knowledge, and non-transparent, untraceable reasoning processes. Retrieval-Augmented Generation (RAG) has emerged as a promising solution by incorporating knowledge from external databases. This enhances the accuracy and credibility of the generation, particularly for knowledge-intensive tasks, and allows for continuous knowledge updates and integration of domain-specific information. RAG synergistically merges LLMs' intrinsic knowledge with the vast, dynamic repositories of external databases. This comprehensive review paper offers a detailed examination of the progression of

    Save this paper
  3. A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models

    Wenqi Fan, Yujuan Ding, Liangbo Ning, et al. · 2024 · 729 citations

    As one of the most advanced techniques in AI, Retrieval-Augmented Generation (RAG) can offer reliable and up-to-date external knowledge, providing huge convenience for numerous tasks. Particularly in the era of AI-Generated Content (AIGC), the powerful capacity of retrieval in providing additional knowledge enables RAG to assist existing generative AI in producing high-quality outputs. Recently, Large Language Models (LLMs) have demonstrated revolutionary abilities in language understanding and generation, while still facing inherent limitations such as hallucinations and out-of-date internal knowledge. Given the powerful abilities of RAG in providing the latest and helpful auxiliary informa

    Save this paper
  4. Benchmarking Large Language Models in Retrieval-Augmented Generation

    Jiawei Chen, Hongyu Lin, Xianpei Han, et al. · 2024 · Proceedings of the AAAI Conference on Artificial Intelligence · 373 citations

    Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different large language models, which make it challenging to identify the potential bottlenecks in the capabilities of RAG for different LLMs. In this paper, we systematically investigate the impact of Retrieval-Augmented Generation on large language models. We analyze the performance of different large language models in 4 fundamental abilities required for RAG, including noise robustness, negative rejection, information integration, and counterfactual robustness

    Save this paper
  5. Retrieval augmented generation for large language models in healthcare: A systematic review

    Lameck Mbangula Amugongo, Pietro Mascheroni, Steven E. Brooks, et al. · 2025 · PLOS Digital Health · 213 citations

    Large Language Models (LLMs) have demonstrated promising capabilities to solve complex tasks in critical sectors such as healthcare. However, LLMs are limited by their training data which is often outdated, the tendency to generate inaccurate ("hallucinated") content and a lack of transparency in the content they generate. To address these limitations, retrieval augmented generation (RAG) grounds the responses of LLMs by exposing them to external knowledge sources. However, in the healthcare domain there is currently a lack of systematic understanding of which datasets, RAG methodologies and evaluation frameworks are available. This review aims to bridge this gap by assessing RAG-based appro

    Save this paper
  6. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework

    Simone Kresevic, Mauro Giuffrè, Miloš Ajčević, et al. · 2024 · npj Digital Medicine · 203 citations

    Large language models (LLMs) can potentially transform healthcare, particularly in providing the right information to the right provider at the right time in the hospital workflow. This study investigates the integration of LLMs into healthcare, specifically focusing on improving clinical decision support systems (CDSSs) through accurate interpretation of medical guidelines for chronic Hepatitis C Virus infection management. Utilizing OpenAI's GPT-4 Turbo model, we developed a customized LLM framework that incorporates retrieval augmented generation (RAG) and prompt engineering. Our framework involved guideline conversion into the best-structured format that can be efficiently processed by L

    Save this paper
  7. Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy

    Zhihong Shao, Yeyun Gong, Yelong Shen, et al. · 2023 · 180 citations

    Retrieval-augmented generation has raise extensive attention as it is promising to address the limitations of large language models including outdated knowledge and hallucinations.However, retrievers struggle to capture relevance, especially for queries with complex information needs.Recent work has proposed to improve relevance modeling by having large language models actively involved in retrieval, i.e., to guide retrieval with generation.In this paper, we show that strong performance can be achieved by a method we call ITER-RETGEN, which synergizes retrieval and generation in an iterative manner: a model's response to a task input shows what might be needed to finish the task, and thus ca

    Save this paper
  8. Integrating Retrieval-Augmented Generation with Large Language Models in Nephrology: Advancing Practical Applications

    Jing Miao, Charat Thongprayoon, Supawadee Suppadungsuk, et al. · 2024 · Medicina · 165 citations

    The integration of large language models (LLMs) into healthcare, particularly in nephrology, represents a significant advancement in applying advanced technology to patient care, medical research, and education. These advanced models have progressed from simple text processors to tools capable of deep language understanding, offering innovative ways to handle health-related data, thus improving medical practice efficiency and effectiveness. A significant challenge in medical applications of LLMs is their imperfect accuracy and/or tendency to produce hallucinations-outputs that are factually incorrect or irrelevant. This issue is particularly critical in healthcare, where precision is essenti

    Save this paper
  9. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines

    Siru Liu, Allison B. McCoy, Adam Wright · 2025 · Journal of the American Medical Informatics Association · 164 citations

    OBJECTIVE: The objectives of this study are to synthesize findings from recent research of retrieval-augmented generation (RAG) and large language models (LLMs) in biomedicine and provide clinical development guidelines to improve effectiveness. MATERIALS AND METHODS: We conducted a systematic literature review and a meta-analysis. The report was created in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 analysis. Searches were performed in 3 databases (PubMed, Embase, PsycINFO) using terms related to "retrieval augmented generation" and "large language model," for articles published in 2023 and 2024. We selected studies that compared baseline LLM per

    Save this paper
  10. Development of a liver disease–specific large language model chat interface using retrieval-augmented generation

    Jin Ge, Steve Sun, Joseph F. Owens, et al. · 2024 · Hepatology · 144 citations

    BACKGROUND AND AIMS: Large language models (LLMs) have significant capabilities in clinical information processing tasks. Commercially available LLMs, however, are not optimized for clinical uses and are prone to generating hallucinatory information. Retrieval-augmented generation (RAG) is an enterprise architecture that allows the embedding of customized data into LLMs. This approach "specializes" the LLMs and is thought to reduce hallucinations. APPROACH AND RESULTS: We developed "LiVersa," a liver disease-specific LLM, by using our institution's protected health information-complaint text embedding and LLM platform, "Versa." We conducted RAG on 30 publicly available American Association f

    Save this paper
  11. Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness

    Yu He Ke, Liyuan Jin, Kabilan Elangovan, et al. · 2025 · npj Digital Medicine · 133 citations

    Large Language Models (LLMs) hold promise for medical applications but often lack domain-specific expertise. Retrieval Augmented Generation (RAG) enables customization by integrating specialized knowledge. This study assessed the accuracy, consistency, and safety of LLM-RAG models in determining surgical fitness and delivering preoperative instructions using 35 local and 23 international guidelines. Ten LLMs (e.g., GPT3.5, GPT4, GPT4o, Gemini, Llama2, and Llama3, Claude) were tested across 14 clinical scenarios. A total of 3234 responses were generated and compared to 448 human-generated answers. The GPT4 LLM-RAG model with international guidelines generated answers within 20 s and achieved

    Save this paper
  12. SurgeryLLM: a retrieval-augmented generation large language model framework for surgical decision support and workflow enhancement

    Chin Siang Ong, Nicholas T. Obey, Yanan Zheng, et al. · 2024 · npj Digital Medicine · 44 citations

    SurgeryLLM, a large language model framework using Retrieval Augmented Generation demonstrably incorporated domain-specific knowledge from current evidence-based surgical guidelines when presented with patient-specific data. The successful incorporation of guideline-based information represents a substantial step toward enabling greater surgeon efficiency, improving patient safety, and optimizing surgical outcomes.

    Save this paper
  13. Adaptive Control of Retrieval-Augmented Generation for Large Language Models Through Reflective Tags

    Chengyuan Yao, Satoshi Fujita · 2024 · Electronics · 15 citations

    While retrieval-augmented generation (RAG) enhances large language models (LLMs), it also introduces challenges that can impact accuracy and performance. In practice, RAG can obscure the intrinsic strengths of LLMs. Firstly, LLMs may become too reliant on external retrieval, underutilizing their own knowledge and reasoning, which can diminish responsiveness. Secondly, RAG may introduce irrelevant or low-quality data, adding noise that disrupts generation, especially with complex tasks. This paper proposes an RAG framework that uses reflective tags to manage retrieval, evaluating documents in parallel and applying the chain-of-thought (CoT) technique for step-by-step generation. The model sel

    Save this paper
  14. Information retrieval from textual data: Harnessing large language models, retrieval augmented generation and prompt engineering

    Asen Hikov, Laura Murphy · 2024 · Journal of AI, Robotics & Workplace Automation · 14 citations

    This paper describes how recent advancements in the field of Generative AI (GenAI), and more specifically large language models (LLMs), are incorporated into a practical application that solves the widespread and relevant business problem of information retrieval from textual data in PDF format: searching through legal texts, financial reports, research articles and so on. Marketing research, for example, often requires reading through hundreds of pages of financial reports to extract key information for research on competitors, partners, markets and prospective clients. It is a manual, error-prone and time-consuming task for marketers, where until recently there was little scope for automat

    Save this paper
  15. The Role of Retrieval-Augmented Generation in Improving Factual Accuracy for Medical Large Language Models

    Eason Ni · 2026 · Scholarly Review Journal

    Large Language Models (LLMs) that rely solely on parametric memory learned through training have demonstrated strong performance in biomedical question-answering, but their tendency to hallucinate facts and the difficulty of adjusting and adding to the learned knowledge limit their usefulness in clinical settings. Retrieval-Augmented Generation (RAG) has emerged as a promising solution by adding a non-parametric source of memory and grounding LLM outputs to reputable external sources. This paper aims to survey the evolution of RAG methodologies and highlight the most recent developments and shifts in paradigms like Agentic RAG. We highlight current state-of-the-art biomedical RAG systems, th

    Save this paper
  16. A Framework for Adaptive Knowledge-Augmented Mizo Large Language Models Using Retrieval-Augmented Generation and Continual Learning

    Vanlalropuia Ralte, Abhisake Sinha - · 2026 · International Journal For Multidisciplinary Research

    The deployment of Large Language Models (LLMs) for low-resource languages is challenging due to the lack of linguistic resources, sparse digital content and the absence of structured knowledge bases. In this paper, we present an adaptive knowledge-augmented framework for Mizo Large Language Models by combining Retrieval-Augmented Generation (RAG) with continual learning. This methodology harnesses semantic retrieval with dense embeddings and FAISS indexing, adaptive evidence re-ranking, parameter-efficient fine-tuning, and incremental knowledge updating to enhance factual accuracy and decrease hallucinations. Experimental evaluation shows better retrieval performance, greater text creation q

    Save this paper
  17. Intelligent automation of economic processes based on retrieval-augmented generation and large language models

    Serhii Arefiev, Serhii Hildi · 2026 · Ukrainian Journal of Applied Economics and Technology

    The article is devoted to the theoretical and methodological substantiation of the concept of intelligent automation of economic processes based on the integration of Retrieval-Augmented Generation (RAG), Large Language Models (LLM), Prompt Engineering, and Automatic Engineering technologies. The modern digital economy is moving from technical automation to cognitive automation, in which self-learning systems are emerging that can adapt to environmental changes, analyze the results of their own activities, and generate new economic solutions. RAG acts as a cognitive intermediary between generation and data retrieval, ensuring the factual reliability of analytical results. LLMs provide semant

    Save this paper
  18. Quantum-Enhanced Retrieval-Augmented Generation for Hallucination Reduction in Large Language Models

    Praveenkumar Seepana · 2026 · International Journal of Computational Science and Engineering Research

    Although the performance of LLMs on a wide variety of natural language processing problems demonstrates remarkable ability, hallucinated responses are introduced as one of the key weaknesses of LLMs, especially in knowledge-intensive applications, where fidelity to facts is paramount. While effective in reducing hallucinations, the context retrieved during the RAG operations is still very often suboptimal with respect to the external documents used in the retrieval stage, and being based on cosine similarity and nearest-neighbour search, these models typically do not return optimal context to support factual generation. The authors propose a five-stage solution, called Quantum-Enhanced Retri

    Save this paper
  19. Hallucination in Large Language Models and Retrieval-Augmented Generation: Mechanisms, Mitigation, and Evaluation

    Haopeng Yang · 2026 · Theoretical and Natural Science

    Large language models have demonstrated strong generative capability in question answering, dialogue, and other knowledge-intensive tasks. However, their outputs remain vulnerable to hallucination, including factual errors, unsupported claims, spurious citations, and distorted reasoning. Retrieval-augmented generation (RAG) has been proposed as a practical remedy because it supplements parametric knowledge with external evidence retrieved at inference time. Yet RAG does not guarantee truthfulness or attribution by default. Errors may arise during query formulation, document retrieval, evidence aggregation, and answer grounding. This paper reviews the relationship between hallucination and RA

    Save this paper

Write your paper with these sources

Folio is the integrity-first research workspace: search 200M+ papers, save sources, and write with citations that format themselves. Free for students and researchers.

Start writing free →