# 대규모 언어 모델을 위한 검색-증강 생성(RAG) 기술 현황 - 번외편

**URL:** https://discuss.pytorch.kr/t/rag/3255
**Category:** 읽을거리&정보공유
**Tags:** llm-survey, rag, paper, survey-paper
**Created:** [1월 11, 2024, 12:45오후 UTC](https://discuss.pytorch.kr/t/rag/3255 "2024-01-11T12:45:41Z")
**Posts on this page:** 1
**Page:** 1

<div class="post-metadata">

### Author: ![9bow](https://discuss.pytorch.kr/user_avatar/discuss.pytorch.kr/9bow/32/16301_2.png) [@9bow](https://discuss.pytorch.kr/u/9bow)
#### Post date: [1월 11, 2024, 12:45오후 UTC](https://discuss.pytorch.kr/t/rag/3255/1 "2024-01-11T12:45:41Z")

</div>

### PyTorchKR​🔥🇰🇷🤔💬

- [12/18~24의 주요 ML 논문](https://discuss.pytorch.kr/t/2023-12-18-12-24-ml-top-ml-papers-of-the-week/3111)에 소개된 [RAG 기술에 대한 서베이 논문](https://discuss.pytorch.kr/t/2023-12-18-12-24-ml-top-ml-papers-of-the-week/3111#retrieval-augmented-generation-for-large-language-models-a-survey-41)을 정리해보았습니다.
- LLM의 활용이 늘어나며, RAG에 대한 연구들 또한 계속되고 있습니다.
- [1부에서는 RAG 기술의 패러다임들](https://discuss.pytorch.kr/t/rag-1-2/3135)을, [2부에서는 주요 구성요소들](https://discuss.pytorch.kr/t/rag-2-2/3160)에 대해서 알아보았습니다. [번외편에서는 저자들이 GitHub 저장소에 공개된 발표 슬라이드(영문)와 참고 기술/코드를 정리한 내용](https://discuss.pytorch.kr/t/rag/3255)을 소개하려고 합니다.
  - [1부: RAG에 대한 간단한 소개와 함께 주요 패러다임들과 각 세부 기술, 방법론들에 대한 소개](https://discuss.pytorch.kr/t/rag-1-2/3135)
  - [2부: RAG의 구성 요소들에 대한 소개와 고려해야 할 주요 주제들과 방법론들에 대한 소개](https://discuss.pytorch.kr/t/rag-2-2/3160)
  - [번외편: 서베이 논문 발표 자료 및 참고 논문/코드](https://discuss.pytorch.kr/t/rag/3255) 👈

> **[Retrieval-Augmented Generation for Large Language Models: A Survey](https://arxiv.org/abs/2312.10997v1)**
>
> Large Language Models (LLMs) showcase impressive capabilities but encounter challenges like hallucination, outdated knowledge, and non-transparent, untraceable reasoning processes. Retrieval-Augmented Generation (RAG) has emerged as a promising...

⚠ hoxy... [RAG 기술 현황 1편](https://discuss.pytorch.kr/t/rag-1-2/3135)과 [2편](https://discuss.pytorch.kr/t/rag-2-2/3160)을 읽고 오셨나요? 아직 읽지 않으셨다면 [1편](https://discuss.pytorch.kr/t/rag-1-2/3135)부터 읽으시는 것을 추천드립니다!

* * *

## RAG 기술 전체 보기

 ![RAG 기술 전체 보기](https://discuss.pytorch.kr/uploads/default/original/2X/f/f8eb513f74ec454730646da7d2d5a5689d5e6cbe.png)

## 관련 논문 및 코드 목록

### Augmentation Stage

#### Pre-training

1.Improving language models by retrieving from trillions of tokens [[paper]](https://markdown.com.cn)[[code]](https://markdown.com.cn)

2.Few-shot Learning with Re-trieval Augmented Language Models [[paper]](https://arxiv.org/pdf/2208.03299.pdf)

3.Toolformer: Language Models Can Teach Themselves to Use Tools[[paper]](https://arxiv.org/abs/2302.04761)

4.Copy is all you need[[paper]](https://openreview.net/pdf?id=CROlOA9Nd8C)

5.In-context learning with retrieval augmented encoder-decoder language model[[paper]](https://arxiv.org/abs/2308.07922)

6.Shall we pretrain autoregressive language models with retrieval?[[paper]](https://arxiv.org/abs/2304.06762)

7.Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP[[paper]](https://arxiv.org/abs/2212.14024)

#### Fine-tuning

1.Dense Passage Retrieval for Open-Domain Question Answering[[paper]](https://arxiv.org/abs/2004.04906)

2.UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation[[paper]](https://arxiv.org/abs/2303.08518)[[code]](https://github.com/microsoft/LMOps)

3.Distilling knowledge from reader to retriever for question answering[[paper]](https://arxiv.org/abs/2012.04584)

4.RA-DIT: Retrieval-Augmented Dual Instruction Tuning[[paper]](https://arxiv.org/abs/2310.01352)

5.Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection[[paper]](https://arxiv.org/abs/2310.11511)

6.Knowledge Graph-Augmented Language Models for Knowledge-Grounded Dialogue Generation[[paper]](https://arxiv.org/abs/2305.18846)

7.Structure-Aware Language Model Pretraining Improves Dense Retrieval on Structured Data [[paper]](https://aclanthology.org/2023.findings-acl.734.pdf) [[code]](https://github.com/OpenMatch/SANTA)

8.Replug: Retrieval-augmented black-box language models [[paper]](https://arxiv.org/pdf/2301.12652.pdf)

9.Augmentation-Adapted Retriever Improves Generalization of Language  
Models as Generic Plug-In [[paper]](https://arxiv.org/abs/2305.17331)[[code]](https://github.com/OpenMatch/Augmentation-Adapted-Retriever)

#### Inference

1.Generalization through Memorization: Nearest Neighbor Language Models[[paper]](https://arxiv.org/abs/1911.00172)

2.DEMONSTRATE–SEARCH–PREDICT:  
Composing retrieval and language models for knowledge-intensive NLP [[paper]](https://arxiv.org/abs/2212.14024)[[code]](https://github.com/stanfordnlp/dspy)

3.Keyword Augmented Retrieval: Novel framework for Information Retrieval integrated with speech interface. [[paper]](https://arxiv.org/abs/2310.04205)

4.Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. [[paper]](https://arxiv.org/pdf/2212.10509.pdf)[[code]](https://github.com/stonybrooknlp/ircot)

5.Generate rather than Retrieve: Large Language Models are Strong Context Generators [[paper]](https://arxiv.org/abs/2209.10063) [[code]](https://github.com/wyu97/GenRead)

6.In-Context Retrieval-Augmented Language Models [[paper]](https://arxiv.org/abs/2302.00083)

### Augmentation Source

#### Unstructured Data

1.UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation[[paper]](https://arxiv.org/abs/2303.08518)[[code]](https://github.com/microsoft/LMOps)

2.From Classification to Generation: Insights into Crosslingual Retrieval Augmented ICL [[paper]](https://arxiv.org/abs/2311.06595)

3.Copy is all you need [[paper]](https://openreview.net/pdf?id=CROlOA9Nd8C)

#### Structured Data

1.FABULA: Intelligence Report Generation Using Retrieval-Augmented Narrative Construction [[paper]](https://arxiv.org/abs/2310.13848)

2.Knowledge Graph-Augmented Language Models for Knowledge-Grounded Dialogue Generation [[paper]](https://arxiv.org/abs/2305.18846)

3.KnowledGPT: Enhancing Large Language Models with Retrieval and Storage Access on Knowledge Bases [[paper]](https://arxiv.org/abs/2308.11761)

4.Graph-ToolFormer: To Empower LLMs with Graph Reasoning Ability via Prompt Augmented by ChatGPT [[paper]](https://arxiv.org/abs/2304.11116)

#### LLM Generated Content

1.Lift Yourself Up: Retrieval-augmented Text Generation with Self-Memory [[paper]](https://arxiv.org/abs/2305.02437)

2.DEMONSTRATE–SEARCH–PREDICT:  
Composing retrieval and language models for knowledge-intensive NLP [[paper]](https://arxiv.org/abs/2212.14024)

3.Recitation-augmented language models[[paper]](https://arxiv.org/pdf/2210.01296.pdf)

4.Generate rather than Retrieve: Large Language Models are Strong Context Generators [[paper]](https://arxiv.org/abs/2209.10063)

5.Self-Knowledge Guided Retrieval Augmentation for Large Language Models [[paper]](https://arxiv.org/abs/2310.05002)

### Augmentation Process

#### Once Retrieval

1.Retrieval-augmented generation for knowledge-intensive nlp tasks [[paper]](https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html)

2.UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation [[paper]](https://arxiv.org/abs/2303.08518)

3.Augmented Large Language Models with Parametric Knowledge Guiding [[paper]](https://arxiv.org/abs/2305.04757)

4.Learning to Retrieve In-Context Examples for Large Language Models.[[paper]](https://arxiv.org/pdf/2307.07164.pdf)

5.Few-shot Learning with Re-trieval Augmented Language Models [[paper]](https://arxiv.org/pdf/2208.03299.pdf)  
6.Replug: Retrieval-augmented black-box language models [[paper]](https://arxiv.org/pdf/2301.12652.pdf)

7.Recitation-augmented language models[[paper]](https://arxiv.org/pdf/2210.01296.pdf)

#### Iterative Retrieval

1.DEMONSTRATE–SEARCH–PREDICT:  
Composing retrieval and language models for knowledge-intensive NLP [[paper]](https://arxiv.org/abs/2212.14024)[[code]](https://github.com/stanfordnlp/dspy)

2.Retrieve-and-Sample: Document-level Event Argument Extraction via Hybrid Retrieval Augmentation [[paper]](https://aclanthology.org/2023.acl-long.17/)

3.Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy[[paper]](https://arxiv.org/abs/2305.15294)

4.RETRIEVAL-GENERATION SYNERGY AUGMENTED LARGE LANGUAGE MODELS [[paper]](https://arxiv.org/abs/2310.05149)

#### Recursive Retrieval

1.Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions [[paper]](https://arxiv.org/abs/2212.10509)[[code]](https://github.com/stonybrooknlp/ircot)

2.Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models [[paper]](https://arxiv.org/abs/2310.14696)

#### Adaptive Retrieval

1.Active Retrieval Augmented Generation[[paper]](https://arxiv.org/abs/2305.06983)[[code]](https://github.com/jzbjyb/FLARE)

2.Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection [[paper]](https://arxiv.org/abs/2310.11511)

3.In-context learning with retrieval augmented encoder-decoder language model [[paper]](https://arxiv.org/abs/2308.07922)

## 더 읽어보기

### 서베이 논문 저자들의 GitHub 저장소

[https://github.com/tongji-kgllm/rag-survey](https://github.com/tongji-kgllm/rag-survey)

### 발표 슬라이드 (PDF/영문)

> **[RAG\_Slide\_ENG.pdf](https://github.com/Tongji-KGLLM/RAG-Survey/blob/main/assets/RAG_Slide_ENG.pdf)**

[RAG\_Slide\_ENG.pdf](https://discuss.pytorch.kr/uploads/short-url/kyUWgYEFD76LM1h8VrR8ACMNTqc.pdf) (8.8 MB)

### 원본 서베이 논문

> **[Retrieval-Augmented Generation for Large Language Models: A Survey](https://arxiv.org/abs/2312.10997v1)**
>
> Large Language Models (LLMs) showcase impressive capabilities but encounter challenges like hallucination, outdated knowledge, and non-transparent, untraceable reasoning processes. Retrieval-Augmented Generation (RAG) has emerged as a promising...

* * *

[🔥파이토치 한국 사용자 모임:kr:](https://pytorch.kr/)이 정리한 이 글이 유용하셨나요? [회원으로 가입](https://discuss.pytorch.kr/signup)하시면 주요 글들을 이메일📨로 보내드립니다! (기본은 Weekly지만 [Daily로 변경도 가능](https://discuss.pytorch.kr/my/preferences/emails)합니다.)

🎁 아래↘쪽에 좋아요❤를 눌러주시면 뉴스 발행에 힘이 됩니다~ 🙇‍♂️
