# \[2023/08/07 ~ 08/13\] 이번 주의 주요 ML 논문 (Top ML Papers of the Week)

**URL:** https://discuss.pytorch.kr/t/2023-08-07-08-13-ml-top-ml-papers-of-the-week/2293
**Category:** 읽을거리&정보공유
**Tags:** trustworthy-llm, llm-generalization, paper, llm-agent, nlp, top-ml-papers-of-the-week, synjax, pug-dataset, llm-as-dba, agentbench, llm, synthetic-data
**Created:** [8월 16, 2023, 7:02오전 UTC](https://discuss.pytorch.kr/t/2023-08-07-08-13-ml-top-ml-papers-of-the-week/2293 "2023-08-16T07:02:09Z")
**Posts on this page:** 1
**Page:** 1

<div class="post-metadata">

### Author: ![9bow](https://discuss.pytorch.kr/user_avatar/discuss.pytorch.kr/9bow/32/16301_2.png) [@9bow](https://discuss.pytorch.kr/u/9bow)
#### Post date: [8월 16, 2023, 7:02오전 UTC](https://discuss.pytorch.kr/t/2023-08-07-08-13-ml-top-ml-papers-of-the-week/2293/1 "2023-08-16T07:02:10Z")

</div>

- _이 글은 GPT 모델로 자동 요약한 설명으로, 잘못된 내용이 있을 수 있으니 원문을 참고해주세요! 😄_

- _읽으시면서 어색하거나 잘못된 내용을 발견하시면 덧글로 알려주시기를 부탁드립니다! 🙇‍♂️_

> **[🥇Top ML Papers of the Week](https://nlp.elvissaravia.com/p/top-ml-papers-of-the-week-f42)**
>
> The top ML Papers of the Week (August 7 - August 13)

* * *

## 서론

- 이번 주에는 대부분의 논문이 대형 언어 모델(LLMs)에 초점을 맞추고 있습니다. 이는 최근 AI 분야에서 LLM의 중요성이 증가하고 있음을 반영한 것으로 보입니다.

- LLM은 텍스트 데이터를 이해하고 생성하는 능력을 가지고 있어, 다양한 분야에서 활용 가능성이 무궁무진합니다.

- 그러나 LLM의 편향성, 신뢰성, 안전성 등에 대한 이슈도 동시에 제기되고 있습니다. 이번 주의 논문들은 이러한 이슈를 해결하고, LLM의 성능을 향상시키는 방법에 대한 연구를 제시하고 있습니다.

## 요약

### 1. LLMs as Database Administrators

- 이 논문은 텍스트 소스로부터 데이터베이스 유지 관리 경험을 지속적으로 획득하는 LLM 기반 프레임워크인 D-Bot을 제시합니다. D-Bot은 문서와 도구에서 데이터베이스 유지 관리 지식을 탐지하고, 원인 분석을 위한 생각의 트리 추론을 수행하며, 여러 LLMs 간의 협력적 진단을 돕습니다.

> **[LLM As DBA](https://arxiv.org/abs/2308.05481)**
>
> Database administrators (DBAs) play a crucial role in managing, maintaining and optimizing a database system to ensure data availability, performance, and reliability. However, it is hard and tedious for DBAs to manage a large number of database...

- 더보기: [https://twitter.com/omarsar0/status/1689811820272353280?s=20](https://twitter.com/omarsar0/status/1689811820272353280?s=20)

### 2. Political Biases Found in NLP Models

- 이 논문은 LLMs의 미디어 편향을 측정하는 방법을 개발하고, 정치적으로 편향된 LLMs 위에 조정된 하위 NLP 모델의 공정성을 포함합니다. 연구 결과, LLM는 기존 말뭉치에서의 극단화를 강화하는 정치적 경향성을 가지고 있음을 발견했습니다.

> **[From Pretraining Data to Language Models to Downstream Tasks: Tracking the...](https://aclanthology.org/2023.acl-long.656/)**
>
> Shangbin Feng, Chan Young Park, Yuhan Liu, Yulia Tsvetkov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.

- 더보기: [https://twitter.com/AiBreakfast/status/1688939983468453888?s=20](https://twitter.com/AiBreakfast/status/1688939983468453888?s=20)

### 3. Evaluating LLMs as Agents

- 이 논문은 LLM-as-Agent의 추론 및 의사결정 능력을 평가하기 위한 다차원 벤치마크인 AgentBench를 제시합니다. 결과적으로, 상업용 LLM와 오픈소스 LLMs 간에 에이전트로서의 역량을 테스트할 때 성능 차이가 크게 나타났으며, GPT-4는 지속적으로 학습하는 에이전트를 구축하는 데 잠재력을 보였습니다.

> **[AgentBench: Evaluating LLMs as Agents](https://arxiv.org/abs/2308.03688v1)**
>
> The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively \\textit{evaluate LLMs as agents} on challenging tasks in interactive environments. We present AgentBench, a...

- 더보기: [https://twitter.com/arankomatsuzaki/status/1688719837760000000?s=20](https://twitter.com/arankomatsuzaki/status/1688719837760000000?s=20)

### 4. Studying LLM Generalization with Influence Functions

- 이 논문은 영향 함수를 활용하여 LLM의 일반화 패턴을 조사하는 효율적인 방법을 소개합니다. 이 방법은 최대 52억 개의 파라미터를 가진 LLMs에 영향 함수를 확장하는 데 사용되며, 네트워크의 중간 계층이 가장 추상적인 일반화 패턴을 담당하는 것으로 나타났습니다.

> **[Studying Large Language Model Generalization with Influence Functions](https://arxiv.org/abs/2308.03296)**
>
> When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which training examples most contribute to a given behavior? Influence...

- 더보기: [https://twitter.com/AnthropicAI/status/1688946685937090560?s=20](https://twitter.com/AnthropicAI/status/1688946685937090560?s=20)

### 5. Seeing Through the Brain

- 이 논문은 EEG 신호로부터 시각 자극 이미지를 재구성하는 파이프라인인 NeuroImagen을 제안합니다. 잠재 확산 모델은 EEG 데이터를 취하고 고해상도 시각 자극 이미지를 재구성합니다.

> **[Seeing through the Brain: Image Reconstruction of Visual Perception from...](https://arxiv.org/abs/2308.02510)**
>
> Seeing is believing, however, the underlying mechanism of how human visual perceptions are intertwined with our cognitions is still a mystery. Thanks to the recent advances in both neuroscience and artificial intelligence, we have been able to record...

- 더보기: [https://twitter.com/\_akhaliq/status/1688787286807228416?s=20](https://twitter.com/_akhaliq/status/1688787286807228416?s=20)

### 6. SynJax

- 이 논문은 구조화된 분포에 대한 추론 알고리즘의 효율적인 벡터화 구현을 제공하는 새로운 라이브러리인 SynJax를 소개합니다. 이는 태깅, 세분화, 구성 트리, 스패닝 트리와 같은 데이터의 구조를 명시적으로 모델링하는 대규모 차별화 모델을 구축하는 데 사용될 수 있습니다.

> **[SynJax: Structured Probability Distributions for JAX](https://arxiv.org/abs/2308.03291v1)**
>
> The development of deep learning software libraries enabled significant progress in the field by allowing users to focus on modeling, while letting the library to take care of the tedious and time-consuming task of optimizing execution for modern...

- 더보기: [https://twitter.com/milosstanojevic/status/1688896558790520832?s=20](https://twitter.com/milosstanojevic/status/1688896558790520832?s=20)

### 7. Synthetic Data Reduces Sycophancy in LLMs

- 이 논문은 LLMs의 아첨을 줄이기 위해 간단한 합성 데이터에 미세 조정하는 방법을 제안합니다. 아첨은 LLMs가 사용자의 견해를 과도하게 따르려는 현상을 의미하며, 이는 사용자의 의견이 객관적으로 틀렸을 때에도 LLMs가 사용자의 견해를 반복하게 됩니다.

> **[Simple synthetic data reduces sycophancy in large language models](https://arxiv.org/abs/2308.03958)**
>
> Sycophancy is an undesirable behavior where models tailor their responses to follow a human user's view even when that view is not objectively correct (e.g., adapting liberal views once a user reveals that they are liberal). In this paper, we study...

- 더보기: [https://twitter.com/JerryWeiAI/status/1689340237993185280?s=20](https://twitter.com/JerryWeiAI/status/1689340237993185280?s=20)

### 8. Photorealistic Unreal Graphics (PUG)

- 이 논문은 Unreal Engine을 사용하여 표현 학습을 위한 사실적이고 의미론적으로 제어 가능한 합성 데이터셋을 제시합니다. 이는 사실적인 합성 데이터를 민주화하고, 비전 모델의 평가를 더 엄격하게 수행하는 것이 목표입니다.

> **[PUG: Photorealistic and Semantically Controllable Synthetic Data for...](https://arxiv.org/abs/2308.03977)**
>
> Synthetic image datasets offer unmatched advantages for designing and evaluating deep neural networks: they make it possible to (i) render as many data samples as needed, (ii) precisely control each scene and yield granular ground truth labels (and...

- 더보기: [https://twitter.com/MetaAI/status/1689316127846109184?s=20](https://twitter.com/MetaAI/status/1689316127846109184?s=20)

> [@Meta AI, Vision 모델을 위한 PUG(Photorealistic Unreal Graphics) 데이터셋 공개](https://discuss.pytorch.kr/t/meta-ai-vision-pug-photorealistic-unreal-graphics/2271):
>
> 최근 Meta에서 여러가지 모델들과 함께 데이터셋까지 활발하게 공개하고 있네요. sunglasses 이번에는 PUG라는, Photorealistic Unreal Graphics 데이터셋을 공개하였습니다. tada소개 역시나 제가 잘 모르는 분야라 김 굽듯이 살짝 설명드리자면 sweat_smile Photorealistic Unreal Graphics (PUG) 는 Meta AI에 의해 연구된 프로젝트로, 학습 및 테스트 시의 분포 이동을 제어하여 뉴럴넷이 변동 요인에 어떻게 일반화되는지를 알아보는데 도움을 줄 수 있다고 합니다. eyes (혹시 이 분야를 아시는 분 계시면 보충 설명 부탁드립니다 sweat_smile ) [CC-BY-NC 라이선스](https://creativecommons.org/licenses/by-nc/2.0/)로 배포하는데요, 설명이 어려우니;; 일단 데모 영상 보시고 가시죠! DeepL로 번역한 [PUG Datasheet](https://pug.metademolab.com/faq.html)의 데이터 소개 내용은 …

### 9. LLMs for Industrial Control

- 이 논문은 건물의 난방, 환기, 에어컨 등을 제어하는 등의 작업을 수행하는 데 GPT를 사용하는 데 필요한 데모와 프롬프트를 선택하고 생성하는 방법을 개발합니다. GPT-4는 RL 방법과 비교해 볼 때 적은 수의 샘플과 낮은 기술 부채를 사용하면서도 비슷한 성능을 보였습니다.

> **[Pre-Trained Large Language Models for Industrial Control](https://arxiv.org/abs/2308.03028)**
>
> For industrial control, developing high-performance controllers with few samples and low technical debt is appealing. Foundation models, possessing rich prior knowledge obtained from pre-training with Internet-scale corpus, have the potential to be a...

- 더보기: [https://twitter.com/emollick/status/1688760539441217536?s=20](https://twitter.com/emollick/status/1688760539441217536?s=20)

### 10. Trustworthy LLMs

- 이 논문은 LLM 신뢰성을 평가하는 데 중요한 카테고리와 하위 카테고리에 대한 포괄적인 개요를 제시합니다. 이러한 차원에는 신뢰성, 안전성, 공정성, 오용 저항성, 설명 가능성 및 추론, 사회 규범 준수, 강건성 등이 포함됩니다. 연구 결과, 정렬된 모델이 신뢰성 면에서 더 나은 성능을 보였지만, 정렬의 효과는 다양했습니다.

> **[Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language...](https://arxiv.org/abs/2308.05374)**
>
> Ensuring alignment, which refers to making models behave in accordance with human intentions \[1,2\], has become a critical task before deploying large language models (LLMs) in real-world applications. For instance, OpenAI devoted six months to...

- 더보기: [https://twitter.com/\_akhaliq/status/1689818964669390848?s=20](https://twitter.com/_akhaliq/status/1689818964669390848?s=20)

## 출처

> **[🥇Top ML Papers of the Week](https://nlp.elvissaravia.com/p/top-ml-papers-of-the-week-f42)**
>
> The top ML Papers of the Week (August 7 - August 13)
