# \[GN\] Otter: 컨텍스트 내에서 명령어 튜닝이 가능한 멀티모달 모델

**URL:** https://discuss.pytorch.kr/t/gn-otter/1813
**Category:** 읽을거리&정보공유
**Created:** [6월 14, 2023, 8:31오전 UTC](https://discuss.pytorch.kr/t/gn-otter/1813 "2023-06-14T08:31:06Z")
**Posts on this page:** 1
**Page:** 1

<div class="post-metadata">

### Author: ![9bow](https://discuss.pytorch.kr/user_avatar/discuss.pytorch.kr/9bow/32/16301_2.png) [@9bow](https://discuss.pytorch.kr/u/9bow)
#### Post date: [6월 14, 2023, 8:31오전 UTC](https://discuss.pytorch.kr/t/gn-otter/1813/1 "2023-06-14T08:31:07Z")

</div>

_[GeekNews](https://news.hada.io/)의 [xguru](https://news.hada.io/user?id=xguru)님께 허락을 받고 GN에 올라온 글들 중에 AI 관련된 소식들을 공유하고 있습니다. 😺_

* * *

### 소개

 ![image](https://discuss.pytorch.kr/uploads/default/original/2X/f/f6981c102a94bf7151f15267dc20a4400132e6c1.png)

- LLM의 제로샷 성능이 좋으려면 고품질 인스트럭션 셋이 필수적이고, VLM(시각-언어 모델)도 마찬가지
- 하지만 현재 vision-language 인스트럭션 셋은 수량/다양성/창의성 면에서 매우 제한적
- MIMIC-IT(MultI-Modal In-Context Instruction Tuning)을 제시
- 이미지 & 비디오 에서 가져온 220만개의 고유명령과, 280만개의 멀티모달 명령-응답 쌍으로 구성된 데이터 셋
- MIMIC-IT 데이터셋으로 훈련한 대규모 VLM이 Otter
- 8개 언어 지원: 영어, 중국어, 한국어, 일본어, 독일어, 프랑스어, 스페인어, 아랍어

### 원문

#### 데모 영상

![](https://img.youtube.com/vi/K8o_LKGQJhs/maxresdefault.jpg)https://www.youtube.com/embed/K8o_LKGQJhs?feature=oembed&wmode=opaque

#### 데모 사이트

[https://ottervideo.cliangyu.com/](https://ottervideo.cliangyu.com/)

#### 프로젝트 홈페이지

[https://otter-ntu.github.io](https://otter-ntu.github.io)

#### 논문

> **[MIMIC-IT: Multi-Modal In-Context Instruction Tuning](https://arxiv.org/abs/2306.05425)**
>
> High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language tasks involving intricate visual scenes, a large quantity of diverse...

### 출처

> **[Otter: 컨텍스트 내에서 명령어 튜닝이 가능한 멀티모달 모델 | GeekNews](https://news.hada.io/topic?id=9388)**
>
> LLM의 제로샷 성능이 좋으려면 고품질 인스트럭션 셋이 필수적이고, VLM(시각-언어 모델)도 마찬가지하지만 현재 vision-language 인스트럭션 셋은 수량/다양성/창의성 면에서 매우 제한적MIMIC-IT(MultI-Modal In-Context Instruction Tuning)을 제시이미지 & 비디오 에서 가져온 220만개의 고유명령과, 280만
