Attention Is All You Need — Vaswani et al., 2017
Transformer의 시작. Self-Attention, Multi-Head Attention, positional encoding, encoder/decoder가 왜 나왔는지 보는 논문. 현대 LLM 구조의 출발점이라 무조건 1순위. (arXiv)
BERT: Pre-training of Deep Bidirectional Transformers — Devlin et al., 2018
Transformer 이후 pre-training → downstream fine-tuning 패러다임을 이해하기 좋음. GPT 계열과 구조가 다르지만 “왜 사전학습을 하는가?”를 이해하는 데 중요함. (arXiv)
Scaling Laws for Neural Language Models — Kaplan et al., 2020
이건 네가 요즘 보는 7B / 32B / 80B, GPU cost, inference cost 같은 얘기의 기초. Model size / dataset size / compute를 늘리면 loss가 어떻게 변하는지 empirical scaling law를 제시함. (arXiv)
Language Models are Few-Shot Learners (GPT-3) — Brown et al., 2020
현대적인 LLM이라는 개념을 사실상 정립한 논문 중 하나. Zero-shot / one-shot / few-shot, in-context learning, 거대한 autoregressive LM이 왜 범용 모델처럼 작동하는지를 보여줌. (arXiv)
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., 2020
바로 RAG 원조 논문. Parametric memory(모델 weight) + non-parametric memory(vector index/retriever)를 결합한다는 핵심 아이디어가 여기서 나옴. 지금 Pinecone/embedding/RAG 구조 이해하려면 읽을 가치 높음. (arXiv)
LoRA: Low-Rank Adaptation of Large Language Models — Hu et al., 2021
“왜 모델 전체를 fine-tuning하지 않고 adapter만 학습하지?”를 이해하는 논문. Pretrained weight를 freeze하고 작은 low-rank matrix만 학습해서 trainable parameter와 GPU memory를 크게 줄임. 현재 PEFT/QLoRA 계열의 핵심 기반. (arXiv)
Training Compute-Optimal Large Language Models (Chinchilla) — Hoffmann et al., 2022
LLM infra 쪽이면 특히 중요. “파라미터만 크게 하면 되는가?” → 아니다. 같은 compute라면 model size와 training tokens의 균형이 중요하다는 걸 보여줌. 70B Chinchilla가 같은 training compute에서 280B Gopher보다 좋은 결과를 냈음. (arXiv)
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Wei et al., 2022
CoT의 대표 논문. 단순히 답을 생성시키는 대신 intermediate reasoning steps를 생성시키면 복잡한 reasoning 성능이 크게 개선될 수 있다는 걸 보여줌. Prompt engineering / reasoning model 계보를 이해하는 출발점. (arXiv)
Training Language Models to Follow Instructions with Human Feedback (InstructGPT) — Ouyang et al., 2022
ChatGPT류 모델을 이해하려면 필수. Pretraining만 한 GPT가 어떻게 “사용자 지시를 따르는 assistant”가 되는지를 설명함. SFT → Reward Model → PPO/RLHF 파이프라인이 핵심. 1.3B InstructGPT가 human preference 평가에서 175B GPT-3보다 선호되는 결과도 보여줌. (arXiv)
Direct Preference Optimization (DPO) — Rafailov et al., 2023
RLHF 다음 세대 alignment를 이해하기 좋음. 복잡한 reward model + PPO 대신 preference pair를 이용해 훨씬 단순한 objective로 preference optimization을 수행하는 방법. 요즘 open-weight 모델 post-training을 볼 때 자주 등장함. (arXiv)
이걸 개념 계보로 보면 훨씬 쉬움.
Transformer
→ Attention Is All You Need
→ Pretraining
→ BERT
→ Scaling
→ Scaling Laws → GPT-3 → Chinchilla
→ LLM 활용
→ RAG / LoRA
→ Reasoning