Schedule
-
EventDateDescriptionCourse Material
-
Lecture08/25/2026
TuesdayIntroduction & Fundamentals[Slides]Main readings:
-
Lecture08/27/2026
ThursdayFundamentals: Learned Representations[Slides]Main readings:
Additional references
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (Reimers & Gurevych 2019)
- SimCSE: Simple Contrastive Learning of Sentence Embeddings (Gao et al 2021)
-
Lecture09/01/2026
TuesdayFundamentals: Autoregressive Language Modeling[Slides]Main readings:
Additional references
- A Neural Probabilistic Language Model (Bengio et al 2003)
- Understanding the difficulty of training deep feedforward neural networks (Glorot & Bengio 2010)
-
Assignment09/01/2026
TuesdayAssignment #1 released -
Lecture09/03/2026
ThursdayRecurrent Neural NetworksMain readings:
- Natural Language Understanding with Distributed Representation (Ch. 4, Ch. 5.5-5.6, Ch. 6) (Cho 2015)
Additional references
- Recurrent neural network based language model (Mikolov et al 2010)
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation (Cho et al 2014)
- Why LSTMs Stop Your Gradients From Vanishing: A View from the Backwards Pass (Weber 2017)
- Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau et al 2015)
-
Lecture09/08/2026
TuesdayAttention and TransformersMain readings:
- Attention Is All You Need (Vaswani et al 2017)
- The Annotated Transformer (Rush et al 2018)
Additional references
- Root Mean Square Layer Normalization (Zhang & Sennrich 2019)
- On Layer Normalization in the Transformer Architecture (Xiong et al 2020)
- RoFormer: Enhanced Transformer with Rotary Position Embedding (Su et al 2021)
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (Ainslie et al 2023)
- (Helpful Blog Post): Why Are Sines and Cosines Used For Positional Encoding? (Muhammad 2023)
- The Illustrated Transformer (Alammar 2018)
- (Interactive) Transformer Explainer: LLM Transformer Model Visually Explained (Cho et al 2024)
- (Interactive) LLM Visualization (Bycroft 2023)
-
Lecture09/10/2026
ThursdayTokenization and Decoding AlgorithmsMain readings:
- Neural Machine Translation of Rare Words with Subword Units (Sennrich et al 2016)
- From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models (Sections 1-3) (Welleck et al 2024)
Additional references
- (Video) Letโs build the GPT Tokenizer (Karpathy 2024)
-
Lecture09/15/2026
TuesdayPretrainingMain readings:
- Language Models are Unsupervised Multitask Learners (Radford et al 2019)
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale (Penedo et al 2024)
Additional references
- OLMo 3 (AI2 2025)
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (Devlin et al 2018)
- LLaMA: Open and Efficient Foundation Language Models (Touvron et al 2023)
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text (Paster et al 2023)
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research (Soldaini et al 2024)
- Language Modeling Is Compression (Delรฉtang et al 2023)
-
Lecture09/17/2026
ThursdayScaling Laws and In-Context LearningMain readings:
- Language Models are Few-Shot Learners (Brown et al 2020)
- Deep Learning Scaling is Predictable, Empirically (Hestness et al 2017)
Additional references
- Scaling Laws for Neural Language Models (Kaplan et al 2020)
- Training Compute-Optimal Large Language Models (Hoffmann et al 2022)
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism (Deepseek AI 2024)
-
Lecture09/22/2026
TuesdayFine-tuning and DistillationMain readings:
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al 2021)
- Sequence-Level Knowledge Distillation (Kim & Rush 2016)
Additional references
- Universal Language Model Fine-tuning for Text Classification (Howard & Ruder 2018)
- Cross-Task Generalization via Natural Language Crowdsourcing Instructions (Mishra et al 2021)
- Finetuned Language Models Are Zero-Shot Learners (Wei et al 2021)
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks (Wang et al 2022)
- Self-Instruct: Aligning Language Models with Self-Generated Instructions (Wang et al 2023)
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4 (Mukherjee et al 2023)
- Symbolic Knowledge Distillation: from General Language Models to Commonsense Models (West et al 2022)
- QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al 2023)
-
Lecture09/24/2026
ThursdayReasoning and Test-Time ScalingMain readings:
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (Wei et al 2022)
- From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models (Sections 4-7) (Welleck et al 2024)
Additional references
- NeurIPS 2024 LLM Inference Tutorial (Reading List)
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI 2025)
- s1: Simple test-time scaling (Muennighoff et al 2025)
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning (Aggarwal & Welleck 2025)
-
Assignment09/25/2026
FridayAssignment #2 released -
Due09/24/2026 11:59 PM
ThursdayAssignment #1 due -
Lecture09/29/2026
TuesdayModeling I: Retrieval and RAGMain readings:
- Retrieval-based Language Models and Applications (Asai et al 2023)
- Dense Passage Retrieval for Open-Domain Question Answering (Karpukhin et al 2020)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al 2020)
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection (Asai et al 2024)
Additional references
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories (Mallen & Asai et al 2023)
- Reliable, Adaptable, and Attributable Language Models with Retrieval (Asai et al 2024)
- Scaling Retrieval-Based Language Models with a Trillion-Token Datastore (Shao et al 2024)
-
Lecture10/01/2026
ThursdayString searchMain readings:
-
Lecture10/06/2026
TuesdayMultimodal models -
Exam10/08/2026 00:00
Thursday -
No class10/13/2026
Tuesday๐ Fall Break ๐ -
No class10/15/2026
Thursday๐ Fall Break ๐ -
Lecture10/20/2026
TuesdayResearch Skills and Experimental DesignMain readings:
-
Lecture10/22/2026
ThursdayBenchmarking and Evaluation TechniquesMain readings:
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (Miller 2024)
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference (Chiang et al 2024)
Additional references
- Holistic Evaluation of Language Models (Liang et al 2023)
- Measuring Massive Multitask Language Understanding (Hendrycks et al 2021)
-
Assignment10/23/2026
FridayAssignment #3 released -
Assignment10/23/2026
FridayAssignment #4 released -
Due10/22/2026 11:59 PM
ThursdayAssignment #2 due -
Lecture10/27/2026
TuesdayDiffusion and FlowsMain readings:
- Denoising Diffusion Probabilistic Models (Ho et al 2020)
- Flow Matching for Generative Modeling (Lipman et al 2022)
-
Lecture10/29/2026
ThursdayReinforcement Learning I: FundamentalsMain readings:
- Deep Reinforcement Learning: Pong from Pixels (Karpathy 2016)
- Spinning Up in Deep RL (Part 1, Part 3, Vanilla PG, PPO) (OpenAI)
Additional references
- Proximal Policy Optimization Algorithms (Schulman et al 2017)
- High-Dimensional Continuous Control Using Generalized Advantage Estimation (Schulman et al 2015)
-
Lecture11/03/2026
TuesdayReinforcement Learning II: ApplicationsMain readings:
- Training language models to follow instructions with human feedback (Ouyang et al 2022)
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek-AI 2025)
Additional references
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
- Deep reinforcement learning from human preferences (Christiano et al 2017)
- Fine-Tuning Language Models from Human Preferences (Ziegler et al 2019)
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
-
No class11/05/2026
ThursdayDemocracy Day -
Lecture11/10/2026
TuesdayLanguage Model-Based AgentsMain readings:
- World of Bits: An Open-Domain Platform for Web-Based Agents (Shi et al 2017)
- WebGPT: Browser-assisted question-answering with human feedback (Nakano et al 2022)
Additional references
- ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al 2023)
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents (Yao et al 2022)
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (Yang et al 2024)
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks (Koh et al 2024)
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (Xie et al 2024)
- ฯ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (Yao et al 2024)
- Programming with Pixels: Computer-Use Meets Software Engineering (Aggarwal & Welleck 2025)
-
Lecture11/12/2026
ThursdayEfficiency: Quantization, Parallelism, and Distributed TrainingMain readings:
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (Dettmers et al 2022)
- QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al 2023)
- The Ultra-Scale Playbook: Training LLMs on GPU Clusters (Tazi et al 2025)
Additional references
- 8-bit Optimizers via Block-wise Quantization (Dettmers et al 2021)
- The case for 4-bit precision: k-bit Inference Scaling Laws (Dettmers & Zettlemoyer 2022)
-
Due11/12/2026 11:59 PM
ThursdayAssignment #3 due -
Lecture11/17/2026
TuesdayLooking at dataMain readings:
- Deduplicating Training Data Makes Language Models Better (Lee et al 2022)
- Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics (Swayamdipta et al 2020)
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (Dodge et al 2021)
-
Lecture11/19/2026
ThursdayMixture of ExpertsMain readings:
- A Review of Sparse Expert Models in Deep Learning (Fedus et al 2022)
- OLMoE: Open Mixture-of-Experts Language Models (Muennighoff et al 2024)
Additional references
-
Lecture11/24/2026
TuesdayScaling Sequence LengthMain readings:
- Self-attention Does Not Need O(n2) Memory (Rabe & Staats 2021)
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces (Gu & Dao 2023)
Additional references
-
No class11/26/2026
Thursday๐ฆ Thanksgiving ๐ฆ -
Lecture12/01/2026
TuesdayMultilingual NLPMain readings:
- Unsupervised Cross-lingual Representation Learning at Scale (Conneau et al 2020)
- mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer (Xue et al 2021)
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages (Chang et al 2024)
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models (Ahia et al 2023)
-
Lecture12/03/2026
ThursdayPoster Session -
Due12/10/2026 11:59 PM
ThursdayAssignment #4 due
