D02-02PM: Large Language Models – Tune and Train your own Models

Overview

This statistical approaches course empowers researchers to take control of these powerful tools, moving beyond off-the-shelf solutions to create customized models that directly serve their research objectives.

Applications are closed

About the Course

​​As Large Language Models (LLMs) increasingly become indispensable tools in academic research across disciplines, from computational linguistics to political research, the ability to adapt these models to specific research contexts has become a crucial skill for modern researchers. While pre-trained LLMs offer impressive general capabilities, their true potential in research applications is unlocked through careful tuning and adaptation to domain-specific tasks and datasets. This intensive five-day summer school course empowers researchers to take control of these powerful tools, moving beyond off-the-shelf solutions to create customized models that directly serve their research objectives. Designed for researchers with basic machine learning knowledge, this hands-on course bridges the gap between theoretical understanding and practical implementation. Participants will learn about the architectural foundations of modern LLMs, explore state-of-the-art open-source models, and master parameter-efficient techniques for model optimization. The course emphasizes practical skills vital for research applications, with participants gaining hands-on experience using the HuggingFace ecosystem and implementing various tuning strategies that can be applied to their specific research domains. Through a combination of lectures, interactive sessions, and guided practical exercises, attendees will learn to prepare custom research datasets, implement efficient training procedures, and evaluate model performance. Special attention is given to resource-efficient approaches, making LLM development accessible even within typical academic computing constraints. By the course’s end, participants will have the practical skills needed to adapt LLMs to their research questions, understand when and how to apply different optimization techniques, and evaluate model performance in their specific research contexts.

Day 1 (Foundations of Large Language Models)

Morning session focuses on understanding the fundamental architecture of modern LLMs, exploring key components like attention mechanisms, decoder architectures, and their distinct advantages over encoder-only models like BERT. Participants will dive into the evolution of language models, examining breakthrough architectures and their impact on current state-of-the-art systems. The afternoon session surveys leading open-source models (like Llama, Mistral, and Pythia), analyzing their unique characteristics, strengths, and potential applications. Practical examples demonstrate how different architectures influence model behavior and performance.

Day 2 (Development Environment and Optimization Techniques)

The day begins with a comprehensive setup of the development environment using HuggingFace’s ecosystem. Participants learn to navigate the Model Hub, use the Transformers library, and set up efficient workflows for model development. The second half introduces parameter-efficient techniques, with detailed coverage of various quantization approaches (like 4-bit and 8-bit quantization). Hands-on sessions include implementing these techniques, understanding their trade-offs, and measuring their impact on model performance and resource usage. Special attention is given to QLoRA and other efficient fine-tuning methods.

Day 3 (Dataset Preparation and Customization)

This day focuses on data – the cornerstone of effective model tuning. Morning sessions explore existing instruction-tuning datasets, analyzing their structure and characteristics. Participants learn to evaluate dataset quality and identify potential biases. The afternoon is dedicated to practical exercises in dataset preparation, including data cleaning, formatting, and augmentation techniques. Attendees work on creating custom validation sets aligned with their research interests, learning best practices for dataset splitting and evaluation metric selection. Special emphasis is placed on creating high-quality instruction-following datasets.

Day 4 (Implementation and Training)

A highly practical day centered on implementing training pipelines. Morning sessions cover code setup for efficient training, including configuration of training parameters, loss functions, and optimization strategies. Participants learn to implement various fine-tuning approaches, from full fine-tuning to parameter-efficient methods. The afternoon focuses on launching and monitoring training runs, implementing evaluation loops, and creating comprehensive benchmarks using custom validation sets. Practical exercises include debugging common training issues and optimizing training efficiency.

Day 5 (Analysis and Strategic Planning)

The final day synthesizes the week’s learning through critical analysis and strategic planning. Morning sessions focus on evaluating training results, comparing different approaches, and understanding the trade-offs between model size, computational resources, and performance. Participants analyze when smaller, efficient models might outperform larger ones for specific tasks. The afternoon includes a reflective workshop where attendees develop strategies for implementing LLMs in their own research or applications, considering factors like computational constraints, task requirements, and deployment scenarios. The day concludes with presentations and discussions of participant-specific use cases and implementation plans.

This description is subject to change at the discretion of the Instructor

2 Credits

For completion of all work before and during the course, as outlined by the Instructor, and 90% participation and attendance of the course.

2 Additional Credits

Course specific extra assignments. These can include submitting assignments before the course, daily assignments, and/or a final assignment to be completed after the course as decided by the Instructor.  

Instructor

Christopher Klamm

christopher@klamm.info

Christopher is a researcher at the University of Cologne in Germany, with a background that spans across several disciplines. His research is focused on analyzing rhetoric, framing, and populism in text, at the intersection of NLP and Computational Political Science. Christopher holds two master’s degrees - one in Computer Science with a minor in Philosophy, and another in Political Science, both from the Technical University of Darmstadt. During his academic journey, he spent a semester abroad studying at the ETH Zurich and the University of Zurich. Christopher is an active participant in various open-source initiatives, such as the BigScience BLOOM project, the Data Provenance Initiative, and the Aya Expedition. He advocates for open research across all disciplines. Christopher is also a co-organizer of the tada.cool speaker series and organizes a specialized workshop (CPSS) on NLP for social and political sciences. The course emphasizes practical skills vital for research applications, with participants gaining hands-on experience using the HuggingFace ecosystem and implementing various tuning strategies that can be applied to their specific research domains.

Klamm

Pricing

15% off During Early Bird!

  • Student Member

    699.30
  • Student Non-Member

    999.00
  • Other Member

    849.15
  • Other Non-Member

    999.00

Secure Your Place!

Please complete this webform for your registration.

Registration

Important Information

  1. Complete in English only 
  2. Do not complete in capital letters
  3. Course fees are reduced for MethodsNET members
  4. If you are not a MethodsNET Member at the time of completing this form you will pay the non-member fee
  5. If your institution or organization is paying for your course, complete the correct invoice information 
  6. Please note that your seat in a course is only reserved and guaranteed after you submit full payment of the registration fee.
Course Selection

You can select one Online + one All Day in each week or one Online + one AM & PM in each week

Address
Invoice Address (if different from above)

By submitting the form, you are agreeing with our Terms & Conditions, Code of conduct, and Privacy Policy.