D02-02AM: Advanced Quantitative Text Analysis

Overview

This statistical approaches course covers advanced methods in quantitative text analysis. Through a combination of lectures and workshops, students will be introduced to a range of some of the latest tools and techniques, apply them to a variety of tasks, and learn how to critically reflect on both the methods themselves and their own results.

Applications are closed

About the Course

The course covers advanced methods in quantitative text analysis. Through a combination of lectures and workshops, students will be introduced to a range of some of the latest tools and techniques, apply them to a variety of tasks, and learn how to critically reflect on both the methods themselves and their own results.

This intensive course explores the latest methods for quantitative text analysis, with a particular focus on deep learning techniques, multilingual models, word embeddings and large language models (LLMs). The course is specifically designed for those who already have a basic understanding of text analysis and natural language processing (NLP) techniques. While the course touches on various aspects of machine learning and programming, there are no prerequisites other than a basic understanding of a programming language (preferably R or Python). The course includes a mix of lectures, where new topics are introduced, illustrated and discussed, and workshops, where students work on hands-on applications using real-world text analysis challenges and tools. By the end of the course, students will have the knowledge and skills to tackle advanced text analysis problems using the latest techniques.

Day 1: Introduction & Fundamentals

  • Review of basic NLP and text analysis concepts, such as tokenization, part-of-speech tagging, and syntactic parsing.
  • Introduction to deep learning fundamentals and neural networks: architecture, backpropagation, and gradient descent.
  • Exercise: Training a basic text classification model using deep learning.
  • Reading: Egami, Fong, Grimmer, Roberts and Stewart (2022) “How to make causal inferences using texts” 

Day 2: Word Embeddings

  • Introduction to word embeddings
  • Overview of models like Word2Vec, GloVe, and FastText.
  • Limitations of word embeddings and the evolution to contextual embeddings.
  • Exercise: Training custom word embeddings on a small corpus.
  • Reading: “Word Embeddings: What Works, What Doesn’t, and How to Tell the Difference for Applied Research” (Rodriguez and Spirling, 2022).

Day 3: Contextual Embeddings

  • Introduction to ELMo, BERT, and GPT.
  • Discuss promises and pitfalls of transfer learning and fine-tuning pre-trained models on specific tasks.
  • Benefits of pre-trained models for reducing the need for large labeled datasets. 
  • Exercise: Implementing BERT (and TopicBERT) for text classification and sentiment analysis tasks.
  • Reading: “BERT: a sentiment analysis odyssey ” (Alaparthi and Mishra, 2021).

Day 4: Multilingual Text Analysis and Large Language Models (LLMs)

  • Challenges and benefits of multilingual text analysis.
  • Introduction of mBERT (and XML-R).
  • Discuss how to use LLMs for advanced text analysis, such as text generation, summarization, translation, and question answering.
  • Ethical considerations of LLMs.
  • Exercise: Using mBERT for multilingual text classification.
  • Reading: “How multilingual is Multilingual BERT?” (Pires, Schlinger, Garrette, 2019).

Day 5: Fine-Tuning, Evaluation, and Real-World Applications

  • Techniques for fine-tuning pre-trained models on domain-specific tasks.
  • Evaluation metrics for text analysis models: precision, recall, F1 score, BLEU, ROUGE.
  • Performance benchmarking 
  • Applications of deep learning models (text generation, summarization, sentiment analysis, and chatbots).
  • Reading: “QuaLLM: An LLM-based Framework to Extract Quantitative Insights from Online Forums”

This description is subject to change at the discretion of the Instructor

2 Credits

For completion of all work before and during the course, as outlined by the Instructor, and 90% participation and attendance of the course.

2 Additional Credits

Course specific extra assignments. These can include submitting assignments before the course, daily assignments, and/or a final assignment to be completed after the course as decided by the Instructor.

Instructor

Bastiaan Bruinsma

sebastianus.bruinsma@chalmers.se

Bastiaan Bruinsma is a postdoctoral researcher focusing on computational social science and bias in AI systems, especially large language models, improving automated text analysis for low-resource languages, and the impact of policy decisions on AI governance.

Bruinsma Pic

Pricing

15% off During Early Bird!

  • Student Member

    699.30
  • Student Non-Member

    999.00
  • Other Member

    849.15
  • Other Non-Member

    999.00

Secure Your Place!

Please complete this webform for your registration.

Registration

Important Information

  1. Complete in English only 
  2. Do not complete in capital letters
  3. Course fees are reduced for MethodsNET members
  4. If you are not a MethodsNET Member at the time of completing this form you will pay the non-member fee
  5. If your institution or organization is paying for your course, complete the correct invoice information 
  6. Please note that your seat in a course is only reserved and guaranteed after you submit full payment of the registration fee.
Course Selection

You can select one Online + one All Day in each week or one Online + one AM & PM in each week

Address
Invoice Address (if different from above)

By submitting the form, you are agreeing with our Terms & Conditions, Code of conduct, and Privacy Policy.