D01-07PM: Introduction to R

Overview

This statistical approaches course covers the basics of programming in R: importing data, performing the necessary data cleaning and transformation steps, creating subsets or merging datasets with other external sources of data. Some of these steps might be a bit different depending on the type of data you intend to work with, but we discuss these differences and show examples of working with all of them.

Applications are closed

About the Course

The course covers the basics of programming in R: importing data, performing the necessary data cleaning and transformation steps, creating subsets or merging datasets with other external sources of data. Some of these steps might be a bit different depending on the type of data you intend to work with, but we discuss these differences and show examples of working with all of them. Knowing that some participants will come with some experience using other statistical software, we cover how to import and export data to proprietary (STATA, SPSS, etc.) formats. The course does not cover statistical theory, but we’ll go through some applied tools for data analysis and hypothesis testing, such as comparing groups and regressions. Topics on data visualization form an inherent part of this course: relying on ggplot2, we’ll create scatterplots, illustrate temporal trends, present distributions in intuitive ways and express regression results concisely with coefficient plots. A session is dedicated to using RMarkdown, an essential workflow allowing you to generate an entire research paper or report. Following a learning-by-doing approach, course participants will have a chance to solve coding exercises both in teams and alone, and to test their newly acquired knowledge via assignments between sessions.  

Introduction to R

Day 1

Why use R? How do you get help when feeling stuck? We’ll cover the basics of R and RStudio, learn what object-oriented programming is all about, perform basic operations, and review the fundamental data types. The first session will already enable you to import and export datasets, even if they come from proprietary software (STATA, SPSS, etc.). By the end of Monday afternoon, you’ll also be able to have a glimpse of (potential) relationships between variables by generating a histogram or a scatterplot. As most of the course will rely on packages (collections of useful functions for data wrangling, visualization and more), the session will end with a hands-on guide on how to install and load them.

Day 2

he course will provide you with the skills necessary for data cleaning. In this session, we will merge and join separate data frames, create subsets, filter out specific observations, delete and create new variables, and remove duplicates. We will cover the fundamentals of working with missing data, dates, and character input. Converting long data frames into wide ones (or the other way around) will help us restructure our datasets into a form that best suits our analytical purposes.

Day 3

The course will give you the applied skills to perform basic statistical tests and to run some quantitative models. We will compare group means, fit OLS regressions, and learn how to interpret the output R generates for us. In order to make such output submission-ready for journals and dissertations, we will harness the power of stargazer, a package that provides you with neatly formatted tabular outputs from regression models. Some time will be dedicated to assessing (potential) violations of regression assumptions using both visual and numeric information.

Day 4

After this session, course participants will be able to write their very own functions. The material of the afternoon will also cover how to iterate the same action over a series of inputs via for-loops and the purrr package. In the second half of the afternoon, we will start our introduction to data visualization with ggplot2. The session demonstrates how to make some of the most common plots, which can effectively convey underlying patterns, associations, and differences in your data.

Day 5

The remainder of the course is dedicated to the best practices for data visualization, as well as an introduction to R Markdown, which allows you to neatly combine code and output with text and references, providing an all-in-one solution for writing up research reports. The session covers box plots, violin plots, line charts, visualizations of regression coefficients, predicted probabilities, and marginal effects. Additionally, examples will illustrate how to modify labels and themes, as well as how to apply different colors, shapes, or types to a particular subset(s) of your data, highlighting subgroup heterogeneity.

Further readings (none are required, but you might want to consult them before / after the course): 

Kieran Healy: Data visualization. A practical introduction. Princeton University Press

Robert I. Kabacoff: R in Action. Data analysis and graphics with R and Tidyverse. Manning. 

Hadley Wickham, Garrett Grolemund, Mine Cetinkaya–Rundel: R for Data Science. O′Reilly.

This description is subject to change at the discretion of the Instructor

2 Credits

For completion of all work before and during the course, as outlined by the Instructor, and 90% participation and attendance of the course.

2 Additional Credits

Course specific extra assignments. These can include submitting assignments before the course, daily assignments, and/or a final assignment to be completed after the course as decided by the Instructor.  

Instructor

Daniel Kovarek

daniel.kovarek@eui.eu

Daniel Kovarek is a Postdoctoral Research Fellow at the Robert Schuman Centre for Advanced Studies at the European University Institute in Florence. He received his PhD in Political Science from the Central European University in Vienna, where his dissertation was recognized with the Outstanding Dissertation Award in 2023. Daniel studies political behavior at the voter and the elite level, applying surveys, experimental and big data methods. Substantively, his main research interests include distributive politics, voting behavior and representation. His work has appeared in journals such as Political Geography, Research & Politics and Democratization. Daniel has been teaching a wide variety of graduate-level courses on applied statistics, research design, data visualization and programming.

Kovarek Pic

Pricing

15% off During Early Bird!

  • Student Member

    699.30
  • Student Non-Member

    999.00
  • Other Member

    849.15
  • Other Non-Member

    999.00

Secure Your Place!

Please complete this webform for your registration.

Registration

Important Information

  1. Complete in English only 
  2. Do not complete in capital letters
  3. Course fees are reduced for MethodsNET members
  4. If you are not a MethodsNET Member at the time of completing this form you will pay the non-member fee
  5. If your institution or organization is paying for your course, complete the correct invoice information 
  6. Please note that your seat in a course is only reserved and guaranteed after you submit full payment of the registration fee.
Course Selection

You can select one Online + one All Day in each week or one Online + one AM & PM in each week

Address
Invoice Address (if different from above)

By submitting the form, you are agreeing with our Terms & Conditions, Code of conduct, and Privacy Policy.