← R for Data Analysis Expert Β· Lesson 14 of 17

Module Thirteen

πŸ“– Every lesson in this course is free to read right here, no account needed. Create a free account to track your progress, take the exam, and earn your certificate.
1

Course Outline

R for Data Analysis Expert Β· Course Outline

πŸ“Š R for Data Analysis Expert Expert Β· R Β· Statistics

Master R for data analysis β€” data manipulation, visualization, statistics & machine learning 8 modules
🎯 target: data analysts, statisticians, data scientists ⏳ duration: 8 weeks Β· hybrid hands‑on Β· project‑based

1. Introduction to R – foundation getting started

Understand the R ecosystem – language, RStudio, packages, and the data analysis workflow.

  • What is R? vs. Python vs. Excel
  • RStudio IDE: Console, Editor, Environment, Plots
  • R Packages: installing, loading, and managing packages
  • Basic R syntax: objects, functions, vectors, data types
  • Nigerian context: analysing local business and government data

2. Data Import & Cleaning wrangle

Import and clean data – readr, tidyr, and the tidyverse.

  • Importing data: CSV, Excel, databases, JSON, web APIs
  • Data cleaning: handling missing values, outliers, and inconsistent formats
  • Data reshaping: pivot_longer, pivot_wider, separate, unite
  • dplyr basics: select, filter, mutate, summarise, arrange
  • Nigerian context: cleaning Nigerian business and government datasets

3. Data Visualization with ggplot2 communicate

Master data visualization – ggplot2, themes, and advanced graphics.

  • Grammar of Graphics: ggplot, geoms, aesthetics, facets, themes
  • Common plots: bar charts, scatter plots, line charts, histograms, boxplots
  • Advanced visuals: faceting, custom themes, annotations, and labels
  • Saving and exporting: ggsave and customizing for presentations
  • Nigerian context: visualizing Nigerian economic and demographic data

4. Data Manipulation with dplyr & tidyr transform

Advanced data wrangling – dplyr, tidyr, and the tidyverse.

  • dplyr verbs: select, filter, mutate, summarise, arrange, group_by
  • Joins: inner, left, right, full, semi, anti
  • Tidyr: pivot_longer, pivot_wider, separate, unite, complete
  • Window functions: row_number, rank, lag, lead
  • Nigerian context: wrangling Nigerian sales, customer, and financial data

5. Statistical Analysis & Hypothesis Testing infer

Perform statistical analysis – descriptive statistics, hypothesis tests, and confidence intervals.

  • Descriptive statistics: mean, median, mode, sd, summary, quantile
  • Hypothesis testing: t-test, ANOVA, chi-square, correlation tests
  • Confidence intervals: calculating and interpreting
  • Linear regression: lm, summary, diagnostic plots
  • Nigerian context: analysing business, health, and economic data

6. Advanced Modeling & Machine Learning predict

Build predictive models – regression, classification, and model evaluation.

  • Linear and logistic regression: glm, lm, model diagnostics
  • Decision trees: rpart, randomForest, xgboost
  • Model evaluation: train/test split, cross-validation, confusion matrix, ROC
  • Feature engineering: creating and selecting features
  • Nigerian context: predicting customer churn, sales, and risk

7. Reporting with R Markdown & Shiny communicate

Share your analysis – R Markdown, Shiny, and interactive dashboards.

  • R Markdown: creating reproducible reports, HTML, PDF, Word
  • Shiny basics: ui, server, reactive expressions
  • Interactive dashboards: Shiny dashboards and flexdashboard
  • Parameterized reports: dynamic reporting with parameters
  • Nigerian context: building reports and dashboards for Nigerian businesses

8. Career & Certification success

Prepare for certification – R certifications, portfolio building, and career paths.

  • R certifications: DataCamp, Coursera, and other certifications
  • Building a portfolio: GitHub, R Pubs, and personal projects
  • Career paths: Data Analyst β†’ Data Scientist β†’ Analytics Manager
  • Interview preparation: common R and statistics questions
  • Nigerian context: opportunities and demand for R experts
🎯 Capstone: End‑to‑End R Project β€” import, clean, visualize, model, and report on a real Nigerian dataset hands‑on Β· portfolio
πŸ“š included: datasets Β· templates Β· case studies πŸŽ“ certification: aligned with DataCamp and Coursera R certifications

bold & italic used for emphasis Β· course outline v1.0
2

Module One

Module One: Introduction to R

πŸ“Š Module One: Introduction to R

Welcome, young data explorer! Let us begin your journey into the world of R.

🌟 Module Introduction

Welcome to the Introduction to R module! Have you ever wondered how businesses turn piles of numbers into beautiful charts and insights? The answer is R.

R is a powerful tool that helps people understand their data. It is like a super-smart assistant that takes your data and turns it into colorful charts, graphs, and dashboards that tell a story.

Imagine you have a big box of Lego bricks. You can build anything you want with them. R is like that β€” it takes your data (the Lego bricks) and helps you build amazing visual stories.

In this module, you will learn what R is, why it is so important, and how to get started. You will discover the R ecosystem, the interface, and how to create your first report. By the end, you will understand the basics of R and be ready to explore more.

πŸ’‘ Think about it: Have you ever seen a chart or a graph that made you understand something better? That is what R does β€” it turns numbers into pictures that make sense.

🎯 Learning Objectives

By the time you finish this module, you will be able to:

  • Explain what R is in simple words.
  • Understand why R is important for data analysis.
  • Install R and RStudio on your computer.
  • Navigate the RStudio interface.
  • Write and run basic R code.
  • Import data into R.
  • Create a simple vector and perform basic operations.
  • Give examples of how R is used in Nigeria.

πŸ“– Warm-up Story: Chidi's Big Discovery

Chidi is a 12-year-old boy living in Enugu. He loves to collect data β€” he counts how many cars pass by his house, how many mangoes fall from the tree, and how many goals his football team scores. He writes everything down in his notebook.

One day, his uncle visited him. His uncle works in a big company in Lagos. Chidi showed his uncle all the numbers in his notebook. His uncle smiled and said, "Chidi, these numbers are amazing! But they are hard to understand. What if you could turn them into beautiful pictures?"

His uncle opened his laptop and showed Chidi a tool called R. He took Chidi's numbers and turned them into colorful charts and graphs. Chidi could see at a glance how many cars passed on different days, which month had the most mangoes, and when his team scored the most goals.

Chidi was fascinated. He said, "Uncle, this is like magic! Numbers become pictures!" His uncle replied, "It is not magic β€” it is R. And you can learn it too."

🧠 Think about it: Have you ever collected data? What did you do with it? Could you turn it into a picture?

πŸ“š Main Lessons

1. What is R?

Definition: R is a programming language used for data analysis, statistics, and creating charts.

Why it is important: R helps people make better decisions by understanding their data. It is like a magnifying glass for numbers!

Simple explanation: Think of R as a magical paintbrush. You give it data, and it paints you a picture that tells a story.

🏫 School example: A teacher can use R to see which students are improving and which need extra help.

🏠 Home example: Your parents could use R to see how much they spend on groceries each month.

πŸ‡³πŸ‡¬ Nigerian example: A Nigerian business owner can use R to see which products are selling best.

Illustration:

    +-------------------+
    |  YOUR DATA        |  ← Numbers, sales, records
    +-------------------+
           |
           V
    +-------------------+
    |  R                |  ← The magic tool
    +-------------------+
           |
           V
    +-------------------+
    |  BEAUTIFUL CHARTS |  ← Pictures that tell a story
    +-------------------+
    

πŸ“Œ Mini summary: R is a tool that turns data into pictures to help you understand it better.

2. Why is R Important?

Definition: R is important because it helps people make sense of data. Data is everywhere, but understanding it can be hard. R makes it easy.

Why it is important: Without tools like R, businesses would have to look at thousands of numbers in spreadsheets. That takes a long time and is hard to understand.

Simple explanation: Imagine reading a book with no pictures. It is still a good book, but pictures make it more interesting and easier to understand. R is like the pictures for your data.

  • It saves time: You can see trends and patterns quickly.
  • It helps make decisions: You can see what is working and what is not.
  • It is interactive: You can click and explore your data.
  • It is everywhere: Businesses all over the world use R.

🏫 School example: A school principal can use R to see attendance patterns and plan better.

🏠 Home example: Your family can use R to track savings and spending.

πŸ‡³πŸ‡¬ Nigerian example: A Nigerian bank can use R to see which branches are performing best.

πŸ“Œ Mini summary: R is important because it helps people understand data quickly and easily.

3. Installing R and RStudio

Definition: R is the programming language. RStudio is the tool we use to write R code.

Why it is important: You need both to start using R.

Simple explanation: Think of R like the engine of a car and RStudio like the steering wheel and dashboard.

  • Step 1: Download R from the R Project website.
  • Step 2: Install R on your computer.
  • Step 3: Download RStudio from the RStudio website.
  • Step 4: Install RStudio on your computer.
  • Step 5: Open RStudio and start exploring!

🏫 School example: A teacher installs R and RStudio to start analyzing student grades.

🏠 Home example: Your parents install R and RStudio to track their expenses.

πŸ‡³πŸ‡¬ Nigerian example: A business analyst installs R and RStudio to analyze sales data.

πŸ“Œ Mini summary: R is the language; RStudio is the tool you use to write R code.

4. Navigating RStudio

Definition: RStudio is the tool you use to write and run R code.

Why it is important: Knowing your way around RStudio makes it easier to work with R.

Simple explanation: Think of RStudio like the cockpit of a plane. Each window has a job.

  • Console: Where you type commands and see results.
  • Editor: Where you write and save scripts.
  • Environment: Shows your data and variables.
  • Plots: Shows your charts and graphs.
  • Help: Provides documentation and help for functions.
  • Files: Shows the files in your working directory.

🏫 School example: A teacher uses the Editor to write code and the Console to see results.

🏠 Home example: Your parents use the Environment to see their data.

πŸ‡³πŸ‡¬ Nigerian example: A business analyst uses the Plots pane to see charts.

    +---------------------------------------------------+
    |  EDITOR          |  CONSOLE                       |
    |  (Write code)    |  (Run code)                   |
    +---------------------------------------------------+
    |  ENVIRONMENT     |  PLOTS                        |
    |  (View data)     |  (View charts)               |
    +---------------------------------------------------+
    

πŸ“Œ Mini summary: RStudio has different panes for writing code, viewing data, and seeing results.

5. Basic R Syntax

Definition: Syntax is the set of rules for writing R code.

Why it is important: You need to learn the rules to write R code correctly.

Simple explanation: Think of syntax like the grammar of a language.

  • Assignment: Use <- to assign values to variables.
  • Functions: Use functions like sum(), mean(), and plot().
  • Comments: Use # to add notes to your code.
  • Vectors: Use c() to create a vector.
  • Data types: numeric, character, logical, factor, date.
# Example R code x <- 5 y <- 10 z <- x + y print(z) # Creating a vector numbers <- c(1, 2, 3, 4, 5) mean(numbers)

🏫 School example: A teacher uses <- to store student scores.

🏠 Home example: Your parents use <- to store expense amounts.

πŸ‡³πŸ‡¬ Nigerian example: A business uses <- to store sales data.

πŸ“Œ Mini summary: R syntax is the set of rules for writing R code.

6. Importing Data

Definition: Importing data means loading data into R.

Why it is important: You cannot analyze data without loading it first.

Simple explanation: Think of importing data like opening a book before you read it.

  • CSV: read.csv() β€” comma-separated values.
  • Excel: readxl package β€” read_excel().
  • Text: read.table() β€” text files.
  • Web: read.csv() from a URL.
  • Database: DBI package for databases.
# Importing a CSV file sales_data <- read.csv("sales.csv") # View the data head(sales_data) summary(sales_data)

🏫 School example: A teacher imports student grades from a CSV file.

🏠 Home example: Your parents import expense data from an Excel file.

πŸ‡³πŸ‡¬ Nigerian example: A business imports sales data from a CSV file.

πŸ“Œ Mini summary: Importing data is the first step in analysis.

7. Creating Vectors

Definition: A vector is a collection of values of the same type.

Why it is important: Vectors are the building blocks of data in R.

Simple explanation: Think of a vector like a list of numbers.

  • Create: Use c() to create a vector.
  • Example: numbers <- c(1, 2, 3, 4, 5).
  • Operations: You can add, subtract, multiply, and divide vectors.
# Creating vectors numbers <- c(1, 2, 3, 4, 5) doubled <- numbers * 2 sum_numbers <- sum(numbers) mean_numbers <- mean(numbers)

🏫 School example: A teacher creates a vector of student grades.

🏠 Home example: Your parents create a vector of expense amounts.

πŸ‡³πŸ‡¬ Nigerian example: A business creates a vector of sales figures.

πŸ“Œ Mini summary: Vectors are collections of values.

8. Basic Operations

Definition: Basic operations include addition, subtraction, multiplication, and division.

Why it is important: You need to perform basic operations to analyze data.

Simple explanation: Think of operations like doing math with your data.

  • Addition: +
  • Subtraction: -
  • Multiplication: *
  • Division: /
  • Mean: mean()
  • Sum: sum()
# Basic operations a <- 10 b <- 5 sum_ab <- a + b diff_ab <- a - b prod_ab <- a * b quot_ab <- a / b

🏫 School example: A teacher calculates the sum and mean of student grades.

🏠 Home example: Your parents calculate the sum of expenses.

πŸ‡³πŸ‡¬ Nigerian example: A business calculates the mean of sales figures.

πŸ“Œ Mini summary: Basic operations help you analyze data.

9. R in Nigeria

Definition: R is used by many Nigerian businesses and organizations to understand their data.

Why it is important: Seeing local examples helps you understand how R is used in your country.

Simple explanation: Think of it like seeing your favourite food at a local restaurant. It makes you feel connected.

  • Banks: Nigerian banks use R to analyze customer data and detect fraud.
  • Telecom: Companies like MTN use R to analyze network performance.
  • Government: The government uses R to track health and education data.
  • Retail: Shops use R to track sales and inventory.
  • Startups: Nigerian startups use R to understand their customers.

πŸ‡³πŸ‡¬ Nigerian example: A Lagos-based supermarket uses R to see which products are selling best.

πŸ“Œ Mini summary: R is helping Nigerian businesses grow and make better decisions.

10. Getting Help

Definition: Getting help means finding information about functions and packages.

Why it is important: You will need help as you learn R.

Simple explanation: Think of getting help like asking a teacher for help.

  • Help function: ?function_name
  • Example: ?mean
  • Online resources: Stack Overflow, R Documentation, and R-bloggers.
  • Community: Join R communities to ask questions.
# Getting help ?mean ?sum ?plot

🏫 School example: A teacher uses ?mean to understand the mean function.

🏠 Home example: Your parents use online resources to learn R.

πŸ‡³πŸ‡¬ Nigerian example: A business analyst uses Stack Overflow to solve problems.

πŸ“Œ Mini summary: Getting help is an important part of learning R.

πŸ“– Key Vocabulary

Word Simple Meaning
R A tool that turns data into pictures.
RStudio The tool used to write R code.
Data Information, like numbers and words.
Visualization A chart or graph that shows data in a picture.
Package A collection of R functions.
Function A command that performs a task.
Vector A collection of values.
Data Frame A table of data.
Variable A container for a value.
Import Load data into R.
Console Where you type commands.
Editor Where you write scripts.
Environment Shows your data and variables.
Plot A chart or graph.
Syntax The rules for writing code.

🧩 Important Concepts

  • R turns data into pictures.
  • RStudio is the tool used to write R code.
  • Packages are collections of functions.
  • Vectors are collections of values.
  • Importing data is the first step in analysis.
  • Basic operations help you analyze data.
  • R is used by Nigerian businesses.
  • Getting help is an important part of learning R.

πŸ“Œ Step-by-Step Explanations

How to install R and RStudio

  1. Go to the R website: Visit cran.r-project.org.
  2. Download R: Choose your operating system and download R.
  3. Install R: Follow the installation instructions.
  4. Go to the RStudio website: Visit rstudio.com.
  5. Download RStudio: Choose the free version and download.
  6. Install RStudio: Follow the installation instructions.
  7. Open RStudio: Click the RStudio icon.

How to create a vector

  1. Open RStudio: Click the RStudio icon.
  2. Open the Console: The Console is on the left.
  3. Type the code: Type numbers <- c(1, 2, 3, 4, 5).
  4. Press Enter: The vector is created.
  5. View the vector: Type numbers and press Enter.
  6. Do operations: Type sum(numbers) and press Enter.

🌍 Real-life Examples

  • School: A teacher creates a report showing student performance.
  • Hospital: A hospital uses R to track patient data.
  • Restaurant: A restaurant uses R to track sales and inventory.
  • Shop: A shop uses R to see which products sell best.

πŸ‡³πŸ‡¬ Nigerian Examples

  • Paystack: Uses R to analyze payment data.
  • Flutterwave: Uses R to track transaction trends.
  • MTN Nigeria: Uses R to monitor network performance.
  • A Lagos supermarket: Uses R to track sales by product.
  • A Nigerian bank: Uses R to analyze branch performance.

🎈 Fun Examples for You

  • Your pocket money: You can use R to track how you spend your pocket money.
  • Your game scores: You can track your video game scores in R.
  • Your reading log: You can track how many books you read each month.
  • Your chores: You can track how many chores you complete each week.

🏠 Everyday Examples

  • At home: Your parents can use R to track family expenses.
  • At school: Your teacher can use R to track class performance.
  • In your community: Local businesses use R to track sales.
  • In your own life: You can use R to track your goals.

πŸ‘©β€πŸ« Teacher Notes

  • Encourage students to think about data they collect in their daily lives.
  • Use the warm-up story to spark curiosity about R.
  • Demonstrate RStudio by running simple code.
  • Discuss why R is important for businesses in Nigeria.
  • Ask students to think about what they would like to visualize.

πŸ‘ͺ Parent Tips

  • Talk to your child about how you use data in your work or daily life.
  • Show your child how you track expenses or other information.
  • Install R and RStudio and explore it together.
  • Encourage your child to think about what data they would like to visualize.
  • Share examples of Nigerian businesses using R.

🧠 Interesting Facts

  • R was first released in 1995.
  • R is used by over 1 million people worldwide.
  • R is free to download and use.
  • R can connect to over 100 different data sources.
  • R is updated every year.

πŸ’‘ Did You Know?

  • Did you know that R can connect to data from Nigerian banks?
  • Did you know that you can use R to create maps of Nigeria?
  • Did you know that R is free and open-source?
  • Did you know that many Nigerian startups use R to understand their customers?
  • Did you know that R can be used to track the Nigerian economy?

πŸ”” Remember This

  • R turns data into pictures.
  • RStudio is the tool for writing R code.
  • Packages are collections of functions.
  • Vectors are collections of values.
  • Importing data is the first step in analysis.
  • Basic operations help you analyze data.
  • Nigerian businesses use R to grow.
  • You can learn R too!

⚠️ Common Mistakes

  • Not installing R first: You need R before RStudio.
  • Forgetting library(): Packages must be loaded with library().
  • Not using <-: Use <- to assign variables.
  • Case sensitivity: R is case-sensitive.
  • Not saving your work: Always save your script.

✨ Best Practices for R

  • Always use <- to assign variables.
  • Use comments (#) to explain your code.
  • Save your scripts regularly.
  • Use library() to load packages.
  • Keep your code clean and organized.
  • Practice with real data.
  • Keep learning and exploring new packages.

πŸ“Š Clear Illustrations

1. What is R?

    +-------------------+
    |  YOUR DATA        |  ← Numbers, sales, records
    +-------------------+
           |
           V
    +-------------------+
    |  R                |  ← The magic tool
    +-------------------+
           |
           V
    +-------------------+
    |  BEAUTIFUL CHARTS |  ← Pictures that tell a story
    +-------------------+
    

2. RStudio Panes

    +---------------------------------------------------+
    |  EDITOR          |  CONSOLE                       |
    |  (Write code)    |  (Run code)                   |
    +---------------------------------------------------+
    |  ENVIRONMENT     |  PLOTS                        |
    |  (View data)     |  (View charts)               |
    +---------------------------------------------------+
    

3. Creating a Vector

    +-------------------+
    |  numbers <- c(1,2,3,4,5)  ← Create a vector
    +-------------------+
           |
           V
    +-------------------+
    |  sum(numbers)     ← Add them up
    +-------------------+
           |
           V
    +-------------------+
    |  mean(numbers)    ← Find the average
    +-------------------+
    

4. Comparison: R vs Excel

R Excel
Free and open-source Paid
Handles large data Limited to small data
Powerful visualizations Basic charts
Reproducible Not reproducible
Programming language Spreadsheet tool

πŸ“ Lesson Summaries

Lesson 1: R is a tool that turns data into pictures.

Lesson 2: R is important because it helps people understand data.

Lesson 3: R is the language; RStudio is the tool.

Lesson 4: RStudio has different panes for writing code and viewing data.

Lesson 5: R syntax is the set of rules for writing R code.

Lesson 6: Importing data is the first step in analysis.

Lesson 7: Vectors are collections of values.

Lesson 8: Basic operations help you analyze data.

Lesson 9: R is used by Nigerian businesses.

Lesson 10: Getting help is an important part of learning R.

πŸ“˜ End-of-Module Summary

In this module, you learned about R β€” a tool that turns data into beautiful pictures. You discovered the R ecosystem, RStudio, and how to create your first report. You also learned about the importance of R and how it is used in Nigeria.

🎯 You can now:

  • Explain what R is in simple words.
  • Understand why R is important for data analysis.
  • Install R and RStudio on your computer.
  • Navigate the RStudio interface.
  • Write and run basic R code.
  • Import data into R.
  • Create a simple vector and perform basic operations.
  • Give examples of how R is used in Nigeria.

❓ Frequently Asked Questions

1. What is R?
R is a tool that turns data into pictures.
2. Is R free?
Yes, R is free and open-source.
3. What is RStudio?
RStudio is the tool used to write R code.
4. What can I do with R?
You can create charts, graphs, and analyze data.
5. Do I need to know programming to use R?
No, you can learn R without prior programming experience.
6. What is a vector?
A vector is a collection of values.
7. What is a package?
A package is a collection of R functions.
8. Can I use R on my phone?
No, R is typically used on a computer.
9. Is R used in Nigeria?
Yes, many Nigerian businesses use R.
10. Can I learn R?
Yes, anyone can learn R!

πŸ“ Review Questions (15)

  1. What is R?
  2. Why is R important?
  3. What is RStudio?
  4. What are the different panes in RStudio?
  5. What is the assignment operator in R?
  6. How do you create a vector?
  7. What is a function?
  8. How do you import data into R?
  9. What is the mean function?
  10. What is the sum function?
  11. Give an example of a Nigerian business using R.
  12. What is the difference between R and Excel?
  13. What is a package in R?
  14. What is a data frame?
  15. What is the first step in data analysis?

✏️ Fill-in-the-Blank Exercises

  1. R is a tool that turns __________ into pictures.
  2. RStudio is the tool used to write __________ code.
  3. A __________ is a collection of values.
  4. __________ is the first step in data analysis.
  5. R is used by __________ businesses.

βœ… True or False Exercises

  1. R is only for adults. (False)
  2. R is free and open-source. (True)
  3. RStudio is the programming language. (False)
  4. A vector is a collection of values. (True)
  5. Nigerian businesses use R. (True)

πŸ”˜ Multiple Choice Questions

  1. What is R?
    A) A tool that turns data into pictures B) A game C) A type of food D) A sport
    Answer: A
  2. What is RStudio?
    A) The tool used to write R code B) A game C) A type of food D) A sport
    Answer: A
  3. What is a vector?
    A) A collection of values B) A chart C) A database D) A filter
    Answer: A
  4. What is the assignment operator in R?
    A) <- B) = C) == D) !=
    Answer: A
  5. What is the mean function?
    A) mean() B) sum() C) sd() D) median()
    Answer: A
  6. What is the sum function?
    A) sum() B) mean() C) sd() D) median()
    Answer: A
  7. How do you create a vector?
    A) c() B) v() C) vec() D) vector()
    Answer: A
  8. What is a package?
    A) A collection of functions B) A chart C) A database D) A filter
    Answer: A
  9. What is a data frame?
    A) A table of data B) A chart C) A database D) A filter
    Answer: A
  10. Which of these is a Nigerian business using R?
    A) Paystack B) Flutterwave C) MTN D) All of the above
    Answer: D
  11. What is the first step in data analysis?
    A) Import data B) Create a chart C) Save the report D) Share the report
    Answer: A
  12. What is the difference between R and Excel?
    A) R is free; Excel is paid B) They are the same C) R is paid; Excel is free D) R is for charts; Excel is for data
    Answer: A
  13. What is the Console in RStudio?
    A) Where you type commands B) Where you write scripts C) Where you view data D) Where you view charts
    Answer: A
  14. What is the Editor in RStudio?
    A) Where you write scripts B) Where you type commands C) Where you view data D) Where you view charts
    Answer: A
  15. What is the Environment in RStudio?
    A) Shows your data and variables B) Where you type commands C) Where you write scripts D) Where you view charts
    Answer: A

πŸ”— Matching Exercises

Match the word on the left with the correct meaning on the right:

Word Meaning
R A tool for writing R code
RStudio A collection of values
Vector A tool that turns data into pictures
Package A collection of functions

Answers: R β†’ A tool that turns data into pictures; RStudio β†’ A tool for writing R code; Vector β†’ A collection of values; Package β†’ A collection of functions.

πŸ“ Short Answer Questions

  1. Explain R in your own words.
  2. Why is R important?
  3. What is the difference between R and RStudio?
  4. How do you import data into R?
  5. Give an example of a Nigerian business using R.

🎭 Scenario-based Exercises

Scenario 1: Chidi has collected data on how many cars pass his house each day. He wants to turn this data into a chart so he can see patterns. What should Chidi do? How can R help him?

Scenario 2: A Lagos supermarket wants to see which products are selling best. They have a lot of sales data in Excel. How can R help them?

πŸ‘₯ Group Activity

In groups of 4–5, discuss a type of data you could collect (e.g., sales, attendance, expenses). Think about how you could use R to visualize this data. Present your ideas to the class.

πŸ§‘ Individual Activity

Think about data you collect in your daily life (e.g., how much time you spend on homework, how many books you read). Write a short paragraph (about 100 words) about how you could use R to visualize this data.

πŸ—£οΈ Classroom Discussion Questions

  1. Why do you think businesses need tools like R?
  2. What kind of data would you like to visualize?
  3. How can R help Nigerian businesses?
  4. What is the most interesting thing you learned about R?
  5. How do you think R will change in the future?

πŸ› οΈ Mini Project

Create a Simple R Report

Find a simple dataset (e.g., a list of sales or expenses). Use R to create a report with at least one chart. Save and share your report with the class.

πŸ“‹ Practical Assignment

Install R and RStudio on your computer. Import a simple CSV file and create a bar chart. Write a short report (about 150 words) about what you did and what you learned.

πŸ† Challenge Exercise

The Challenge: Imagine you are a business analyst in Lagos. You have sales data for four regions: Lagos, Abuja, Kano, and Port Harcourt. Create an R script that imports the data, creates a bar chart, and calculates the average sales per region.

πŸ” Quiz Answers

Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A

Fill-in-the-Blank: 1. data, 2. R, 3. vector, 4. Importing, 5. Nigerian

True or False: 1. False, 2. True, 3. False, 4. True, 5. True

Matching: R β†’ A tool that turns data into pictures; RStudio β†’ A tool for writing R code; Vector β†’ A collection of values; Package β†’ A collection of functions.

🎯 Key Takeaways

  • R is a tool that turns data into pictures.
  • RStudio is the tool for writing R code.
  • Packages are collections of functions.
  • Vectors are collections of values.
  • Importing data is the first step in analysis.
  • Basic operations help you analyze data.
  • Nigerian businesses use R to grow.
  • Anyone can learn R!

πŸ”œ Preparation for Module Two

In the next module, we will explore Data Import and Cleaning. You will learn how to clean and prepare your data for analysis.


πŸŽ‰ Congratulations! You have completed Module One of the R for Data Analysis course.

πŸ‘ You are now ready to move to Module Two: Data Import and Cleaning.

3

Module Two

Module Two: Data Import and Cleaning

πŸ“₯ Module Two: Data Import and Cleaning

Welcome back, young data cleaner! Today we learn how to import and clean our data.

🌟 Module Introduction

In Module One, you learned what R is and how to use RStudio. Now you are ready to work with real data. But real data is often messy. It has mistakes, missing values, and inconsistent formats.

Imagine you have a box of toys. Some toys are broken, some are missing pieces, and some are dirty. Before you can play with them, you need to clean them. Data is the same. You need to clean it before you can analyze it.

In this module, you will learn how to import data into R and how to clean it. You will learn about missing values, duplicates, and inconsistent formats. By the end, you will be able to turn messy data into clean, analysis-ready data.

πŸ’‘ Think about it: Have you ever had to clean a messy room? Data cleaning is like that β€” you organize and tidy up your data.

🎯 Learning Objectives

By the time you finish this module, you will be able to:

  • Explain what data cleaning is and why it is important.
  • Import data from CSV, Excel, and other sources.
  • View and explore your data.
  • Handle missing values.
  • Remove duplicates.
  • Fix inconsistent formats.
  • Use basic dplyr functions for data manipulation.
  • Give examples of data cleaning in Nigerian businesses.

πŸ“– Warm-up Story: Nneka's Messy Data

Nneka is a 12-year-old girl who loves to collect data. She asked her classmates about their favourite foods and wrote down their answers. But her notebook was messy. Some names were spelled wrong. Some answers were missing. Some were written in different ways.

Nneka tried to make a chart, but it was hard. She showed her notebook to her uncle, who works with data. Her uncle said, "Nneka, your data is messy. You need to clean it before you can make a chart."

He opened R and showed her how to import the data and clean it. Together, they fixed the misspelled names, filled in missing answers, and made everything consistent. Then they created a beautiful chart showing the favourite foods of the class.

Nneka learned that cleaning data is just as important as making charts. Without clean data, your charts can be wrong.

🧠 Think about it: Have you ever had data that was messy or hard to understand? What did you do?

πŸ“š Main Lessons

1. What is Data Cleaning?

Definition: Data cleaning is the process of fixing errors and making data consistent.

Why it is important: Clean data leads to accurate analysis. Dirty data leads to wrong conclusions.

Simple explanation: Think of data cleaning like washing your hands before eating. You need to remove the dirt so you don't get sick.

🏫 School example: A teacher cleans student data by fixing misspelled names.

🏠 Home example: Your parents clean expense data by removing duplicate entries.

πŸ‡³πŸ‡¬ Nigerian example: A business cleans sales data by fixing inconsistent region names.

  • Missing values: When data is missing (e.g., a blank cell).
  • Duplicates: When the same data appears more than once.
  • Inconsistent formats: When data is written in different ways (e.g., "Lagos" vs "lagos").
  • Errors: When data is wrong (e.g., a typo).
    +-------------------+
    |  MESSY DATA       |  ← Data with errors, blanks, and mistakes
    +-------------------+
           |
           V
    +-------------------+
    |  DATA CLEANING    |  ← The cleaning process
    +-------------------+
           |
           V
    +-------------------+
    |  CLEAN DATA       |  ← Ready for analysis
    +-------------------+
    

πŸ“Œ Mini summary: Data cleaning is the process of fixing errors and making data consistent.

2. Importing Data

Definition: Importing data means loading data from a file into R.

Why it is important: You need to import data before you can analyze it.

Simple explanation: Think of importing data like opening a book before you read it.

  • CSV: read.csv("filename.csv")
  • Excel: read_excel("filename.xlsx")
  • Text: read.table("filename.txt")
  • Web: read.csv("https://example.com/data.csv")
# Importing a CSV file sales_data <- read.csv("sales.csv") # Importing an Excel file library(readxl) sales_data <- read_excel("sales.xlsx") # View the data head(sales_data) View(sales_data)

🏫 School example: A teacher imports student grades from a CSV file.

🏠 Home example: Your parents import expense data from an Excel file.

πŸ‡³πŸ‡¬ Nigerian example: A business imports sales data from a CSV file.

πŸ“Œ Mini summary: Importing data is the first step in analysis.

3. Exploring Your Data

Definition: Exploring your data means looking at it to understand what you have.

Why it is important: You need to know what your data looks like before you can clean it.

Simple explanation: Think of exploring your data like looking at a map before a journey.

  • head(): Shows the first few rows.
  • tail(): Shows the last few rows.
  • summary(): Shows a summary of each column.
  • str(): Shows the structure of the data.
  • View(): Opens the data in a spreadsheet view.
# Exploring data head(sales_data) # First 6 rows tail(sales_data) # Last 6 rows summary(sales_data) # Summary statistics str(sales_data) # Structure of the data View(sales_data) # Open in a spreadsheet view

🏫 School example: A teacher uses head() to see the first few student grades.

🏠 Home example: Your parents use summary() to see expense statistics.

πŸ‡³πŸ‡¬ Nigerian example: A business uses head() to see the first few sales records.

πŸ“Œ Mini summary: Exploring your data helps you understand what you have.

4. Handling Missing Values

Definition: Missing values are blank cells in your data.

Why it is important: Missing values can cause errors in your analysis.

Simple explanation: Think of missing values like missing pieces in a puzzle.

  • is.na(): Checks for missing values.
  • na.omit(): Removes rows with missing values.
  • replace(): Replaces missing values with a value.
  • mutate(): Creates a new column with cleaned data.
# Checking for missing values sum(is.na(sales_data$Sales)) # Removing rows with missing values clean_data <- na.omit(sales_data) # Replacing missing values with 0 sales_data$Sales[is.na(sales_data$Sales)] <- 0 # Using dplyr to handle missing values library(dplyr) clean_data <- sales_data %>% filter(!is.na(Sales))

🏫 School example: A teacher replaces missing grades with 0.

🏠 Home example: Your parents remove rows with missing expense amounts.

πŸ‡³πŸ‡¬ Nigerian example: A business replaces missing sales with 0.

πŸ“Œ Mini summary: Handling missing values prevents errors in your analysis.

5. Removing Duplicates

Definition: Duplicates are rows that appear more than once.

Why it is important: Duplicates can make your analysis inaccurate.

Simple explanation: Think of duplicates like copying a page in a book.

  • duplicated(): Checks for duplicates.
  • distinct(): Removes duplicate rows.
# Checking for duplicates sum(duplicated(sales_data)) # Removing duplicates clean_data <- distinct(sales_data)

🏫 School example: A teacher removes duplicate student entries.

🏠 Home example: Your parents remove duplicate expense entries.

πŸ‡³πŸ‡¬ Nigerian example: A business removes duplicate customer records.

πŸ“Œ Mini summary: Removing duplicates keeps your data accurate.

6. Fixing Inconsistent Formats

Definition: Inconsistent formats are when data is written in different ways.

Why it is important: Inconsistent formats make it hard to analyze data.

Simple explanation: Think of inconsistent formats like different spellings of the same word.

  • tolower(): Converts text to lowercase.
  • toupper(): Converts text to uppercase.
  • str_trim(): Removes extra spaces.
  • mutate(): Creates a new column with cleaned data.
# Converting text to lowercase sales_data$Region <- tolower(sales_data$Region) # Removing extra spaces library(stringr) sales_data$Region <- str_trim(sales_data$Region) # Using dplyr to fix formats library(dplyr) clean_data <- sales_data %>% mutate(Region = tolower(Region), Region = str_trim(Region))

🏫 School example: A teacher fixes inconsistent student names.

🏠 Home example: Your parents fix inconsistent category names.

πŸ‡³πŸ‡¬ Nigerian example: A business fixes inconsistent region names.

πŸ“Œ Mini summary: Fixing inconsistent formats makes your data consistent.

7. Renaming Columns

Definition: Renaming columns means changing the names of columns.

Why it is important: Clear column names make your data easier to understand.

Simple explanation: Think of renaming columns like labeling boxes in a storage room.

  • names(): Shows the column names.
  • rename(): Renames columns.
# Checking column names names(sales_data) # Renaming columns library(dplyr) clean_data <- sales_data %>% rename(SalesAmount = Sales, CustomerName = Customer)

🏫 School example: A teacher renames "Grade" to "Score".

🏠 Home example: Your parents rename "Amount" to "Expense".

πŸ‡³πŸ‡¬ Nigerian example: A business renames "Sales" to "Revenue".

πŸ“Œ Mini summary: Renaming columns makes your data easier to understand.

8. Filtering Data

Definition: Filtering means keeping only the rows that meet certain conditions.

Why it is important: Filtering helps you focus on the data that matters.

Simple explanation: Think of filtering like using a sieve to separate what you want.

  • filter(): Keeps rows that meet conditions.
# Filtering data library(dplyr) lagos_data <- sales_data %>% filter(Region == "lagos") high_sales <- sales_data %>% filter(Sales > 1000)

🏫 School example: A teacher filters students by grade.

🏠 Home example: Your parents filter expenses by category.

πŸ‡³πŸ‡¬ Nigerian example: A business filters sales by region.

πŸ“Œ Mini summary: Filtering helps you focus on the data that matters.

9. Selecting Columns

Definition: Selecting means keeping only the columns you need.

Why it is important: Selecting helps you focus on the columns that matter.

Simple explanation: Think of selecting like choosing the right tools for a job.

  • select(): Keeps only the columns you specify.
# Selecting columns library(dplyr) selected_data <- sales_data %>% select(Region, Sales, Profit)

🏫 School example: A teacher selects only the grade column.

🏠 Home example: Your parents select only the expense column.

πŸ‡³πŸ‡¬ Nigerian example: A business selects only the sales column.

πŸ“Œ Mini summary: Selecting helps you focus on the columns that matter.

10. Data Cleaning in Nigerian Businesses

Definition: Nigerian businesses use data cleaning to prepare their data for analysis.

Why it is important: Data cleaning helps Nigerian businesses make better decisions.

Simple explanation: Think of it like cleaning a shop before customers come in.

  • Banks: Clean customer data by fixing errors.
  • Telecom: Clean customer records by removing duplicates.
  • Retail: Clean sales data by fixing inconsistent product names.
  • Government: Clean citizen data by handling missing values.
  • Startups: Clean user data by fixing inconsistent formats.

πŸ‡³πŸ‡¬ Nigerian example: A Lagos supermarket cleans sales data by fixing inconsistent product names.

πŸ“Œ Mini summary: Nigerian businesses use data cleaning to make better decisions.

πŸ“– Key Vocabulary

Word Simple Meaning
Data Cleaning Fixing errors and making data consistent.
Import Loading data into R.
Missing Value A blank cell in your data.
Duplicate A row that appears more than once.
Inconsistent Format Data written in different ways.
Filter Keeping only certain rows.
Select Keeping only certain columns.
NA A missing value in R.
head() Shows the first few rows.
summary() Shows a summary of the data.
str() Shows the structure of the data.
View() Opens the data in a spreadsheet view.
mutate() Creates a new column.
rename() Changes column names.
distinct() Removes duplicate rows.

🧩 Important Concepts

  • Data cleaning fixes errors and makes data consistent.
  • Importing data is the first step in analysis.
  • Missing values can cause errors.
  • Duplicates can make your analysis inaccurate.
  • Inconsistent formats make data hard to analyze.
  • Filtering helps you focus on the data that matters.
  • Selecting helps you focus on the columns that matter.
  • Nigerian businesses use data cleaning to make better decisions.

πŸ“Œ Step-by-Step Explanations

How to import and clean data

  1. Set working directory: setwd("path/to/your/folder").
  2. Import data: data <- read.csv("filename.csv").
  3. Explore data: head(data), summary(data), str(data).
  4. Handle missing values: na.omit(data) or replace with a value.
  5. Remove duplicates: distinct(data).
  6. Fix inconsistent formats: tolower(), str_trim().
  7. Rename columns: rename(data, new = old).
  8. Filter data: filter(data, condition).
  9. Select columns: select(data, column1, column2).
  10. Save clean data: write.csv(clean_data, "clean_data.csv").

🌍 Real-life Examples

  • School: A teacher cleans student data by fixing names and removing duplicates.
  • Hospital: A hospital cleans patient data by handling missing values.
  • Restaurant: A restaurant cleans sales data by fixing inconsistent product names.
  • Shop: A shop cleans inventory data by removing duplicates.

πŸ‡³πŸ‡¬ Nigerian Examples

  • Paystack: Cleans payment data by fixing errors.
  • Flutterwave: Cleans transaction data by removing duplicates.
  • MTN Nigeria: Cleans customer data by handling missing values.
  • A Lagos supermarket: Cleans sales data by fixing inconsistent product names.
  • A Nigerian bank: Cleans customer data by fixing inconsistent formats.

🎈 Fun Examples for You

  • Your toy collection: Clean your toy list by removing duplicate toys.
  • Your game scores: Clean your game scores by removing errors.
  • Your reading log: Clean your reading log by fixing inconsistent book titles.
  • Your chores: Clean your chore list by removing duplicates.

🏠 Everyday Examples

  • At home: Your parents clean expense data by removing duplicate entries.
  • At school: Your teacher cleans student data by fixing errors.
  • In your community: Local businesses clean sales data by fixing inconsistent product names.
  • In your own life: You can clean your personal data by removing duplicates.

πŸ‘©β€πŸ« Teacher Notes

  • Encourage students to think about messy data they have encountered.
  • Use the warm-up story to spark interest in data cleaning.
  • Demonstrate importing and cleaning data in R.
  • Discuss why data cleaning is important for accurate analysis.
  • Ask students to think about how they would clean different types of data.

πŸ‘ͺ Parent Tips

  • Talk to your child about how you clean data in your work or daily life.
  • Show your child how you clean expense or other data.
  • Open R and explore data cleaning together.
  • Encourage your child to think about how they would clean different types of data.
  • Share examples of Nigerian businesses using data cleaning.

🧠 Interesting Facts

  • Data cleaning can take up to 80% of the time in a data project.
  • Missing values are one of the most common data problems.
  • Duplicates are often caused by human error.
  • Inconsistent formats are common in real-world data.
  • Data cleaning is a skill that is in high demand.

πŸ’‘ Did You Know?

  • Did you know that data cleaning is also called "data munging"?
  • Did you know that many Nigerian businesses have dedicated data cleaning teams?
  • Did you know that data cleaning can improve the accuracy of your analysis by up to 50%?
  • Did you know that R has many packages for data cleaning?
  • Did you know that data cleaning is an essential skill for data analysts?

πŸ”” Remember This

  • Data cleaning fixes errors and makes data consistent.
  • Importing data is the first step in analysis.
  • Missing values can cause errors.
  • Duplicates can make your analysis inaccurate.
  • Inconsistent formats make data hard to analyze.
  • Filtering helps you focus on the data that matters.
  • Selecting helps you focus on the columns that matter.
  • Nigerian businesses use data cleaning to make better decisions.

⚠️ Common Mistakes

  • Not cleaning data: Using dirty data leads to wrong results.
  • Removing the wrong rows: Check before you remove.
  • Not checking for missing values: Missing values can cause errors.
  • Not removing duplicates: Duplicates can skew your analysis.
  • Not fixing inconsistent formats: Inconsistent formats make data hard to analyze.

✨ Best Practices for Data Cleaning

  • Always explore your data before cleaning.
  • Check for missing values and handle them.
  • Remove duplicates to keep your data accurate.
  • Fix inconsistent formats to make your data consistent.
  • Rename columns to make your data understandable.
  • Filter and select to focus on what matters.
  • Save your clean data for future use.
  • Document your cleaning steps.

πŸ“Š Clear Illustrations

1. The Data Cleaning Process

    +-------------------+
    |  MESSY DATA       |  ← Data with errors, blanks, and mistakes
    +-------------------+
           |
           V
    +-------------------+
    |  IMPORT DATA      |  ← Load data into R
    +-------------------+
           |
           V
    +-------------------+
    |  EXPLORE DATA     |  ← Look at the data
    +-------------------+
           |
           V
    +-------------------+
    |  HANDLE MISSING   |  ← Fix missing values
    +-------------------+
           |
           V
    +-------------------+
    |  REMOVE DUPLICATES |  ← Remove duplicates
    +-------------------+
           |
           V
    +-------------------+
    |  FIX INCONSISTENT |  ← Fix inconsistent formats
    +-------------------+
           |
           V
    +-------------------+
    |  CLEAN DATA       |  ← Ready for analysis
    +-------------------+
    

2. Filtering Data

    +-------------------+
    |  ALL DATA         |  ← All rows
    +-------------------+
           |
           V
    +-------------------+
    |  FILTER           |  ← Keep only rows that meet conditions
    +-------------------+
           |
           V
    +-------------------+
    |  FILTERED DATA    |  ← Only the rows you want
    +-------------------+
    

3. Selecting Data

    +-------------------+
    |  ALL COLUMNS      |  ← All columns
    +-------------------+
           |
           V
    +-------------------+
    |  SELECT           |  ← Keep only certain columns
    +-------------------+
           |
           V
    +-------------------+
    |  SELECTED DATA    |  ← Only the columns you want
    +-------------------+
    

4. Comparison: Dirty vs Clean Data

Dirty Data Clean Data
Has missing values No missing values
Has duplicates No duplicates
Inconsistent formats Consistent formats
Unclear column names Clear column names
Hard to analyze Easy to analyze

πŸ“ Lesson Summaries

Lesson 1: Data cleaning fixes errors and makes data consistent.

Lesson 2: Importing data is the first step in analysis.

Lesson 3: Exploring your data helps you understand what you have.

Lesson 4: Handling missing values prevents errors.

Lesson 5: Removing duplicates keeps your data accurate.

Lesson 6: Fixing inconsistent formats makes your data consistent.

Lesson 7: Renaming columns makes your data easier to understand.

Lesson 8: Filtering helps you focus on the data that matters.

Lesson 9: Selecting helps you focus on the columns that matter.

Lesson 10: Nigerian businesses use data cleaning to make better decisions.

πŸ“˜ End-of-Module Summary

In this module, you learned about data import and cleaning. You discovered how to import data from different sources and explore it. You learned how to handle missing values, remove duplicates, and fix inconsistent formats. You also learned how to rename columns, filter data, and select columns.

🎯 You can now:

  • Explain what data cleaning is and why it is important.
  • Import data from CSV, Excel, and other sources.
  • View and explore your data.
  • Handle missing values.
  • Remove duplicates.
  • Fix inconsistent formats.
  • Use basic dplyr functions for data manipulation.
  • Give examples of data cleaning in Nigerian businesses.

❓ Frequently Asked Questions

1. What is data cleaning?
Data cleaning is the process of fixing errors and making data consistent.
2. Why is data cleaning important?
Data cleaning leads to accurate analysis.
3. How do I import data into R?
Use read.csv() for CSV files or read_excel() for Excel files.
4. What is a missing value?
A missing value is a blank cell in your data.
5. How do I handle missing values?
Use na.omit() to remove rows or replace with a value.
6. What is a duplicate?
A duplicate is a row that appears more than once.
7. How do I remove duplicates?
Use distinct() to remove duplicate rows.
8. What is an inconsistent format?
Data written in different ways.
9. How do I fix inconsistent formats?
Use tolower() and str_trim() to standardize.
10. How do Nigerian businesses use data cleaning?
They use it to prepare data for analysis and make better decisions.

πŸ“ Review Questions (15)

  1. What is data cleaning?
  2. Why is data cleaning important?
  3. How do you import a CSV file into R?
  4. What function shows the first few rows of data?
  5. What is a missing value?
  6. How do you handle missing values?
  7. What is a duplicate?
  8. How do you remove duplicates?
  9. What is an inconsistent format?
  10. How do you fix inconsistent formats?
  11. How do you rename a column?
  12. How do you filter data?
  13. How do you select columns?
  14. Give an example of a Nigerian business using data cleaning.
  15. What is the best practice for data cleaning?

✏️ Fill-in-the-Blank Exercises

  1. Data __________ fixes errors and makes data consistent.
  2. __________ data is the first step in analysis.
  3. A __________ value is a blank cell in your data.
  4. __________ removes duplicate rows.
  5. __________ helps you focus on the data that matters.

βœ… True or False Exercises

  1. Data cleaning is not important. (False)
  2. Importing data is the first step in analysis. (True)
  3. Missing values can cause errors. (True)
  4. Duplicates are not a problem. (False)
  5. Filtering helps you focus on the data that matters. (True)

πŸ”˜ Multiple Choice Questions

  1. What is data cleaning?
    A) Fixing errors and making data consistent B) Deleting data C) Adding data D) Moving data
    Answer: A
  2. How do you import a CSV file?
    A) read.csv() B) read_excel() C) read.table() D) read.txt()
    Answer: A
  3. What is a missing value?
    A) A blank cell B) A duplicate C) A number D) A date
    Answer: A
  4. How do you remove duplicates?
    A) distinct() B) duplicate() C) remove() D) delete()
    Answer: A
  5. What is an inconsistent format?
    A) Data written in different ways B) Missing data C) Duplicate data D) Correct data
    Answer: A
  6. How do you fix inconsistent formats?
    A) tolower() B) toupper() C) str_trim() D) All of the above
    Answer: D
  7. How do you rename a column?
    A) rename() B) change() C) update() D) modify()
    Answer: A
  8. How do you filter data?
    A) filter() B) select() C) mutate() D) rename()
    Answer: A
  9. How do you select columns?
    A) select() B) filter() C) mutate() D) rename()
    Answer: A
  10. Which of these is a Nigerian business using data cleaning?
    A) Paystack B) Flutterwave C) MTN D) All of the above
    Answer: D
  11. What does head() do?
    A) Shows the first few rows B) Shows the last few rows C) Shows a summary D) Shows the structure
    Answer: A
  12. What does summary() do?
    A) Shows a summary B) Shows the first few rows C) Shows the last few rows D) Shows the structure
    Answer: A
  13. What does str() do?
    A) Shows the structure B) Shows the first few rows C) Shows the last few rows D) Shows a summary
    Answer: A
  14. What is the best practice for data cleaning?
    A) Explore your data first B) Delete everything C) Ignore missing values D) Keep duplicates
    Answer: A
  15. What is NA in R?
    A) A missing value B) A number C) A date D) A string
    Answer: A

πŸ”— Matching Exercises

Match the word on the left with the correct meaning on the right:

Word Meaning
read.csv() Removes duplicates
distinct() Shows the first few rows
head() Imports a CSV file
filter() Keeps only certain rows

Answers: read.csv() β†’ Imports a CSV file; distinct() β†’ Removes duplicates; head() β†’ Shows the first few rows; filter() β†’ Keeps only certain rows.

πŸ“ Short Answer Questions

  1. Explain data cleaning in your own words.
  2. Why is data cleaning important?
  3. How do you import a CSV file into R?
  4. How do you handle missing values?
  5. Give an example of a Nigerian business using data cleaning.

🎭 Scenario-based Exercises

Scenario 1: Nneka has a list of students with their grades. Some students are listed twice, and some grades are missing. What should Nneka do? How can R help her?

Scenario 2: A Lagos supermarket has sales data in different formats. Some product names are written in uppercase, some in lowercase. Some have extra spaces. How can R help clean this data?

πŸ‘₯ Group Activity

In groups of 4–5, discuss a type of messy data you have encountered. Think about how you could use R to clean it. Present your ideas to the class.

πŸ§‘ Individual Activity

Think about a dataset you could clean. Write a short paragraph (about 100 words) about how you would use R to clean it.

πŸ—£οΈ Classroom Discussion Questions

  1. Why is data cleaning important?
  2. What kind of messy data have you seen?
  3. How can R help Nigerian businesses clean their data?
  4. What is the most interesting thing you learned about data cleaning?
  5. How do you think data cleaning will change in the future?

πŸ› οΈ Mini Project

Clean a Dataset

Find a messy dataset (e.g., a list of sales with duplicates and missing values). Use R to clean it. Document the steps you took and the final result.

πŸ“‹ Practical Assignment

Download a sample dataset with errors. Use R to clean it by handling missing values, removing duplicates, and fixing inconsistent formats. Write a short report (about 150 words) about what you did and what you learned.

πŸ† Challenge Exercise

The Challenge: Imagine you are a business analyst in Lagos. You have sales data from four different stores, each in a different format. Use R to clean and combine the data. Create a single dataset with consistent formatting.

πŸ” Quiz Answers

Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-D, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A

Fill-in-the-Blank: 1. cleaning, 2. Importing, 3. missing, 4. distinct(), 5. Filtering

True or False: 1. False, 2. True, 3. True, 4. False, 5. True

Matching: read.csv() β†’ Imports a CSV file; distinct() β†’ Removes duplicates; head() β†’ Shows the first few rows; filter() β†’ Keeps only certain rows.

🎯 Key Takeaways

  • Data cleaning fixes errors and makes data consistent.
  • Importing data is the first step in analysis.
  • Missing values can cause errors.
  • Duplicates can make your analysis inaccurate.
  • Inconsistent formats make data hard to analyze.
  • Filtering helps you focus on the data that matters.
  • Selecting helps you focus on the columns that matter.
  • Nigerian businesses use data cleaning to make better decisions.

πŸ”œ Preparation for Module Three

In the next module, we will explore Data Visualization with ggplot2. You will learn how to create beautiful charts and graphs to visualize your data.


πŸŽ‰ Congratulations! You have completed Module Two of the R for Data Analysis course.

πŸ‘ You are now ready to move to Module Three: Data Visualization with ggplot2.

4

Module Three

Module Three: Data Visualization with ggplot2

πŸ“Š Module Three: Data Visualization with ggplot2

Welcome back, young artist! Today we learn how to turn data into beautiful pictures.

🌟 Module Introduction

In Module Two, you learned how to clean your data. Now you have clean, ready-to-use data. But data is hard to understand when it is just numbers. You need pictures to tell the story.

Imagine you have a story to tell. You could just read it out loud, but it would be much more interesting with pictures. Data visualizations are the pictures for your data.

ggplot2 is a package in R that helps you create beautiful charts and graphs. It is like a paintbrush for your data.

In this module, you will learn about the grammar of graphics, how to create bar charts, scatter plots, line charts, and more. You will also learn how to customize your charts to make them look professional.

πŸ’‘ Think about it: Have you ever seen a chart that told you a story? That is what data visualization does.

🎯 Learning Objectives

By the time you finish this module, you will be able to:

  • Explain what data visualization is and why it is important.
  • Understand the grammar of graphics.
  • Create bar charts to compare categories.
  • Create scatter plots to show relationships.
  • Create line charts to show trends over time.
  • Customize charts with titles, colours, and themes.
  • Give examples of data visualization in Nigerian businesses.

πŸ“– Warm-up Story: Ada's Art Gallery

Ada is a 12-year-old girl who loves art. She had a lot of data about her class's favourite foods. She wanted to show it to her teacher, but numbers were boring.

Her uncle, who works with data, said, "Ada, you need to visualize your data. Turn it into a picture. That way, everyone can understand it."

Ada learned how to use ggplot2 in R. She created a beautiful bar chart showing the favourite foods of her class. Her teacher was amazed. The chart told a story that numbers alone could not.

🧠 Think about it: Have you ever turned numbers into a picture? That is what data visualization is all about.

πŸ“š Main Lessons

1. What is Data Visualization?

Definition: Data visualization is the process of turning data into pictures like charts and graphs.

Why it is important: Pictures make it easier to understand data. A picture is worth a thousand numbers.

Simple explanation: Think of data visualization like drawing a picture of your data.

🏫 School example: A teacher uses a bar chart to show student grades.

🏠 Home example: Your parents use a pie chart to show expenses by category.

πŸ‡³πŸ‡¬ Nigerian example: A business uses a line chart to show sales over time.

  • Bar chart: Compares categories.
  • Scatter plot: Shows relationships.
  • Line chart: Shows trends over time.
  • Pie chart: Shows proportions.
    +-------------------+
    |  DATA             |  ← Numbers
    +-------------------+
           |
           V
    +-------------------+
    |  VISUALIZATION    |  ← Pictures
    +-------------------+
           |
           V
    +-------------------+
    |  UNDERSTANDING    |  ← Insights
    +-------------------+
    

πŸ“Œ Mini summary: Data visualization turns numbers into pictures.

2. What is ggplot2?

Definition: ggplot2 is a package in R that helps you create beautiful charts.

Why it is important: ggplot2 is the most popular package for data visualization in R.

Simple explanation: Think of ggplot2 like a paintbrush for your data.

  • Grammar of Graphics: A system for building charts layer by layer.
  • Components: Data, aesthetics, geometries, facets, themes.
  • Syntax: ggplot(data) + geom_function() + options.
# Basic ggplot2 syntax library(ggplot2) ggplot(data = my_data) + geom_bar(mapping = aes(x = category))

🏫 School example: A teacher uses ggplot2 to create a bar chart of grades.

🏠 Home example: Your parents use ggplot2 to create a pie chart of expenses.

πŸ‡³πŸ‡¬ Nigerian example: A business uses ggplot2 to create a line chart of sales.

πŸ“Œ Mini summary: ggplot2 is a package for creating charts in R.

3. Installing and Loading ggplot2

Definition: Installing means downloading the package. Loading means making it available to use.

Why it is important: You need to install and load ggplot2 before you can use it.

Simple explanation: Think of installing like buying a new tool. Loading is like taking it out of the toolbox.

  • Install: install.packages("ggplot2").
  • Load: library(ggplot2).
# Installing ggplot2 install.packages("ggplot2") # Loading ggplot2 library(ggplot2)

🏫 School example: A teacher installs ggplot2 on the school computer.

🏠 Home example: Your parents install ggplot2 on their computer.

πŸ‡³πŸ‡¬ Nigerian example: A business installs ggplot2 on the company computer.

πŸ“Œ Mini summary: You need to install and load ggplot2 before using it.

4. The Grammar of Graphics

Definition: The grammar of graphics is a system for building charts layer by layer.

Why it is important: It makes it easy to build complex charts.

Simple explanation: Think of it like building a house β€” you start with the foundation and add layers.

  • Data: The data you want to visualize.
  • Aesthetics: What you map to the axes (e.g., x, y, colour).
  • Geometries: The type of chart (e.g., bar, point, line).
  • Facets: Multiple charts side by side.
  • Themes: Styling and appearance.
# Grammar of graphics ggplot(data = my_data) + geom_bar(mapping = aes(x = category, fill = group)) + facet_wrap(~region) + theme_minimal()

🏫 School example: A teacher uses facets to show grades by class.

🏠 Home example: Your parents use themes to make charts look professional.

πŸ‡³πŸ‡¬ Nigerian example: A business uses facets to show sales by region.

πŸ“Œ Mini summary: The grammar of graphics builds charts layer by layer.

5. Creating a Bar Chart

Definition: A bar chart uses bars to compare categories.

Why it is important: Bar charts are the most common type of chart.

Simple explanation: Think of a bar chart like blocks of different heights.

  • geom_bar(): Creates a bar chart.
  • aes(x): The category on the x-axis.
  • aes(y): The value on the y-axis (optional).
# Bar chart library(ggplot2) ggplot(data = sales_data) + geom_bar(mapping = aes(x = Region))

🏫 School example: A teacher creates a bar chart of student grades.

🏠 Home example: Your parents create a bar chart of expenses by category.

πŸ‡³πŸ‡¬ Nigerian example: A business creates a bar chart of sales by region.

    +-------------------+
    |  SALES BY REGION  |
    +-------------------+
    |  Lagos     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
    |  Abuja     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
    |  Kano      β–ˆβ–ˆβ–ˆβ–ˆ
    |  Port H    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
    +-------------------+
    

πŸ“Œ Mini summary: Bar charts compare categories.

6. Creating a Scatter Plot

Definition: A scatter plot uses points to show relationships between two variables.

Why it is important: Scatter plots help you see patterns and correlations.

Simple explanation: Think of a scatter plot like dots on a map.

  • geom_point(): Creates a scatter plot.
  • aes(x): The variable on the x-axis.
  • aes(y): The variable on the y-axis.
# Scatter plot library(ggplot2) ggplot(data = sales_data) + geom_point(mapping = aes(x = Sales, y = Profit))

🏫 School example: A teacher creates a scatter plot of study time vs grades.

🏠 Home example: Your parents create a scatter plot of income vs expenses.

πŸ‡³πŸ‡¬ Nigerian example: A business creates a scatter plot of advertising spend vs sales.

    +-------------------+
    |  PROFIT VS SALES  |
    +-------------------+
    |  *    *           |
    |  *  *   *         |
    |  *   *    *       |
    |  *     *   *      |
    |  *  *     *       |
    +-------------------+
    

πŸ“Œ Mini summary: Scatter plots show relationships between variables.

7. Creating a Line Chart

Definition: A line chart uses lines to show trends over time.

Why it is important: Line charts are the best way to show how things change over time.

Simple explanation: Think of a line chart like connecting the dots.

  • geom_line(): Creates a line chart.
  • aes(x): The time variable on the x-axis.
  • aes(y): The value on the y-axis.
# Line chart library(ggplot2) ggplot(data = sales_data) + geom_line(mapping = aes(x = Date, y = Sales))

🏫 School example: A teacher creates a line chart of grades over time.

🏠 Home example: Your parents create a line chart of expenses over months.

πŸ‡³πŸ‡¬ Nigerian example: A business creates a line chart of sales over quarters.

    +-------------------+
    |  SALES OVER TIME  |
    +-------------------+
    |  200|    *--*--*  |
    |  150|  *      *   |
    |  100|*        *   |
    |   50|          *  |
    |    0|___________  |
    +-------------------+
    

πŸ“Œ Mini summary: Line charts show trends over time.

8. Customizing Your Charts

Definition: Customizing means changing the look of your chart.

Why it is important: Customization makes your charts look professional.

Simple explanation: Think of customization like decorating your chart.

  • labs(): Adds titles and labels.
  • theme(): Changes the appearance.
  • scale_color_manual(): Changes colours.
# Customizing a chart library(ggplot2) ggplot(data = sales_data) + geom_bar(mapping = aes(x = Region, fill = Region)) + labs(title = "Sales by Region", x = "Region", y = "Sales") + theme_minimal() + scale_fill_manual(values = c("blue", "green", "red", "orange"))

🏫 School example: A teacher adds a title to a chart.

🏠 Home example: Your parents change the colours of a chart.

πŸ‡³πŸ‡¬ Nigerian example: A business uses a professional theme for a report.

πŸ“Œ Mini summary: Customizing makes your charts look professional.

9. Faceting β€” Multiple Charts

Definition: Faceting creates multiple charts side by side.

Why it is important: Faceting helps you compare different groups.

Simple explanation: Think of faceting like multiple windows into your data.

  • facet_wrap(): Creates a grid of charts.
  • facet_grid(): Creates a grid of charts.
# Faceting library(ggplot2) ggplot(data = sales_data) + geom_bar(mapping = aes(x = Product)) + facet_wrap(~Region)

🏫 School example: A teacher creates charts for each class.

🏠 Home example: Your parents create charts for each month.

πŸ‡³πŸ‡¬ Nigerian example: A business creates charts for each region.

πŸ“Œ Mini summary: Faceting creates multiple charts.

10. Data Visualization in Nigerian Businesses

Definition: Nigerian businesses use data visualization to make decisions.

Why it is important: Visualizations help Nigerian businesses understand their data.

Simple explanation: Think of it like seeing a map before a journey.

  • Banks: Use charts to show branch performance.
  • Telecom: Use charts to show network performance.
  • Retail: Use charts to show sales by product.
  • Government: Use charts to show budget allocation.
  • Startups: Use charts to show growth.

πŸ‡³πŸ‡¬ Nigerian example: A Lagos supermarket uses a bar chart to show sales by product.

πŸ“Œ Mini summary: Nigerian businesses use data visualization to make decisions.

πŸ“– Key Vocabulary

Word Simple Meaning
Visualization A picture of your data.
ggplot2 A package for creating charts.
Bar Chart A chart with bars.
Scatter Plot A chart with points.
Line Chart A chart with lines.
Grammar of Graphics A system for building charts.
Aesthetics What you map to axes.
Geometries The type of chart.
Facets Multiple charts.
Themes Styling and appearance.
geom_bar() Creates a bar chart.
geom_point() Creates a scatter plot.
geom_line() Creates a line chart.
labs() Adds titles and labels.
facet_wrap() Creates multiple charts.

🧩 Important Concepts

  • Data visualization turns numbers into pictures.
  • ggplot2 is a package for creating charts.
  • The grammar of graphics builds charts layer by layer.
  • Bar charts compare categories.
  • Scatter plots show relationships.
  • Line charts show trends over time.
  • Customizing makes your charts look professional.
  • Faceting creates multiple charts.
  • Nigerian businesses use data visualization.

πŸ“Œ Step-by-Step Explanations

How to create a bar chart

  1. Load ggplot2: library(ggplot2).
  2. Prepare your data: Make sure your data is clean.
  3. Create the chart: ggplot(data = my_data) + geom_bar(mapping = aes(x = category)).
  4. Customize: Add titles, colours, and themes.
  5. Display: The chart appears in the Plots pane.

How to create a scatter plot

  1. Load ggplot2: library(ggplot2).
  2. Prepare your data: Make sure your data is clean.
  3. Create the chart: ggplot(data = my_data) + geom_point(mapping = aes(x = x_var, y = y_var)).
  4. Customize: Add titles, colours, and themes.
  5. Display: The chart appears in the Plots pane.

🌍 Real-life Examples

  • School: A teacher creates a bar chart of student grades.
  • Hospital: A hospital creates a line chart of patient admissions.
  • Restaurant: A restaurant creates a scatter plot of sales vs customer satisfaction.
  • Shop: A shop creates a bar chart of sales by product.

πŸ‡³πŸ‡¬ Nigerian Examples

  • Paystack: Uses bar charts to show transaction volumes.
  • Flutterwave: Uses line charts to show payment trends.
  • MTN Nigeria: Uses scatter plots to show network performance.
  • A Lagos supermarket: Uses bar charts to show sales by product.
  • A Nigerian bank: Uses line charts to show branch performance.

🎈 Fun Examples for You

  • Your pocket money: Use a bar chart to show your spending.
  • Your game scores: Use a line chart to show your progress.
  • Your reading log: Use a bar chart to show books read.
  • Your chores: Use a bar chart to show completion rates.

🏠 Everyday Examples

  • At home: Your parents use charts to track expenses.
  • At school: Your teacher uses charts to show grades.
  • In your community: Local businesses use charts to track sales.
  • In your own life: You can use charts to track your goals.

πŸ‘©β€πŸ« Teacher Notes

  • Encourage students to think about charts they see in everyday life.
  • Use the warm-up story to spark interest in data visualization.
  • Demonstrate creating charts in R.
  • Discuss the importance of choosing the right chart.
  • Ask students to think about how charts are used in Nigerian businesses.

πŸ‘ͺ Parent Tips

  • Talk to your child about charts you see in the news or at work.
  • Show your child how you create charts in Excel or other tools.
  • Open R and explore creating charts together.
  • Encourage your child to create their own charts.
  • Share examples of charts from Nigerian businesses.

🧠 Interesting Facts

  • The first bar chart was created in 1786.
  • Pie charts were invented in 1801.
  • Line charts are the oldest type of chart.
  • ggplot2 was created in 2005.
  • Data visualization is used in almost every industry.

πŸ’‘ Did You Know?

  • Did you know that charts help you understand data faster?
  • Did you know that you can create maps of Nigeria in R?
  • Did you know that ggplot2 has over 50 different geometries?
  • Did you know that Nigerian businesses use R for data visualization?
  • Did you know that the grammar of graphics was invented by Leland Wilkinson?

πŸ”” Remember This

  • Data visualization turns numbers into pictures.
  • ggplot2 is a package for creating charts.
  • Bar charts compare categories.
  • Scatter plots show relationships.
  • Line charts show trends over time.
  • Customizing makes your charts look professional.
  • Faceting creates multiple charts.
  • Nigerian businesses use data visualization.

⚠️ Common Mistakes

  • Using the wrong chart: Choose the right chart for your data.
  • Not loading ggplot2: Always load ggplot2 with library(ggplot2).
  • Forgetting to map aesthetics: Use aes() to map variables.
  • Not adding titles: Titles make charts understandable.
  • Using too many colours: Use colours wisely.

✨ Best Practices for Data Visualization

  • Choose the right chart for your data.
  • Use clear titles and labels.
  • Use colours wisely.
  • Keep it simple.
  • Use themes for a professional look.
  • Add annotations to highlight important points.
  • Test your charts with others.

πŸ“Š Clear Illustrations

1. Bar Chart

    +-------------------+
    |  SALES BY REGION  |
    +-------------------+
    |  Lagos     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
    |  Abuja     β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
    |  Kano      β–ˆβ–ˆβ–ˆβ–ˆ
    |  Port H    β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ
    +-------------------+
    

2. Scatter Plot

    +-------------------+
    |  PROFIT VS SALES  |
    +-------------------+
    |  *    *           |
    |  *  *   *         |
    |  *   *    *       |
    |  *     *   *      |
    |  *  *     *       |
    +-------------------+
    

3. Line Chart

    +-------------------+
    |  SALES OVER TIME  |
    +-------------------+
    |  200|    *--*--*  |
    |  150|  *      *   |
    |  100|*        *   |
    |   50|          *  |
    |    0|___________  |
    +-------------------+
    

4. Comparison: Chart Types

Chart Type Best For Example
Bar Chart Comparing categories Sales by region
Scatter Plot Showing relationships Sales vs profit
Line Chart Showing trends over time Sales over time
Pie Chart Showing proportions Market share

πŸ“ Lesson Summaries

Lesson 1: Data visualization turns numbers into pictures.

Lesson 2: ggplot2 is a package for creating charts.

Lesson 3: You need to install and load ggplot2 before using it.

Lesson 4: The grammar of graphics builds charts layer by layer.

Lesson 5: Bar charts compare categories.

Lesson 6: Scatter plots show relationships.

Lesson 7: Line charts show trends over time.

Lesson 8: Customizing makes your charts look professional.

Lesson 9: Faceting creates multiple charts.

Lesson 10: Nigerian businesses use data visualization.

πŸ“˜ End-of-Module Summary

In this module, you learned about data visualization with ggplot2. You discovered the grammar of graphics and how to create bar charts, scatter plots, and line charts. You also learned how to customize your charts and use faceting to create multiple charts. You saw examples from Nigeria and learned best practices for data visualization.

🎯 You can now:

  • Explain what data visualization is and why it is important.
  • Understand the grammar of graphics.
  • Create bar charts to compare categories.
  • Create scatter plots to show relationships.
  • Create line charts to show trends over time.
  • Customize charts with titles, colours, and themes.
  • Give examples of data visualization in Nigerian businesses.

❓ Frequently Asked Questions

1. What is data visualization?
Data visualization is turning numbers into pictures.
2. What is ggplot2?
ggplot2 is a package for creating charts in R.
3. What is the grammar of graphics?
A system for building charts layer by layer.
4. What is a bar chart?
A chart with bars that compares categories.
5. What is a scatter plot?
A chart with points that shows relationships.
6. What is a line chart?
A chart with lines that shows trends over time.
7. How do you customize a chart?
Use labs(), theme(), and scale_color_manual().
8. What is faceting?
Creating multiple charts side by side.
9. Why is data visualization important?
It makes data easier to understand.
10. How do Nigerian businesses use data visualization?
They use it to make better decisions.

πŸ“ Review Questions (15)

  1. What is data visualization?
  2. What is ggplot2?
  3. What is the grammar of graphics?
  4. What is a bar chart?
  5. What is a scatter plot?
  6. What is a line chart?
  7. How do you customize a chart?
  8. What is faceting?
  9. Why is data visualization important?
  10. How do you create a bar chart in ggplot2?
  11. How do you create a scatter plot in ggplot2?
  12. How do you create a line chart in ggplot2?
  13. Give an example of a Nigerian business using data visualization.
  14. What is the best practice for data visualization?
  15. What is the difference between a bar chart and a line chart?

✏️ Fill-in-the-Blank Exercises

  1. Data __________ turns numbers into pictures.
  2. __________ is a package for creating charts in R.
  3. A __________ chart compares categories.
  4. A __________ plot shows relationships.
  5. A __________ chart shows trends over time.

βœ… True or False Exercises

  1. Data visualization is not important. (False)
  2. ggplot2 is a package for creating charts. (True)
  3. Bar charts show trends over time. (False)
  4. Scatter plots show relationships. (True)
  5. Line charts show trends over time. (True)

πŸ”˜ Multiple Choice Questions

  1. What is data visualization?
    A) Turning numbers into pictures B) Deleting data C) Adding data D) Moving data
    Answer: A
  2. What is ggplot2?
    A) A package for creating charts B) A game C) A type of food D) A sport
    Answer: A
  3. What is a bar chart?
    A) A chart that compares categories B) A chart that shows relationships C) A chart that shows trends D) A chart that shows proportions
    Answer: A
  4. What is a scatter plot?
    A) A chart that shows relationships B) A chart that compares categories C) A chart that shows trends D) A chart that shows proportions
    Answer: A
  5. What is a line chart?
    A) A chart that shows trends over time B) A chart that compares categories C) A chart that shows relationships D) A chart that shows proportions
    Answer: A
  6. What is faceting?
    A) Creating multiple charts B) Creating one chart C) Deleting a chart D) Moving a chart
    Answer: A
  7. How do you customize a chart?
    A) Use labs(), theme(), and scale_color_manual() B) Use filter() C) Use select() D) Use mutate()
    Answer: A
  8. What is the grammar of graphics?
    A) A system for building charts B) A system for cleaning data C) A system for importing data D) A system for analyzing data
    Answer: A
  9. Why is data visualization important?
    A) It makes data easier to understand B) It makes data harder to understand C) It makes data larger D) It makes data smaller
    Answer: A
  10. Which of these is a Nigerian business using data visualization?
    A) Paystack B) Flutterwave C) MTN D) All of the above
    Answer: D
  11. What does geom_bar() do?
    A) Creates a bar chart B) Creates a scatter plot C) Creates a line chart D) Creates a pie chart
    Answer: A
  12. What does geom_point() do?
    A) Creates a scatter plot B) Creates a bar chart C) Creates a line chart D) Creates a pie chart
    Answer: A
  13. What does geom_line() do?
    A) Creates a line chart B) Creates a bar chart C) Creates a scatter plot D) Creates a pie chart
    Answer: A
  14. What does facet_wrap() do?
    A) Creates multiple charts B) Creates one chart C) Deletes a chart D) Moves a chart
    Answer: A
  15. What is the best practice for data visualization?
    A) Choose the right chart B) Use many colours C) No titles D) No labels
    Answer: A

πŸ”— Matching Exercises

Match the word on the left with the correct meaning on the right:

Word Meaning
Bar Chart Shows trends over time
Scatter Plot Compares categories
Line Chart Shows relationships
Faceting Creates multiple charts

Answers: Bar Chart β†’ Compares categories; Scatter Plot β†’ Shows relationships; Line Chart β†’ Shows trends over time; Faceting β†’ Creates multiple charts.

πŸ“ Short Answer Questions

  1. Explain data visualization in your own words.
  2. What is the difference between a bar chart and a scatter plot?
  3. How do you create a bar chart in ggplot2?
  4. Why is data visualization important?
  5. Give an example of a Nigerian business using data visualization.

🎭 Scenario-based Exercises

Scenario 1: Ada has data on her class's favourite foods. She wants to show it to her teacher. What chart should she use? Why?

Scenario 2: A Lagos supermarket wants to see how sales have changed over the last year. What chart should they use? Why?

πŸ‘₯ Group Activity

In groups of 4–5, discuss how you would visualize sales data for a Nigerian business. What charts would you create? Present your ideas to the class.

πŸ§‘ Individual Activity

Think about a dataset you would like to visualize. Write a short paragraph (about 100 words) about the chart you would create.

πŸ—£οΈ Classroom Discussion Questions

  1. Why are charts important?
  2. What is your favourite type of chart and why?
  3. How can Nigerian businesses benefit from data visualization?
  4. What is the most interesting thing you learned about data visualization?
  5. How do you think data visualization will change in the future?

πŸ› οΈ Mini Project

Create a Visualization

Find a dataset and create a visualization using ggplot2. Include a title, labels, and a theme. Save and share your visualization with the class.

πŸ“‹ Practical Assignment

Use ggplot2 to create a bar chart and a scatter plot. Customize both charts with titles, colours, and themes. Write a short report (about 150 words) about what you did and what you learned.

πŸ† Challenge Exercise

The Challenge: Imagine you are a data analyst in Lagos. You have sales data for four regions. Create a bar chart, a scatter plot, and a line chart. Customize all charts and use faceting.

πŸ” Quiz Answers

Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A

Fill-in-the-Blank: 1. visualization, 2. ggplot2, 3. bar, 4. scatter, 5. line

True or False: 1. False, 2. True, 3. False, 4. True, 5. True

Matching: Bar Chart β†’ Compares categories; Scatter Plot β†’ Shows relationships; Line Chart β†’ Shows trends over time; Faceting β†’ Creates multiple charts.

🎯 Key Takeaways

  • Data visualization turns numbers into pictures.
  • ggplot2 is a package for creating charts.
  • Bar charts compare categories.
  • Scatter plots show relationships.
  • Line charts show trends over time.
  • Customizing makes your charts look professional.
  • Faceting creates multiple charts.
  • Nigerian businesses use data visualization to make better decisions.

πŸ”œ Preparation for Module Four

In the next module, we will explore Data Manipulation with dplyr. You will learn how to filter, select, mutate, and summarise data.


πŸŽ‰ Congratulations! You have completed Module Three of the R for Data Analysis course.

πŸ‘ You are now ready to move to Module Four: Data Manipulation with dplyr.

5

Module Four

Module Four: Data Manipulation with dplyr

πŸ”„ Module Four: Data Manipulation with dplyr

Welcome back, young data wrangler! Today we learn how to manipulate and transform our data.

🌟 Module Introduction

In Module Three, you learned how to create beautiful charts with ggplot2. But sometimes your data is not in the right shape for analysis. You need to manipulate it β€” filter, select, mutate, and summarise.

Imagine you have a big box of Lego bricks. You need to find the right pieces, sort them, and build something new. Data manipulation is like that β€” you take your data and shape it into what you need.

dplyr is a package in R that makes data manipulation easy. It is like a toolbox for your data.

In this module, you will learn about the dplyr verbs: filter(), select(), mutate(), summarise(), arrange(), and group_by(). You will also learn how to join datasets together.

πŸ’‘ Think about it: Have you ever had to sort through a big pile of things to find what you need? That is what data manipulation is all about.

🎯 Learning Objectives

By the time you finish this module, you will be able to:

  • Explain what dplyr is and why it is important.
  • Use filter() to select rows based on conditions.
  • Use select() to choose specific columns.
  • Use mutate() to create new columns.
  • Use summarise() to summarise data.
  • Use arrange() to sort data.
  • Use group_by() to group data for analysis.
  • Join datasets together.
  • Give examples of data manipulation in Nigerian businesses.

πŸ“– Warm-up Story: Chidi's Toy Box

Chidi is a 12-year-old boy who loves Lego. He has a big box of Lego bricks. He wanted to build a spaceship, but he needed to find the right pieces. He had to sort through all the bricks to find the ones he needed.

His uncle, who works with data, said, "Chidi, sorting through Lego is like sorting through data. You need to filter out the pieces you don't need, select the ones you want, and mutate them into something new."

Chidi learned how to use dplyr in R. He could filter, select, and mutate his data just like he sorted his Lego bricks.

🧠 Think about it: Have you ever had to sort through a big pile of things? That is what data manipulation is all about.

πŸ“š Main Lessons

1. What is dplyr?

Definition: dplyr is a package in R that helps you manipulate data.

Why it is important: dplyr makes data manipulation easy and fast.

Simple explanation: Think of dplyr like a toolbox for your data. It has tools to filter, select, mutate, and summarise.

🏫 School example: A teacher uses dplyr to filter student grades.

🏠 Home example: Your parents use dplyr to summarise expenses.

πŸ‡³πŸ‡¬ Nigerian example: A business uses dplyr to analyse sales data.

  • filter(): Select rows based on conditions.
  • select(): Choose specific columns.
  • mutate(): Create new columns.
  • summarise(): Summarise data.
  • arrange(): Sort data.
  • group_by(): Group data for analysis.
    +-------------------+
    |  dplyr TOOLBOX    |
    +-------------------+
    |  filter()         |  ← Select rows
    |  select()         |  ← Select columns
    |  mutate()         |  ← Create new columns
    |  summarise()      |  ← Summarise data
    |  arrange()        |  ← Sort data
    |  group_by()       |  ← Group data
    +-------------------+
    

πŸ“Œ Mini summary: dplyr is a toolbox for manipulating data.

2. Installing and Loading dplyr

Definition: Installing means downloading the package. Loading means making it available to use.

Why it is important: You need to install and load dplyr before you can use it.

Simple explanation: Think of installing like buying a new tool. Loading is like taking it out of the toolbox.

  • Install: install.packages("dplyr").
  • Load: library(dplyr).
# Installing dplyr install.packages("dplyr") # Loading dplyr library(dplyr)

🏫 School example: A teacher installs dplyr on the school computer.

🏠 Home example: Your parents install dplyr on their computer.

πŸ‡³πŸ‡¬ Nigerian example: A business installs dplyr on the company computer.

πŸ“Œ Mini summary: You need to install and load dplyr before using it.

3. The Pipe Operator (%>%)

Definition: The pipe operator (%>%) passes data from one function to the next.

Why it is important: The pipe makes your code easier to read and write.

Simple explanation: Think of the pipe like a conveyor belt that moves data from one step to the next.

  • Syntax: data %>% function1() %>% function2()
  • Example: sales_data %>% filter(Region == "Lagos") %>% summarise(total = sum(Sales)).
# Without pipe filtered <- filter(sales_data, Region == "Lagos") summarised <- summarise(filtered, total = sum(Sales)) # With pipe sales_data %>% filter(Region == "Lagos") %>% summarise(total = sum(Sales))

🏫 School example: A teacher uses the pipe to chain operations.

🏠 Home example: Your parents use the pipe to chain operations.

πŸ‡³πŸ‡¬ Nigerian example: A business uses the pipe to chain operations.

πŸ“Œ Mini summary: The pipe passes data from one function to the next.

4. filter() β€” Selecting Rows

Definition: filter() selects rows that meet certain conditions.

Why it is important: filter() helps you focus on the data that matters.

Simple explanation: Think of filter() like using a sieve to separate what you want.

  • Syntax: filter(data, condition)
  • Example: filter(sales_data, Region == "Lagos")
  • Multiple conditions: filter(sales_data, Region == "Lagos", Sales > 1000)
# Using filter() lagos_data <- sales_data %>% filter(Region == "Lagos") high_sales <- sales_data %>% filter(Sales > 1000) # Multiple conditions lagos_high <- sales_data %>% filter(Region == "Lagos", Sales > 1000)

🏫 School example: A teacher filters students by grade.

🏠 Home example: Your parents filter expenses by category.

πŸ‡³πŸ‡¬ Nigerian example: A business filters sales by region.

πŸ“Œ Mini summary: filter() selects rows based on conditions.

5. select() β€” Selecting Columns

Definition: select() chooses specific columns.

Why it is important: select() helps you focus on the columns that matter.

Simple explanation: Think of select() like choosing the right tools for a job.

  • Syntax: select(data, column1, column2)
  • Example: select(sales_data, Region, Sales)
  • Remove columns: select(sales_data, -Region)
# Using select() selected_data <- sales_data %>% select(Region, Sales, Profit) # Remove a column no_region <- sales_data %>% select(-Region)

🏫 School example: A teacher selects only the grade column.

🏠 Home example: Your parents select only the expense column.

πŸ‡³πŸ‡¬ Nigerian example: A business selects only the sales column.

πŸ“Œ Mini summary: select() chooses specific columns.

6. mutate() β€” Creating New Columns

Definition: mutate() creates new columns from existing columns.

Why it is important: mutate() helps you create new variables.

Simple explanation: Think of mutate() like adding new pieces to your Lego creation.

  • Syntax: mutate(data, new_column = expression)
  • Example: mutate(sales_data, Profit = Sales - Cost)
# Using mutate() sales_data <- sales_data %>% mutate(Profit = Sales - Cost) # Multiple new columns sales_data <- sales_data %>% mutate(Profit = Sales - Cost, Margin = Profit / Sales)

🏫 School example: A teacher creates a new column for letter grades.

🏠 Home example: Your parents create a new column for expense categories.

πŸ‡³πŸ‡¬ Nigerian example: A business creates a new column for profit.

πŸ“Œ Mini summary: mutate() creates new columns.

7. summarise() β€” Summarising Data

Definition: summarise() creates summary statistics for your data.

Why it is important: summarise() helps you understand your data.

Simple explanation: Think of summarise() like getting a summary of a book.

  • Syntax: summarise(data, new_name = function(column))
  • Example: summarise(sales_data, total = sum(Sales), avg = mean(Sales))
# Using summarise() summary <- sales_data %>% summarise(total_sales = sum(Sales), avg_sales = mean(Sales), count = n())

🏫 School example: A teacher summarises grades.

🏠 Home example: Your parents summarise expenses.

πŸ‡³πŸ‡¬ Nigerian example: A business summarises sales.

πŸ“Œ Mini summary: summarise() creates summary statistics.

8. arrange() β€” Sorting Data

Definition: arrange() sorts your data.

Why it is important: arrange() helps you order your data.

Simple explanation: Think of arrange() like organising your books on a shelf.

  • Syntax: arrange(data, column)
  • Example: arrange(sales_data, Sales)
  • Descending: arrange(sales_data, desc(Sales))
# Using arrange() sorted_data <- sales_data %>% arrange(Sales) # Descending order sorted_data <- sales_data %>% arrange(desc(Sales))

🏫 School example: A teacher sorts students by grade.

🏠 Home example: Your parents sort expenses by amount.

πŸ‡³πŸ‡¬ Nigerian example: A business sorts sales by region.

πŸ“Œ Mini summary: arrange() sorts your data.

9. group_by() β€” Grouping Data

Definition: group_by() groups your data by one or more variables.

Why it is important: group_by() allows you to perform operations on groups.

Simple explanation: Think of group_by() like sorting your Lego bricks by colour.

  • Syntax: group_by(data, column)
  • Example: group_by(sales_data, Region)
  • With summarise: group_by(sales_data, Region) %>% summarise(total = sum(Sales))
# Using group_by() with summarise() region_summary <- sales_data %>% group_by(Region) %>% summarise(total_sales = sum(Sales), avg_sales = mean(Sales), count = n())

🏫 School example: A teacher groups students by class.

🏠 Home example: Your parents group expenses by category.

πŸ‡³πŸ‡¬ Nigerian example: A business groups sales by region.

πŸ“Œ Mini summary: group_by() groups your data.

10. Joining Datasets

Definition: Joining combines two datasets based on a common column.

Why it is important: Joining allows you to combine data from different sources.

Simple explanation: Think of joining like connecting two Lego pieces.

  • inner_join(): Keeps only matching rows.
  • left_join(): Keeps all rows from the left table.
  • right_join(): Keeps all rows from the right table.
  • full_join(): Keeps all rows from both tables.
# Using joins sales_data <- read.csv("sales.csv") customer_data <- read.csv("customers.csv") # Inner join joined_data <- sales_data %>% inner_join(customer_data, by = "CustomerID")

🏫 School example: A teacher joins student data with grade data.

🏠 Home example: Your parents join expense data with category data.

πŸ‡³πŸ‡¬ Nigerian example: A business joins sales data with customer data.

πŸ“Œ Mini summary: Joining combines two datasets.

πŸ“– Key Vocabulary

Word Simple Meaning
dplyr A package for manipulating data.
filter() Selects rows based on conditions.
select() Chooses specific columns.
mutate() Creates new columns.
summarise() Creates summary statistics.
arrange() Sorts data.
group_by() Groups data.
Pipe (%>%) Passes data from one function to the next.
Join Combines two datasets.
inner_join() Keeps only matching rows.
left_join() Keeps all rows from the left table.
right_join() Keeps all rows from the right table.
full_join() Keeps all rows from both tables.
summarise() Creates summary statistics.
n() Counts the number of rows.

🧩 Important Concepts

  • dplyr is a toolbox for manipulating data.
  • filter() selects rows based on conditions.
  • select() chooses specific columns.
  • mutate() creates new columns.
  • summarise() creates summary statistics.
  • arrange() sorts your data.
  • group_by() groups your data.
  • The pipe (%>%) passes data from one function to the next.
  • Joining combines two datasets.
  • Nigerian businesses use dplyr to analyse data.

πŸ“Œ Step-by-Step Explanations

How to use dplyr verbs

  1. Load dplyr: library(dplyr).
  2. Import data: data <- read.csv("filename.csv").
  3. Filter rows: data %>% filter(condition).
  4. Select columns: data %>% select(column1, column2).
  5. Create new columns: data %>% mutate(new = expression).
  6. Summarise data: data %>% summarise(total = sum(column)).
  7. Sort data: data %>% arrange(column).
  8. Group data: data %>% group_by(column) %>% summarise(total = sum(column)).

🌍 Real-life Examples

  • School: A teacher filters students by grade.
  • Hospital: A hospital filters patients by condition.
  • Restaurant: A restaurant filters sales by date.
  • Shop: A shop filters sales by product.

πŸ‡³πŸ‡¬ Nigerian Examples

  • Paystack: Filters transactions by status.
  • Flutterwave: Summarises payments by currency.
  • MTN Nigeria: Groups customers by region.
  • A Lagos supermarket: Summarises sales by product.
  • A Nigerian bank: Groups branches by performance.

🎈 Fun Examples for You

  • Your pocket money: Filter expenses by category.
  • Your game scores: Sort scores by date.
  • Your reading log: Group books by genre.
  • Your chores: Summarise chores by completion.

🏠 Everyday Examples

  • At home: Your parents filter expenses by category.
  • At school: Your teacher sorts students by grade.
  • In your community: Local businesses summarise sales.
  • In your own life: You can filter and sort your own data.

πŸ‘©β€πŸ« Teacher Notes

  • Encourage students to think about how they sort and filter things in daily life.
  • Use the warm-up story to spark interest in data manipulation.
  • Demonstrate dplyr verbs in R.
  • Discuss the importance of data manipulation.
  • Ask students to think about how data manipulation is used in Nigerian businesses.

πŸ‘ͺ Parent Tips

  • Talk to your child about how you sort and filter data in your work.
  • Show your child how you organise data in Excel or other tools.
  • Open R and explore dplyr together.
  • Encourage your child to think about how they would manipulate data.
  • Share examples of data manipulation in Nigerian businesses.

🧠 Interesting Facts

  • dplyr was created in 2014.
  • dplyr is part of the tidyverse.
  • dplyr is used by millions of people.
  • dplyr is one of the most popular R packages.
  • dplyr is constantly being updated.

πŸ’‘ Did You Know?

  • Did you know that dplyr can handle millions of rows?
  • Did you know that dplyr is part of the tidyverse?
  • Did you know that the pipe operator was inspired by Unix?
  • Did you know that dplyr has functions for working with databases?
  • Did you know that Nigerian businesses use dplyr for data analysis?

πŸ”” Remember This

  • dplyr is a toolbox for manipulating data.
  • filter() selects rows based on conditions.
  • select() chooses specific columns.
  • mutate() creates new columns.
  • summarise() creates summary statistics.
  • arrange() sorts your data.
  • group_by() groups your data.
  • The pipe (%>%) passes data from one function to the next.
  • Joining combines two datasets.
  • Nigerian businesses use dplyr to analyse data.

⚠️ Common Mistakes

  • Not loading dplyr: Always load dplyr with library(dplyr).
  • Forgetting the pipe: Use the pipe to chain operations.
  • Using the wrong join: Choose the right join for your data.
  • Not using group_by(): Use group_by() before summarise().
  • Not checking data types: Make sure your data is the right type.

✨ Best Practices for Data Manipulation

  • Always load dplyr with library(dplyr).
  • Use the pipe (%>%) to chain operations.
  • Use filter() to select rows.
  • Use select() to choose columns.
  • Use mutate() to create new columns.
  • Use summarise() to summarise data.
  • Use group_by() to group data.
  • Check your data types.
  • Keep your code clean and readable.

πŸ“Š Clear Illustrations

1. dplyr Verbs

    +-------------------+
    |  filter()         |  ← Select rows
    +-------------------+
    |  select()         |  ← Select columns
    +-------------------+
    |  mutate()         |  ← Create new columns
    +-------------------+
    |  summarise()      |  ← Summarise data
    +-------------------+
    |  arrange()        |  ← Sort data
    +-------------------+
    |  group_by()       |  ← Group data
    +-------------------+
    

2. The Pipe Operator

    +-------------------+
    |  data             |  ← Start with data
    +-------------------+
           |
           V
    +-------------------+
    |  filter()         |  ← Filter rows
    +-------------------+
           |
           V
    +-------------------+
    |  select()         |  ← Select columns
    +-------------------+
           |
           V
    +-------------------+
    |  summarise()      |  ← Summarise data
    +-------------------+
    

3. Joining Data

    +-------------------+        +-------------------+
    |  Table A          |        |  Table B          |
    |  ID | Name        |        |  ID | Sales       |
    +-------------------+        +-------------------+
           |                             |
           +-------------+---------------+
                         |
                         V
    +-------------------+
    |  Joined Table     |
    |  ID | Name | Sales|
    +-------------------+
    

4. Comparison: dplyr vs Base R

dplyr Base R
filter() data[data$column == condition, ]
select() data[, c("col1", "col2")]
mutate() data$new <- expression
summarise() sum(data$column)
arrange() data[order(data$column), ]
group_by() aggregate(data, by = list(group), FUN = sum)

πŸ“ Lesson Summaries

Lesson 1: dplyr is a toolbox for manipulating data.

Lesson 2: You need to install and load dplyr before using it.

Lesson 3: The pipe (%>%) passes data from one function to the next.

Lesson 4: filter() selects rows based on conditions.

Lesson 5: select() chooses specific columns.

Lesson 6: mutate() creates new columns.

Lesson 7: summarise() creates summary statistics.

Lesson 8: arrange() sorts your data.

Lesson 9: group_by() groups your data.

Lesson 10: Joining combines two datasets.

πŸ“˜ End-of-Module Summary

In this module, you learned about data manipulation with dplyr. You discovered the dplyr verbs: filter(), select(), mutate(), summarise(), arrange(), and group_by(). You also learned how to use the pipe operator and how to join datasets together.

🎯 You can now:

  • Explain what dplyr is and why it is important.
  • Use filter() to select rows based on conditions.
  • Use select() to choose specific columns.
  • Use mutate() to create new columns.
  • Use summarise() to summarise data.
  • Use arrange() to sort data.
  • Use group_by() to group data for analysis.
  • Join datasets together.
  • Give examples of data manipulation in Nigerian businesses.

❓ Frequently Asked Questions

1. What is dplyr?
dplyr is a package for manipulating data.
2. What is filter()?
filter() selects rows based on conditions.
3. What is select()?
select() chooses specific columns.
4. What is mutate()?
mutate() creates new columns.
5. What is summarise()?
summarise() creates summary statistics.
6. What is arrange()?
arrange() sorts your data.
7. What is group_by()?
group_by() groups your data.
8. What is the pipe operator?
The pipe (%>%) passes data from one function to the next.
9. What is joining?
Joining combines two datasets.
10. How do Nigerian businesses use dplyr?
They use it to analyse data.

πŸ“ Review Questions (15)

  1. What is dplyr?
  2. What does filter() do?
  3. What does select() do?
  4. What does mutate() do?
  5. What does summarise() do?
  6. What does arrange() do?
  7. What does group_by() do?
  8. What is the pipe operator?
  9. What is joining?
  10. How do you load dplyr?
  11. How do you filter data?
  12. How do you select columns?
  13. How do you create a new column?
  14. How do you summarise data?
  15. Give an example of a Nigerian business using dplyr.

✏️ Fill-in-the-Blank Exercises

  1. dplyr is a package for __________ data.
  2. __________ selects rows based on conditions.
  3. __________ chooses specific columns.
  4. __________ creates new columns.
  5. __________ creates summary statistics.

βœ… True or False Exercises

  1. dplyr is a package for manipulating data. (True)
  2. filter() selects columns. (False)
  3. select() chooses specific columns. (True)
  4. mutate() creates new columns. (True)
  5. summarise() sorts data. (False)

πŸ”˜ Multiple Choice Questions

  1. What is dplyr?
    A) A package for manipulating data B) A game C) A type of food D) A sport
    Answer: A
  2. What does filter() do?
    A) Selects rows B) Selects columns C) Creates new columns D) Sorts data
    Answer: A
  3. What does select() do?
    A) Selects rows B) Selects columns C) Creates new columns D) Sorts data
    Answer: B
  4. What does mutate() do?
    A) Selects rows B) Selects columns C) Creates new columns D) Sorts data
    Answer: C
  5. What does summarise() do?
    A) Selects rows B) Selects columns C) Creates new columns D) Creates summary statistics
    Answer: D
  6. What does arrange() do?
    A) Selects rows B) Selects columns C) Creates new columns D) Sorts data
    Answer: D
  7. What does group_by() do?
    A) Selects rows B) Selects columns C) Groups data D) Sorts data
    Answer: C
  8. What is the pipe operator?
    A) %>% B) %in% C) %*% D) %/%
    Answer: A
  9. What is joining?
    A) Combining two datasets B) Selecting rows C) Selecting columns D) Sorting data
    Answer: A
  10. Which of these is a Nigerian business using dplyr?
    A) Paystack B) Flutterwave C) MTN D) All of the above
    Answer: D
  11. How do you load dplyr?
    A) library(dplyr) B) install(dplyr) C) load(dplyr) D) use(dplyr)
    Answer: A
  12. How do you filter data?
    A) data %>% filter(condition) B) data %>% select(condition) C) data %>% mutate(condition) D) data %>% summarise(condition)
    Answer: A
  13. How do you select columns?
    A) data %>% filter(column) B) data %>% select(column) C) data %>% mutate(column) D) data %>% summarise(column)
    Answer: B
  14. How do you create a new column?
    A) data %>% filter(new = expression) B) data %>% select(new = expression) C) data %>% mutate(new = expression) D) data %>% summarise(new = expression)
    Answer: C
  15. How do you summarise data?
    A) data %>% filter(total = sum(column)) B) data %>% select(total = sum(column)) C) data %>% mutate(total = sum(column)) D) data %>% summarise(total = sum(column))
    Answer: D

πŸ”— Matching Exercises

Match the word on the left with the correct meaning on the right:

Word Meaning
filter() Sorts data
select() Selects rows
mutate() Selects columns
arrange() Creates new columns

Answers: filter() β†’ Selects rows; select() β†’ Selects columns; mutate() β†’ Creates new columns; arrange() β†’ Sorts data.

πŸ“ Short Answer Questions

  1. Explain dplyr in your own words.
  2. What is the difference between filter() and select()?
  3. What is the pipe operator and why is it useful?
  4. How do you group data in dplyr?
  5. Give an example of a Nigerian business using dplyr.

🎭 Scenario-based Exercises

Scenario 1: Chidi has a dataset of student grades. He wants to filter students who scored above 80 and create a new column for letter grades. How can dplyr help him?

Scenario 2: A Lagos supermarket has sales data for different products. They want to summarise total sales by product category. How can dplyr help them?

πŸ‘₯ Group Activity

In groups of 4–5, discuss how you would manipulate sales data for a Nigerian business. What dplyr verbs would you use? Present your ideas to the class.

πŸ§‘ Individual Activity

Think about a dataset you would like to manipulate. Write a short paragraph (about 100 words) about what dplyr verbs you would use.

πŸ—£οΈ Classroom Discussion Questions

  1. Why is data manipulation important?
  2. What is your favourite dplyr verb and why?
  3. How can Nigerian businesses benefit from data manipulation?
  4. What is the most interesting thing you learned about dplyr?
  5. How do you think data manipulation will change in the future?

πŸ› οΈ Mini Project

Manipulate a Dataset

Find a dataset and use dplyr to manipulate it. Filter, select, mutate, summarise, and group the data. Save and share your results with the class.

πŸ“‹ Practical Assignment

Use dplyr to filter, select, mutate, summarise, and group a dataset. Write a short report (about 150 words) about what you did and what you learned.

πŸ† Challenge Exercise

The Challenge: Imagine you are a data analyst in Lagos. You have sales data for four regions. Use dplyr to filter each region, summarise total sales, and arrange by sales. Create a new column for profit.

πŸ” Quiz Answers

Multiple Choice: 1-A, 2-A, 3-B, 4-C, 5-D, 6-D, 7-C, 8-A, 9-A, 10-D, 11-A, 12-A, 13-B, 14-C, 15-D

Fill-in-the-Blank: 1. manipulating, 2. filter(), 3. select(), 4. mutate(), 5. summarise()

True or False: 1. True, 2. False, 3. True, 4. True, 5. False

Matching: filter() β†’ Selects rows; select() β†’ Selects columns; mutate() β†’ Creates new columns; arrange() β†’ Sorts data.

🎯 Key Takeaways

  • dplyr is a toolbox for manipulating data.
  • filter() selects rows based on conditions.
  • select() chooses specific columns.
  • mutate() creates new columns.
  • summarise() creates summary statistics.
  • arrange() sorts your data.
  • group_by() groups your data.
  • The pipe (%>%) passes data from one function to the next.
  • Joining combines two datasets.
  • Nigerian businesses use dplyr to analyse data.

πŸ”œ Preparation for Module Five

In the next module, we will explore Statistical Analysis and Hypothesis Testing. You will learn how to perform statistical tests and analyse your data.


πŸŽ‰ Congratulations! You have completed Module Four of the R for Data Analysis course.

πŸ‘ You are now ready to move to Module Five: Statistical Analysis and Hypothesis Testing.

6

Module FIve

Module Five: Statistical Analysis and Hypothesis Testing

πŸ“Š Module Five: Statistical Analysis and Hypothesis Testing

Welcome back, young data scientist! Today we learn how to understand and test our data.

🌟 Module Introduction

In Module Four, you learned how to manipulate data. Now you have clean, organised data. But how do you understand it? How do you know if there is a real pattern or just random chance?

Imagine you are a detective. You have clues (your data). You need to figure out what they mean. Statistics is like your detective toolkit. It helps you understand your data and make decisions.

Hypothesis testing is a way to test if something is true. It is like a lie detector for your data.

In this module, you will learn about descriptive statistics, t-tests, ANOVA, correlation, and linear regression. By the end, you will be able to understand and test your data.

πŸ’‘ Think about it: Have you ever wondered if something was true or just a coincidence? Statistics can help you find out.

🎯 Learning Objectives

By the time you finish this module, you will be able to:

  • Explain what statistics is and why it is important.
  • Calculate descriptive statistics (mean, median, mode, standard deviation).
  • Understand the difference between population and sample.
  • Perform a t-test to compare two groups.
  • Perform ANOVA to compare multiple groups.
  • Calculate correlation to measure relationships.
  • Perform linear regression to predict outcomes.
  • Give examples of statistical analysis in Nigerian businesses.

πŸ“– Warm-up Story: Ada's Mystery

Ada is a 12-year-old girl who loves mysteries. She noticed that her class performed better on tests on Fridays than on Mondays. She wondered, "Is this just a coincidence, or is there a real difference?"

Her uncle, who works with data, said, "Ada, you need to use statistics. Statistics can help you find out if the difference is real or just random chance."

Ada learned about hypothesis testing. She tested her data and found that the difference was real. She was like a detective solving a mystery.

🧠 Think about it: Have you ever wondered if something was true or just a coincidence? Statistics can help you find out.

πŸ“š Main Lessons

1. What is Statistics?

Definition: Statistics is the science of collecting, analysing, and interpreting data.

Why it is important: Statistics helps you make sense of data and make decisions.

Simple explanation: Think of statistics like a magnifying glass for your data. It helps you see what is really there.

🏫 School example: A teacher uses statistics to understand student performance.

🏠 Home example: Your parents use statistics to understand their spending.

πŸ‡³πŸ‡¬ Nigerian example: A business uses statistics to understand sales.

  • Descriptive statistics: Summarise data (mean, median, mode, standard deviation).
  • Inferential statistics: Make predictions and test hypotheses.
    +-------------------+
    |  STATISTICS       |
    +-------------------+
    |  Descriptive      |  ← Summarise data
    |  Inferential      |  ← Make predictions
    +-------------------+
    

πŸ“Œ Mini summary: Statistics is the science of understanding data.

2. Descriptive Statistics

Definition: Descriptive statistics summarise your data.

Why it is important: They help you understand your data at a glance.

Simple explanation: Think of descriptive statistics like a summary of a book.

  • Mean: The average. sum(x) / length(x).
  • Median: The middle value.
  • Mode: The most common value.
  • Standard deviation: How spread out the data is.
  • Range: max(x) - min(x).
  • Quartiles: Divide data into four parts.
# Descriptive statistics in R mean(sales_data$Sales) median(sales_data$Sales) sd(sales_data$Sales) summary(sales_data$Sales)

🏫 School example: A teacher calculates the mean grade.

🏠 Home example: Your parents calculate the mean expense.

πŸ‡³πŸ‡¬ Nigerian example: A business calculates the mean sales.

πŸ“Œ Mini summary: Descriptive statistics summarise your data.

3. Population vs. Sample

Definition: A population is the entire group you are studying. A sample is a subset of the population.

Why it is important: You often can't study the whole population, so you use a sample.

Simple explanation: Think of a population like all the students in a school. A sample is like one class.

  • Population: Everyone in the group.
  • Sample: A smaller part of the group.
  • Parameter: A number that describes a population.
  • Statistic: A number that describes a sample.

🏫 School example: All students in the school (population) vs. one class (sample).

🏠 Home example: All family members (population) vs. one person (sample).

πŸ‡³πŸ‡¬ Nigerian example: All customers (population) vs. a survey group (sample).

πŸ“Œ Mini summary: Population is the whole group; sample is a part of it.

4. Hypothesis Testing

Definition: Hypothesis testing is a way to test if something is true.

Why it is important: It helps you make decisions based on data.

Simple explanation: Think of hypothesis testing like a lie detector for your data.

  • Null hypothesis (H0): There is no effect or difference.
  • Alternative hypothesis (H1): There is an effect or difference.
  • p-value: The probability of seeing the data if the null hypothesis is true.
  • Significance level (alpha): Usually 0.05. If p-value < 0.05, reject H0.
# Hypothesis testing steps # 1. State H0 and H1 # 2. Choose significance level (alpha = 0.05) # 3. Calculate test statistic # 4. Calculate p-value # 5. Make decision (reject H0 if p-value < alpha)

🏫 School example: Testing if students who study more get better grades.

🏠 Home example: Testing if eating breakfast improves test scores.

πŸ‡³πŸ‡¬ Nigerian example: Testing if a new marketing campaign increases sales.

πŸ“Œ Mini summary: Hypothesis testing helps you make decisions.

5. The t-Test

Definition: A t-test compares the means of two groups.

Why it is important: It helps you see if two groups are different.

Simple explanation: Think of a t-test like comparing two things.

  • One-sample t-test: Compares a sample mean to a known value.
  • Two-sample t-test: Compares the means of two groups.
  • Paired t-test: Compares two related groups.
# One-sample t-test t.test(sales_data$Sales, mu = 1000) # Two-sample t-test t.test(sales_data$Sales ~ sales_data$Region) # Paired t-test t.test(before, after, paired = TRUE)

🏫 School example: Comparing grades of two classes.

🏠 Home example: Comparing expenses before and after a budget.

πŸ‡³πŸ‡¬ Nigerian example: Comparing sales of two products.

πŸ“Œ Mini summary: A t-test compares two groups.

6. ANOVA

Definition: ANOVA compares the means of three or more groups.

Why it is important: It helps you see if multiple groups are different.

Simple explanation: Think of ANOVA like comparing many things at once.

  • One-way ANOVA: One independent variable.
  • Two-way ANOVA: Two independent variables.
# One-way ANOVA aov_result <- aov(Sales ~ Region, data = sales_data) summary(aov_result)

🏫 School example: Comparing grades across three classes.

🏠 Home example: Comparing expenses across three categories.

πŸ‡³πŸ‡¬ Nigerian example: Comparing sales across four regions.

πŸ“Œ Mini summary: ANOVA compares three or more groups.

7. Correlation

Definition: Correlation measures the relationship between two variables.

Why it is important: It helps you see if two things are related.

Simple explanation: Think of correlation like a friendship between two variables.

  • Positive correlation: As one goes up, the other goes up.
  • Negative correlation: As one goes up, the other goes down.
  • Zero correlation: No relationship.
  • Pearson correlation: Measures linear relationship.
  • Spearman correlation: Measures rank relationship.
# Correlation cor(sales_data$Sales, sales_data$Profit) # Correlation matrix cor(sales_data[, c("Sales", "Profit", "Cost")])

🏫 School example: Correlation between study time and grades.

🏠 Home example: Correlation between income and savings.

πŸ‡³πŸ‡¬ Nigerian example: Correlation between advertising spend and sales.

πŸ“Œ Mini summary: Correlation measures relationships.

8. Linear Regression

Definition: Linear regression models the relationship between a dependent variable and one or more independent variables.

Why it is important: It helps you predict outcomes.

Simple explanation: Think of linear regression like drawing a line through your data.

  • Simple linear regression: One independent variable.
  • Multiple linear regression: Multiple independent variables.
  • Formula: y = a + b*x
# Simple linear regression model <- lm(Sales ~ Profit, data = sales_data) summary(model) # Multiple linear regression model <- lm(Sales ~ Profit + Cost, data = sales_data) summary(model)

🏫 School example: Predicting grades based on study time.

🏠 Home example: Predicting expenses based on income.

πŸ‡³πŸ‡¬ Nigerian example: Predicting sales based on advertising spend.

πŸ“Œ Mini summary: Linear regression predicts outcomes.

9. Interpreting Results

Definition: Interpreting results means understanding what your analysis tells you.

Why it is important: You need to understand your results to make decisions.

Simple explanation: Think of interpreting results like reading a map.

  • p-value: If p-value < 0.05, the result is significant.
  • R-squared: How well the model fits the data.
  • Coefficient: The strength and direction of the relationship.
  • Confidence interval: The range where the true value lies.

🏫 School example: A teacher interprets test scores.

🏠 Home example: Your parents interpret expense data.

πŸ‡³πŸ‡¬ Nigerian example: A business interprets sales data.

πŸ“Œ Mini summary: Interpreting results helps you make decisions.

10. Statistics in Nigerian Businesses

Definition: Nigerian businesses use statistics to make decisions.

Why it is important: Statistics help Nigerian businesses understand their data.

Simple explanation: Think of it like using a compass to find your way.

  • Banks: Use statistics to analyse customer data.
  • Telecom: Use statistics to analyse network performance.
  • Retail: Use statistics to analyse sales.
  • Government: Use statistics to analyse economic data.
  • Startups: Use statistics to understand their customers.

πŸ‡³πŸ‡¬ Nigerian example: A Lagos supermarket uses statistics to understand sales patterns.

πŸ“Œ Mini summary: Nigerian businesses use statistics to make decisions.

πŸ“– Key Vocabulary

Word Simple Meaning
Statistics The science of understanding data.
Mean The average.
Median The middle value.
Standard Deviation How spread out the data is.
Population The entire group.
Sample A part of the group.
Hypothesis A testable statement.
p-value The probability of seeing the data if the null hypothesis is true.
t-test Compares two groups.
ANOVA Compares three or more groups.
Correlation Measures relationships.
Linear Regression Predicts outcomes.
R-squared How well the model fits.
Confidence Interval The range where the true value lies.
Significance Level Usually 0.05.

🧩 Important Concepts

  • Statistics is the science of understanding data.
  • Descriptive statistics summarise data.
  • Population is the whole group; sample is a part.
  • Hypothesis testing helps you make decisions.
  • A t-test compares two groups.
  • ANOVA compares three or more groups.
  • Correlation measures relationships.
  • Linear regression predicts outcomes.
  • Nigerian businesses use statistics to make decisions.

πŸ“Œ Step-by-Step Explanations

How to perform a t-test

  1. State H0 and H1: H0: no difference, H1: difference.
  2. Choose alpha: Usually 0.05.
  3. Calculate t-statistic: Use t.test().
  4. Calculate p-value: Use t.test().
  5. Make decision: If p-value < 0.05, reject H0.

How to perform linear regression

  1. Choose variables: Dependent and independent.
  2. Create model: lm(y ~ x, data).
  3. View summary: summary(model).
  4. Interpret results: Look at coefficients, R-squared, p-values.
  5. Make predictions: predict(model, newdata).

🌍 Real-life Examples

  • School: A teacher uses statistics to understand grades.
  • Hospital: A hospital uses statistics to understand patient data.
  • Restaurant: A restaurant uses statistics to understand sales.
  • Shop: A shop uses statistics to understand inventory.

πŸ‡³πŸ‡¬ Nigerian Examples

  • Paystack: Uses statistics to analyse payment data.
  • Flutterwave: Uses statistics to analyse transaction trends.
  • MTN Nigeria: Uses statistics to analyse network performance.
  • A Lagos supermarket: Uses statistics to analyse sales.
  • A Nigerian bank: Uses statistics to analyse customer data.

🎈 Fun Examples for You

  • Your pocket money: Calculate the mean of your spending.
  • Your game scores: Calculate the median of your scores.
  • Your reading log: Calculate the standard deviation of books read.
  • Your chores: Calculate the correlation between chores and rewards.

🏠 Everyday Examples

  • At home: Your parents use statistics to understand spending.
  • At school: Your teacher uses statistics to understand grades.
  • In your community: Local businesses use statistics to understand sales.
  • In your own life: You can use statistics to understand your habits.

πŸ‘©β€πŸ« Teacher Notes

  • Encourage students to think about statistics in their daily lives.
  • Use the warm-up story to spark interest in statistics.
  • Demonstrate statistical tests in R.
  • Discuss the importance of statistics.
  • Ask students to think about how statistics are used in Nigerian businesses.

πŸ‘ͺ Parent Tips

  • Talk to your child about how you use statistics in your work.
  • Show your child how you calculate averages and other statistics.
  • Open R and explore statistical tests together.
  • Encourage your child to think about how they would use statistics.
  • Share examples of statistics in Nigerian businesses.

🧠 Interesting Facts

  • Statistics is used in almost every field.
  • The p-value was invented in the 1920s.
  • R was created for statistical analysis.
  • Statistics can help you make better decisions.
  • Nigerian businesses are increasingly using statistics.

πŸ’‘ Did You Know?

  • Did you know that statistics can help you predict the future?
  • Did you know that the t-test was invented by William Gosset?
  • Did you know that ANOVA stands for Analysis of Variance?
  • Did you know that correlation does not imply causation?
  • Did you know that Nigerian banks use statistics to detect fraud?

πŸ”” Remember This

  • Statistics is the science of understanding data.
  • Descriptive statistics summarise data.
  • Population is the whole group; sample is a part.
  • Hypothesis testing helps you make decisions.
  • A t-test compares two groups.
  • ANOVA compares three or more groups.
  • Correlation measures relationships.
  • Linear regression predicts outcomes.
  • Nigerian businesses use statistics to make decisions.

⚠️ Common Mistakes

  • Confusing correlation with causation: Just because two things are related doesn't mean one causes the other.
  • Using the wrong test: Choose the right test for your data.
  • Ignoring assumptions: Make sure your data meets the assumptions of the test.
  • Misinterpreting p-values: A p-value less than 0.05 does not mean the effect is large.
  • Not checking for outliers: Outliers can affect your results.

✨ Best Practices for Statistics

  • Choose the right test for your data.
  • Check your data for outliers.
  • Interpret your results carefully.
  • Visualise your data to understand it better.
  • Document your analysis steps.
  • Share your findings with others.
  • Keep learning and exploring new statistical methods.

πŸ“Š Clear Illustrations

1. Descriptive Statistics

    +-------------------+
    |  DATA             |  ← 10, 20, 30, 40, 50
    +-------------------+
           |
           V
    +-------------------+
    |  MEAN             |  ← 30
    +-------------------+
    |  MEDIAN           |  ← 30
    +-------------------+
    |  STANDARD DEV     |  ← 15.81
    +-------------------+
    

2. t-Test

    +-------------------+
    |  GROUP A          |  ← 10, 20, 30
    +-------------------+
    |  GROUP B          |  ← 40, 50, 60
    +-------------------+
           |
           V
    +-------------------+
    |  t-TEST           |  ← p-value = 0.02
    +-------------------+
           |
           V
    +-------------------+
    |  SIGNIFICANT      |  ← Reject H0
    +-------------------+
    

3. Linear Regression

    +-------------------+
    |  DATA POINTS      |  ← *  *  *
    +-------------------+
           |
           V
    +-------------------+
    |  REGRESSION LINE  |  ← y = 2x + 3
    +-------------------+
           |
           V
    +-------------------+
    |  PREDICT          |  ← y = 2(5) + 3 = 13
    +-------------------+
    

4. Comparison: Statistics Tests

Test Purpose Number of Groups
t-test Compare means 2
ANOVA Compare means 3+
Correlation Measure relationship 2 variables
Linear Regression Predict outcomes 2+ variables

πŸ“ Lesson Summaries

Lesson 1: Statistics is the science of understanding data.

Lesson 2: Descriptive statistics summarise data.

Lesson 3: Population is the whole group; sample is a part.

Lesson 4: Hypothesis testing helps you make decisions.

Lesson 5: A t-test compares two groups.

Lesson 6: ANOVA compares three or more groups.

Lesson 7: Correlation measures relationships.

Lesson 8: Linear regression predicts outcomes.

Lesson 9: Interpreting results helps you make decisions.

Lesson 10: Nigerian businesses use statistics to make decisions.

πŸ“˜ End-of-Module Summary

In this module, you learned about statistical analysis and hypothesis testing. You discovered descriptive statistics, the difference between population and sample, and hypothesis testing. You also learned about t-tests, ANOVA, correlation, and linear regression.

🎯 You can now:

  • Explain what statistics is and why it is important.
  • Calculate descriptive statistics (mean, median, mode, standard deviation).
  • Understand the difference between population and sample.
  • Perform a t-test to compare two groups.
  • Perform ANOVA to compare multiple groups.
  • Calculate correlation to measure relationships.
  • Perform linear regression to predict outcomes.
  • Give examples of statistical analysis in Nigerian businesses.

❓ Frequently Asked Questions

1. What is statistics?
Statistics is the science of understanding data.
2. What is descriptive statistics?
Descriptive statistics summarise data.
3. What is the difference between population and sample?
Population is the whole group; sample is a part.
4. What is hypothesis testing?
Hypothesis testing helps you make decisions.
5. What is a t-test?
A t-test compares two groups.
6. What is ANOVA?
ANOVA compares three or more groups.
7. What is correlation?
Correlation measures relationships.
8. What is linear regression?
Linear regression predicts outcomes.
9. What is a p-value?
The probability of seeing the data if the null hypothesis is true.
10. How do Nigerian businesses use statistics?
They use it to make decisions.

πŸ“ Review Questions (15)

  1. What is statistics?
  2. What is descriptive statistics?
  3. What is the difference between population and sample?
  4. What is hypothesis testing?
  5. What is a t-test?
  6. What is ANOVA?
  7. What is correlation?
  8. What is linear regression?
  9. What is a p-value?
  10. How do you calculate the mean?
  11. How do you calculate the median?
  12. How do you calculate the standard deviation?
  13. Give an example of a Nigerian business using statistics.
  14. What is the best practice for statistics?
  15. What is the difference between correlation and causation?

✏️ Fill-in-the-Blank Exercises

  1. __________ is the science of understanding data.
  2. __________ statistics summarise data.
  3. __________ is the whole group; __________ is a part.
  4. A __________ test compares two groups.
  5. __________ measures relationships.

βœ… True or False Exercises

  1. Statistics is not important. (False)
  2. Descriptive statistics summarise data. (True)
  3. Population is a part of the sample. (False)
  4. A t-test compares two groups. (True)
  5. Correlation implies causation. (False)

πŸ”˜ Multiple Choice Questions

  1. What is statistics?
    A) The science of understanding data B) A game C) A type of food D) A sport
    Answer: A
  2. What is descriptive statistics?
    A) Summarises data B) Makes predictions C) Tests hypotheses D) Analyses relationships
    Answer: A
  3. What is the difference between population and sample?
    A) Population is the whole group; sample is a part B) They are the same C) Sample is the whole group D) Population is a part
    Answer: A
  4. What is hypothesis testing?
    A) Helps you make decisions B) Summarises data C) Makes predictions D) Analyses relationships
    Answer: A
  5. What is a t-test?
    A) Compares two groups B) Compares three or more groups C) Measures relationships D) Predicts outcomes
    Answer: A
  6. What is ANOVA?
    A) Compares three or more groups B) Compares two groups C) Measures relationships D) Predicts outcomes
    Answer: A
  7. What is correlation?
    A) Measures relationships B) Compares groups C) Predicts outcomes D) Summarises data
    Answer: A
  8. What is linear regression?
    A) Predicts outcomes B) Measures relationships C) Compares groups D) Summarises data
    Answer: A
  9. What is a p-value?
    A) The probability of seeing the data if H0 is true B) The mean C) The median D) The standard deviation
    Answer: A
  10. Which of these is a Nigerian business using statistics?
    A) Paystack B) Flutterwave C) MTN D) All of the above
    Answer: D
  11. How do you calculate the mean?
    A) sum(x) / length(x) B) median(x) C) sd(x) D) max(x) - min(x)
    Answer: A
  12. How do you calculate the median?
    A) The middle value B) sum(x) / length(x) C) sd(x) D) max(x) - min(x)
    Answer: A
  13. How do you calculate the standard deviation?
    A) sd(x) B) sum(x) / length(x) C) median(x) D) max(x) - min(x)
    Answer: A
  14. What is the best practice for statistics?
    A) Choose the right test B) Ignore outliers C) Use the wrong test D) Don't interpret results
    Answer: A
  15. What is the difference between correlation and causation?
    A) Correlation does not imply causation B) Correlation implies causation C) They are the same D) Causation implies correlation
    Answer: A

πŸ”— Matching Exercises

Match the word on the left with the correct meaning on the right:

Word Meaning
t-test Measures relationships
ANOVA Compares two groups
Correlation Compares three or more groups
Regression Predicts outcomes

Answers: t-test β†’ Compares two groups; ANOVA β†’ Compares three or more groups; Correlation β†’ Measures relationships; Regression β†’ Predicts outcomes.

πŸ“ Short Answer Questions

  1. Explain statistics in your own words.
  2. What is the difference between population and sample?
  3. What is a t-test and when would you use it?
  4. What is correlation and when would you use it?
  5. Give an example of a Nigerian business using statistics.

🎭 Scenario-based Exercises

Scenario 1: Ada wants to know if students who study more get better grades. What statistical test should she use?

Scenario 2: A Lagos supermarket wants to know if there is a relationship between advertising spend and sales. What statistical test should they use?

πŸ‘₯ Group Activity

In groups of 4–5, discuss how you would use statistics to analyse a dataset. What tests would you use? Present your ideas to the class.

πŸ§‘ Individual Activity

Think about a dataset you would like to analyse. Write a short paragraph (about 100 words) about what statistical tests you would use.

πŸ—£οΈ Classroom Discussion Questions

  1. Why is statistics important?
  2. What is your favourite statistical test and why?
  3. How can Nigerian businesses benefit from statistics?
  4. What is the most interesting thing you learned about statistics?
  5. How do you think statistics will change in the future?

πŸ› οΈ Mini Project

Analyse a Dataset

Find a dataset and perform statistical analysis. Calculate descriptive statistics, perform a t-test, and create a linear regression model. Save and share your results with the class.

πŸ“‹ Practical Assignment

Use R to perform statistical analysis on a dataset. Calculate descriptive statistics, perform a t-test, and create a linear regression model. Write a short report (about 150 words) about what you did and what you learned.

πŸ† Challenge Exercise

The Challenge: Imagine you are a data analyst in Lagos. You have sales data for four regions. Perform statistical analysis to see if there is a significant difference in sales between regions. Use ANOVA and interpret the results.

πŸ” Quiz Answers

Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A

Fill-in-the-Blank: 1. Statistics, 2. Descriptive, 3. Population, sample, 4. t, 5. Correlation

True or False: 1. False, 2. True, 3. False, 4. True, 5. False

Matching: t-test β†’ Compares two groups; ANOVA β†’ Compares three or more groups; Correlation β†’ Measures relationships; Regression β†’ Predicts outcomes.

🎯 Key Takeaways

  • Statistics is the science of understanding data.
  • Descriptive statistics summarise data.
  • Population is the whole group; sample is a part.
  • Hypothesis testing helps you make decisions.
  • A t-test compares two groups.
  • ANOVA compares three or more groups.
  • Correlation measures relationships.
  • Linear regression predicts outcomes.
  • Nigerian businesses use statistics to make decisions.

πŸ”œ Preparation for Module Six

In the next module, we will explore Advanced Modeling and Machine Learning. You will learn how to build predictive models and use machine learning algorithms.


πŸŽ‰ Congratulations! You have completed Module Five of the R for Data Analysis course.

πŸ‘ You are now ready to move to Module Six: Advanced Modeling and Machine Learning.

7

Module SIx

Module Six: Advanced Modeling and Machine Learning

πŸ€– Module Six: Advanced Modeling and Machine Learning

Welcome back, young data scientist! Today we learn how to build models and make predictions.

🌟 Module Introduction

In Module Five, you learned how to test hypotheses and understand relationships. Now you are ready for the next step: predicting the future!

Imagine you are a weather forecaster. You look at past weather data to predict tomorrow's weather. That is what machine learning does β€” it uses past data to predict future outcomes.

Machine learning is a type of artificial intelligence that helps computers learn from data. It is like teaching a computer to recognise patterns.

In this module, you will learn about linear regression, logistic regression, decision trees, random forests, and model evaluation. By the end, you will be able to build your own predictive models.

πŸ’‘ Think about it: Have you ever tried to predict something? Machine learning helps you do that with data.

🎯 Learning Objectives

By the time you finish this module, you will be able to:

  • Explain what machine learning is.
  • Understand the difference between supervised and unsupervised learning.
  • Build a linear regression model.
  • Build a logistic regression model.
  • Build a decision tree model.
  • Build a random forest model.
  • Evaluate model performance.
  • Give examples of machine learning in Nigerian businesses.

πŸ“– Warm-up Story: Chidi's Weather Predictions

Chidi is a 12-year-old boy who loves to predict the weather. He noticed that when it was cloudy in the morning, it often rained in the afternoon. He wanted to predict if it would rain.

His uncle, who works with data, said, "Chidi, you are doing machine learning. You are using past data (cloudy mornings) to predict future outcomes (rain)."

Chidi learned how to use R to build predictive models. He could now predict the weather with data. He was like a real data scientist.

🧠 Think about it: Have you ever tried to predict something? Machine learning helps you do that with data.

πŸ“š Main Lessons

1. What is Machine Learning?

Definition: Machine learning is a type of artificial intelligence that allows computers to learn from data.

Why it is important: Machine learning helps you make predictions and decisions.

Simple explanation: Think of machine learning like teaching a computer to recognise patterns.

🏫 School example: A teacher uses past test scores to predict future performance.

🏠 Home example: Your parents use past spending to predict future expenses.

πŸ‡³πŸ‡¬ Nigerian example: A business uses past sales to predict future sales.

  • Supervised learning: Learning from labelled data.
  • Unsupervised learning: Learning from unlabelled data.
  • Reinforcement learning: Learning from rewards and punishments.
    +-------------------+
    |  MACHINE LEARNING |
    +-------------------+
    |  Supervised       |  ← Labelled data
    |  Unsupervised     |  ← Unlabelled data
    |  Reinforcement    |  ← Rewards and punishments
    +-------------------+
    

πŸ“Œ Mini summary: Machine learning is teaching computers to learn from data.

2. Supervised vs. Unsupervised Learning

Definition: Supervised learning uses labelled data. Unsupervised learning uses unlabelled data.

Why it is important: You need to choose the right type for your problem.

Simple explanation: Think of supervised learning like learning with a teacher. Unsupervised learning is like learning by yourself.

  • Supervised: Regression (predict numbers) and classification (predict categories).
  • Unsupervised: Clustering (grouping) and dimensionality reduction.

🏫 School example: Supervised: predicting grades. Unsupervised: grouping students by interests.

🏠 Home example: Supervised: predicting expenses. Unsupervised: grouping expenses by category.

πŸ‡³πŸ‡¬ Nigerian example: Supervised: predicting sales. Unsupervised: grouping customers by behaviour.

πŸ“Œ Mini summary: Supervised learning uses labelled data; unsupervised learning uses unlabelled data.

3. Linear Regression

Definition: Linear regression predicts a continuous outcome.

Why it is important: It is the simplest and most common predictive model.

Simple explanation: Think of linear regression like drawing a line through your data.

  • Simple linear regression: One predictor.
  • Multiple linear regression: Multiple predictors.
  • Formula: y = a + b*x.
# Linear regression model <- lm(Sales ~ Advertising, data = sales_data) summary(model) # Multiple linear regression model <- lm(Sales ~ Advertising + Price, data = sales_data) summary(model)

🏫 School example: Predicting grades from study time.

🏠 Home example: Predicting expenses from income.

πŸ‡³πŸ‡¬ Nigerian example: Predicting sales from advertising spend.

πŸ“Œ Mini summary: Linear regression predicts continuous outcomes.

4. Logistic Regression

Definition: Logistic regression predicts a binary outcome (yes/no).

Why it is important: It is used for classification problems.

Simple explanation: Think of logistic regression like answering yes/no questions.

  • Binary classification: Two outcomes (yes/no, true/false).
  • Formula: log(p/(1-p)) = a + b*x.
  • glm(): Function for logistic regression in R.
# Logistic regression model <- glm(Churn ~ Tenure + MonthlyCharges, data = customer_data, family = "binomial") summary(model)

🏫 School example: Predicting if a student will pass or fail.

🏠 Home example: Predicting if you will overspend this month.

πŸ‡³πŸ‡¬ Nigerian example: Predicting if a customer will churn.

πŸ“Œ Mini summary: Logistic regression predicts binary outcomes.

5. Decision Trees

Definition: A decision tree is a tree-like model that makes decisions based on rules.

Why it is important: Decision trees are easy to understand and interpret.

Simple explanation: Think of a decision tree like a flowchart.

  • Nodes: Questions or decisions.
  • Branches: Answers to questions.
  • Leaves: Final decisions or predictions.
  • rpart(): Function for decision trees in R.
# Decision tree library(rpart) tree <- rpart(Churn ~ Tenure + MonthlyCharges, data = customer_data) plot(tree) text(tree)

🏫 School example: A decision tree to determine if a student passes.

🏠 Home example: A decision tree to decide if you should buy something.

πŸ‡³πŸ‡¬ Nigerian example: A decision tree to determine if a customer will churn.

    +-------------------+
    |  DECISION TREE    |
    +-------------------+
    |  Tenure > 12?     |
    |  /         \      |
    | Yes         No    |
    | /             \   |
    | Churn?        ... |
    +-------------------+
    

πŸ“Œ Mini summary: Decision trees make decisions based on rules.

6. Random Forest

Definition: Random forest is a collection of many decision trees.

Why it is important: Random forests are more accurate than single decision trees.

Simple explanation: Think of random forest like many experts voting on a decision.

  • Ensemble method: Combines multiple models.
  • randomForest(): Function for random forest in R.
  • Bagging: Bootstrap aggregating.
# Random forest library(randomForest) rf <- randomForest(Churn ~ Tenure + MonthlyCharges, data = customer_data) print(rf)

🏫 School example: Predicting student performance with many teachers.

🏠 Home example: Predicting expenses with many family members.

πŸ‡³πŸ‡¬ Nigerian example: Predicting sales with many predictors.

πŸ“Œ Mini summary: Random forest combines many decision trees.

7. Model Evaluation

Definition: Model evaluation is the process of assessing how well your model performs.

Why it is important: You need to know if your model is any good.

Simple explanation: Think of model evaluation like grading your model.

  • Train/test split: Split data into training and testing sets.
  • Accuracy: Percentage of correct predictions.
  • Confusion matrix: Shows true/false positives and negatives.
  • R-squared: How well the model fits the data.
  • RMSE: Root mean squared error.
# Model evaluation library(caret) set.seed(123) train_index <- createDataPartition(sales_data$Sales, p = 0.8, list = FALSE) train_data <- sales_data[train_index, ] test_data <- sales_data[-train_index, ] model <- lm(Sales ~ Advertising, data = train_data) predictions <- predict(model, test_data) RMSE(predictions, test_data$Sales)

🏫 School example: Testing a student's knowledge with an exam.

🏠 Home example: Testing a recipe before serving it.

πŸ‡³πŸ‡¬ Nigerian example: Testing a sales model before using it.

πŸ“Œ Mini summary: Model evaluation assesses how well your model performs.

8. Overfitting and Underfitting

Definition: Overfitting is when a model learns the training data too well. Underfitting is when a model is too simple.

Why it is important: You want a model that generalises well to new data.

Simple explanation: Think of overfitting like memorising the answers. Underfitting is like not studying enough.

  • Overfitting: High accuracy on training, low accuracy on test.
  • Underfitting: Low accuracy on both training and test.
  • Regularisation: Techniques to prevent overfitting.
  • Cross-validation: A technique to evaluate model performance.

🏫 School example: A student who memorises the textbook but can't answer new questions.

🏠 Home example: A recipe that works only with specific ingredients.

πŸ‡³πŸ‡¬ Nigerian example: A sales model that works only for one region.

πŸ“Œ Mini summary: Overfitting is memorising the training data; underfitting is too simple.

9. Cross-Validation

Definition: Cross-validation is a technique to evaluate model performance by splitting data multiple times.

Why it is important: It gives a more reliable estimate of model performance.

Simple explanation: Think of cross-validation like taking multiple tests instead of just one.

  • k-fold cross-validation: Split data into k folds.
  • Leave-one-out cross-validation: Leave one observation out.
  • trainControl(): Function for cross-validation in R.
# Cross-validation library(caret) train_control <- trainControl(method = "cv", number = 5) model <- train(Sales ~ Advertising, data = sales_data, method = "lm", trControl = train_control) print(model)

🏫 School example: Testing a student with multiple quizzes.

🏠 Home example: Testing a recipe with multiple taste tests.

πŸ‡³πŸ‡¬ Nigerian example: Testing a sales model with multiple time periods.

πŸ“Œ Mini summary: Cross-validation gives a reliable estimate of model performance.

10. Machine Learning in Nigerian Businesses

Definition: Nigerian businesses use machine learning to make predictions and decisions.

Why it is important: Machine learning helps Nigerian businesses grow.

Simple explanation: Think of it like a crystal ball for business.

  • Banks: Use machine learning to detect fraud.
  • Telecom: Use machine learning to predict churn.
  • Retail: Use machine learning to predict sales.
  • Government: Use machine learning to predict economic trends.
  • Startups: Use machine learning to understand customers.

πŸ‡³πŸ‡¬ Nigerian example: A Lagos supermarket uses machine learning to predict sales.

πŸ“Œ Mini summary: Nigerian businesses use machine learning to grow.

πŸ“– Key Vocabulary

Word Simple Meaning
Machine Learning Teaching computers to learn from data.
Supervised Learning Learning from labelled data.
Unsupervised Learning Learning from unlabelled data.
Linear Regression Predicts continuous outcomes.
Logistic Regression Predicts binary outcomes.
Decision Tree Makes decisions based on rules.
Random Forest Combines many decision trees.
Overfitting Memorising the training data.
Underfitting Too simple to learn patterns.
Cross-Validation Evaluates model performance.
Accuracy Percentage of correct predictions.
Confusion Matrix Shows true/false positives and negatives.
R-squared How well the model fits the data.
RMSE Root mean squared error.
Ensemble Method Combines multiple models.

🧩 Important Concepts

  • Machine learning teaches computers to learn from data.
  • Supervised learning uses labelled data.
  • Unsupervised learning uses unlabelled data.
  • Linear regression predicts continuous outcomes.
  • Logistic regression predicts binary outcomes.
  • Decision trees make decisions based on rules.
  • Random forest combines many decision trees.
  • Model evaluation assesses performance.
  • Overfitting and underfitting are common problems.
  • Cross-validation gives reliable performance estimates.
  • Nigerian businesses use machine learning to grow.

πŸ“Œ Step-by-Step Explanations

How to build a machine learning model

  1. Define the problem: What do you want to predict?
  2. Collect data: Gather your data.
  3. Clean data: Clean and prepare your data.
  4. Split data: Split into training and test sets.
  5. Choose model: Choose a model (e.g., linear regression).
  6. Train model: Train the model on the training data.
  7. Evaluate model: Evaluate on the test data.
  8. Deploy model: Use the model to make predictions.

How to evaluate a classification model

  1. Make predictions: Use the model to predict on test data.
  2. Create confusion matrix: Compare predictions to actual values.
  3. Calculate accuracy: (TP + TN) / (TP + TN + FP + FN).
  4. Calculate precision: TP / (TP + FP).
  5. Calculate recall: TP / (TP + FN).
  6. Calculate F1 score: 2 * (precision * recall) / (precision + recall).

🌍 Real-life Examples

  • School: A teacher uses machine learning to predict student performance.
  • Hospital: A hospital uses machine learning to predict patient outcomes.
  • Restaurant: A restaurant uses machine learning to predict sales.
  • Shop: A shop uses machine learning to predict inventory needs.

πŸ‡³πŸ‡¬ Nigerian Examples

  • Paystack: Uses machine learning to detect fraud.
  • Flutterwave: Uses machine learning to predict transaction trends.
  • MTN Nigeria: Uses machine learning to predict churn.
  • A Lagos supermarket: Uses machine learning to predict sales.
  • A Nigerian bank: Uses machine learning to detect anomalies.

🎈 Fun Examples for You

  • Your pocket money: Predict how much you will spend.
  • Your game scores: Predict your next score.
  • Your reading log: Predict how many books you will read.
  • Your chores: Predict when you will finish your chores.

🏠 Everyday Examples

  • At home: Your parents use machine learning to predict expenses.
  • At school: Your teacher uses machine learning to predict grades.
  • In your community: Local businesses use machine learning to predict sales.
  • In your own life: You can use machine learning to predict your habits.

πŸ‘©β€πŸ« Teacher Notes

  • Encourage students to think about predictions in their daily lives.
  • Use the warm-up story to spark interest in machine learning.
  • Demonstrate building models in R.
  • Discuss the importance of model evaluation.
  • Ask students to think about how machine learning is used in Nigerian businesses.

πŸ‘ͺ Parent Tips

  • Talk to your child about how you make predictions in your work.
  • Show your child how you use data to make decisions.
  • Open R and explore machine learning together.
  • Encourage your child to think about how they would use machine learning.
  • Share examples of machine learning in Nigerian businesses.

🧠 Interesting Facts

  • Machine learning is used in almost every industry.
  • The first machine learning algorithm was developed in the 1950s.
  • Machine learning can help predict the future.
  • R is one of the most popular languages for machine learning.
  • Nigerian businesses are increasingly using machine learning.

πŸ’‘ Did You Know?

  • Did you know that machine learning can detect fraud?
  • Did you know that decision trees are used in medicine?
  • Did you know that random forests are very accurate?
  • Did you know that cross-validation prevents overfitting?
  • Did you know that Nigerian banks use machine learning?

πŸ”” Remember This

  • Machine learning teaches computers to learn from data.
  • Supervised learning uses labelled data.
  • Unsupervised learning uses unlabelled data.
  • Linear regression predicts continuous outcomes.
  • Logistic regression predicts binary outcomes.
  • Decision trees make decisions based on rules.
  • Random forest combines many decision trees.
  • Model evaluation assesses performance.
  • Overfitting and underfitting are common problems.
  • Cross-validation gives reliable performance estimates.

⚠️ Common Mistakes

  • Not splitting data: Always split into training and test.
  • Overfitting: Too complex model.
  • Underfitting: Too simple model.
  • Not evaluating: Always evaluate your model.
  • Ignoring assumptions: Check assumptions of your model.

✨ Best Practices for Machine Learning

  • Split data into training and test sets.
  • Use cross-validation for reliable estimates.
  • Avoid overfitting.
  • Avoid underfitting.
  • Evaluate your model with appropriate metrics.
  • Check assumptions of your model.
  • Interpret your results carefully.
  • Keep learning and exploring new models.

πŸ“Š Clear Illustrations

1. Machine Learning Types

    +-------------------+
    |  MACHINE LEARNING |
    +-------------------+
    |  Supervised       |  ← Labelled data
    |  Unsupervised     |  ← Unlabelled data
    |  Reinforcement    |  ← Rewards and punishments
    +-------------------+
    

2. Linear Regression

    +-------------------+
    |  DATA POINTS      |  ← *  *  *
    +-------------------+
           |
           V
    +-------------------+
    |  REGRESSION LINE  |  ← y = 2x + 3
    +-------------------+
           |
           V
    +-------------------+
    |  PREDICT          |  ← y = 2(5) + 3 = 13
    +-------------------+
    

3. Decision Tree

    +-------------------+
    |  Tenure > 12?     |
    |  /         \      |
    | Yes         No    |
    | /             \   |
    | Churn?        ... |
    +-------------------+
    

4. Random Forest

    +-------------------+
    |  Tree 1           |
    +-------------------+
    |  Tree 2           |
    +-------------------+
    |  Tree 3           |
    +-------------------+
           |
           V
    +-------------------+
    |  Vote             |  ← Combine predictions
    +-------------------+
    

πŸ“ Lesson Summaries

Lesson 1: Machine learning teaches computers to learn from data.

Lesson 2: Supervised learning uses labelled data; unsupervised uses unlabelled.

Lesson 3: Linear regression predicts continuous outcomes.

Lesson 4: Logistic regression predicts binary outcomes.

Lesson 5: Decision trees make decisions based on rules.

Lesson 6: Random forest combines many decision trees.

Lesson 7: Model evaluation assesses performance.

Lesson 8: Overfitting is memorising; underfitting is too simple.

Lesson 9: Cross-validation gives reliable performance estimates.

Lesson 10: Nigerian businesses use machine learning to grow.

πŸ“˜ End-of-Module Summary

In this module, you learned about advanced modeling and machine learning. You discovered supervised and unsupervised learning, linear and logistic regression, decision trees, random forest, and model evaluation. You also learned about overfitting, underfitting, and cross-validation.

🎯 You can now:

  • Explain what machine learning is.
  • Understand the difference between supervised and unsupervised learning.
  • Build a linear regression model.
  • Build a logistic regression model.
  • Build a decision tree model.
  • Build a random forest model.
  • Evaluate model performance.
  • Give examples of machine learning in Nigerian businesses.

❓ Frequently Asked Questions

1. What is machine learning?
Machine learning is teaching computers to learn from data.
2. What is supervised learning?
Supervised learning uses labelled data.
3. What is unsupervised learning?
Unsupervised learning uses unlabelled data.
4. What is linear regression?
Linear regression predicts continuous outcomes.
5. What is logistic regression?
Logistic regression predicts binary outcomes.
6. What is a decision tree?
A decision tree makes decisions based on rules.
7. What is random forest?
Random forest combines many decision trees.
8. What is overfitting?
Overfitting is memorising the training data.
9. What is cross-validation?
Cross-validation evaluates model performance.
10. How do Nigerian businesses use machine learning?
They use it to make predictions and decisions.

πŸ“ Review Questions (15)

  1. What is machine learning?
  2. What is the difference between supervised and unsupervised learning?
  3. What is linear regression?
  4. What is logistic regression?
  5. What is a decision tree?
  6. What is random forest?
  7. What is overfitting?
  8. What is underfitting?
  9. What is cross-validation?
  10. How do you evaluate a model?
  11. What is accuracy?
  12. What is a confusion matrix?
  13. Give an example of a Nigerian business using machine learning.
  14. What is the best practice for machine learning?
  15. What is the difference between classification and regression?

✏️ Fill-in-the-Blank Exercises

  1. __________ learning uses labelled data.
  2. __________ regression predicts continuous outcomes.
  3. __________ regression predicts binary outcomes.
  4. A __________ tree makes decisions based on rules.
  5. __________ forest combines many decision trees.

βœ… True or False Exercises

  1. Machine learning is only for experts. (False)
  2. Supervised learning uses labelled data. (True)
  3. Linear regression predicts binary outcomes. (False)
  4. Decision trees are easy to interpret. (True)
  5. Overfitting is a good thing. (False)

πŸ”˜ Multiple Choice Questions

  1. What is machine learning?
    A) Teaching computers to learn from data B) A game C) A type of food D) A sport
    Answer: A
  2. What is supervised learning?
    A) Learning from labelled data B) Learning from unlabelled data C) Learning from rewards D) Learning from punishments
    Answer: A
  3. What is linear regression?
    A) Predicts continuous outcomes B) Predicts binary outcomes C) Makes decisions D) Combines models
    Answer: A
  4. What is logistic regression?
    A) Predicts binary outcomes B) Predicts continuous outcomes C) Makes decisions D) Combines models
    Answer: A
  5. What is a decision tree?
    A) Makes decisions based on rules B) Predicts continuous outcomes C) Predicts binary outcomes D) Combines models
    Answer: A
  6. What is random forest?
    A) Combines many decision trees B) Makes decisions based on rules C) Predicts continuous outcomes D) Predicts binary outcomes
    Answer: A
  7. What is overfitting?
    A) Memorising the training data B) Too simple C) Good performance D) Bad performance
    Answer: A
  8. What is underfitting?
    A) Too simple B) Memorising the training data C) Good performance D) Bad performance
    Answer: A
  9. What is cross-validation?
    A) Evaluates model performance B) Trains a model C) Makes predictions D) Combines models
    Answer: A
  10. Which of these is a Nigerian business using machine learning?
    A) Paystack B) Flutterwave C) MTN D) All of the above
    Answer: D
  11. How do you split data?
    A) Train/test split B) Only train C) Only test D) No split
    Answer: A
  12. What is accuracy?
    A) Percentage of correct predictions B) Percentage of wrong predictions C) Model complexity D) Model simplicity
    Answer: A
  13. What is a confusion matrix?
    A) Shows true/false positives and negatives B) Shows accuracy C) Shows R-squared D) Shows RMSE
    Answer: A
  14. What is the best practice for machine learning?
    A) Split data, evaluate, avoid overfitting B) Don't split data C) Use all data for training D) Ignore evaluation
    Answer: A
  15. What is the difference between classification and regression?
    A) Classification predicts categories; regression predicts numbers B) Classification predicts numbers; regression predicts categories C) They are the same D) Classification is easier
    Answer: A

πŸ”— Matching Exercises

Match the word on the left with the correct meaning on the right:

Word Meaning
Regression Predicts categories
Classification Predicts numbers
Decision Tree Combines many trees
Random Forest Makes decisions based on rules

Answers: Regression β†’ Predicts numbers; Classification β†’ Predicts categories; Decision Tree β†’ Makes decisions based on rules; Random Forest β†’ Combines many trees.

πŸ“ Short Answer Questions

  1. Explain machine learning in your own words.
  2. What is the difference between supervised and unsupervised learning?
  3. What is the difference between regression and classification?
  4. What is overfitting and why is it a problem?
  5. Give an example of a Nigerian business using machine learning.

🎭 Scenario-based Exercises

Scenario 1: Chidi wants to predict if it will rain tomorrow based on temperature and humidity. What type of model should he use?

Scenario 2: A Lagos supermarket wants to predict sales for the next month. What type of model should they use?

πŸ‘₯ Group Activity

In groups of 4–5, discuss how you would use machine learning to solve a problem. What model would you use? Present your ideas to the class.

πŸ§‘ Individual Activity

Think about a problem you could solve with machine learning. Write a short paragraph (about 100 words) about what model you would use.

πŸ—£οΈ Classroom Discussion Questions

  1. Why is machine learning important?
  2. What is your favourite machine learning model and why?
  3. How can Nigerian businesses benefit from machine learning?
  4. What is the most interesting thing you learned about machine learning?
  5. How do you think machine learning will change in the future?

πŸ› οΈ Mini Project

Build a Machine Learning Model

Find a dataset and build a machine learning model. Split the data, train the model, and evaluate its performance. Save and share your results with the class.

πŸ“‹ Practical Assignment

Use R to build a machine learning model on a dataset. Split the data, train the model, and evaluate its performance. Write a short report (about 150 words) about what you did and what you learned.

πŸ† Challenge Exercise

The Challenge: Imagine you are a data scientist in Lagos. You have customer data and want to predict churn. Build a machine learning model to predict churn. Use logistic regression, decision tree, and random forest. Compare their performance.

πŸ” Quiz Answers

Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A

Fill-in-the-Blank: 1. Supervised, 2. Linear, 3. Logistic, 4. decision, 5. Random

True or False: 1. False, 2. True, 3. False, 4. True, 5. False

Matching: Regression β†’ Predicts numbers; Classification β†’ Predicts categories; Decision Tree β†’ Makes decisions based on rules; Random Forest β†’ Combines many trees.

🎯 Key Takeaways

  • Machine learning teaches computers to learn from data.
  • Supervised learning uses labelled data.
  • Unsupervised learning uses unlabelled data.
  • Linear regression predicts continuous outcomes.
  • Logistic regression predicts binary outcomes.
  • Decision trees make decisions based on rules.
  • Random forest combines many decision trees.
  • Model evaluation assesses performance.
  • Overfitting and underfitting are common problems.
  • Cross-validation gives reliable performance estimates.
  • Nigerian businesses use machine learning to grow.

πŸ”œ Preparation for Module Seven

In the next module, we will explore Reporting with R Markdown and Shiny. You will learn how to create dynamic reports and interactive dashboards.


πŸŽ‰ Congratulations! You have completed Module Six of the R for Data Analysis course.

πŸ‘ You are now ready to move to Module Seven: Reporting with R Markdown and Shiny.

8

Module Seven

Module 7 Β· R for Data Analysis Β· Clean HTML

Module 7 Β· Telling Stories with Data

Hello, young data explorer! In this module, we are going to learn how to tell stories with data. Numbers and tables are useful, but sometimes they are hard to understand. That is why we use pictures – also called graphs or charts – to show our data in a fun and clear way.

Imagine you have a big bag of colourful sweets. If I just tell you the numbers, you might forget. But if I draw a picture with coloured bars, you will see which colour is the most popular right away! That is what this module is about: turning numbers into pictures so everyone can understand them easily.

We will use the R programming language to create these pictures. R has special tools – called packages – that help us draw beautiful graphs. We will learn the most important ones: ggplot2 and plotly.

By the end of this module, you will be able to look at a table of numbers, choose the best graph, and draw it using R. You will be a data storyteller!


Learning Objectives

  • Explain why we use graphs and charts to show data.
  • Name at least three types of graphs (bar chart, line chart, pie chart).
  • Understand what the ggplot2 package does.
  • Create a basic bar chart using ggplot2.
  • Create a basic line chart using ggplot2.
  • Know how to add titles and labels to make graphs clear.
  • Learn about interactive graphs using plotly.
  • Choose the right graph for different kinds of data.
  • Recognise common mistakes when making graphs.
  • Practice making graphs from real Nigerian data.

Warm-up Story: The Great School Election

At Sunshine Primary School, the pupils voted for their favourite school lunch. The choices were: Jollof Rice, Fried Plantain, Pounded Yam, and Beans. The head teacher, Mrs. Ade, received the votes as numbers:

  • Jollof Rice – 45 votes
  • Fried Plantain – 30 votes
  • Pounded Yam – 20 votes
  • Beans – 5 votes

She wrote these numbers on the board. The children could not quickly see which food was the winner. It was confusing! Then her assistant, Mr. Bello, said, β€œLet me draw a picture.” He made a bar chart with colourful bars. Instantly, everyone saw that Jollof Rice was the most loved food. The children cheered!

That is the power of data visualisation. It turns boring numbers into a story that everyone can understand in seconds. In this module, you will learn how to create these magical pictures using R.


Main Lessons

Lesson 1: What is Data Visualisation?

Definition: Data visualisation means making pictures (graphs, charts, maps) from data (numbers and facts).

Why it is important: Our brains understand pictures faster than numbers. A picture can show patterns, trends, and outliers (things that don't fit) very quickly.

Simple explanation: Think of your data as a story. The graph is the illustration that makes the story exciting.

Real-life example: Weather forecasters use maps with colours to show temperature – red for hot, blue for cold.

School example: Teachers draw bar charts to show how many pupils got A, B, C in a test.

Home example: You can draw a picture of how you spend your day: school, play, homework, sleep.

Nigerian example: The National Bureau of Statistics uses graphs to show Nigeria's population growth.

Illustration:

   DATA (numbers)   -->   VISUALISATION (pictures)   -->   UNDERSTANDING
      (45, 30, 20, 5)        (bar chart with bars)        (Jollof wins!)

Mini summary: Visualisation turns numbers into pictures so we can see what the data is telling us.


Lesson 2: The Grammar of Graphics – Introducing ggplot2

Definition: ggplot2 is a special R package that helps us create beautiful graphs. It is based on the "grammar of graphics" – a set of rules for building graphs layer by layer.

Why it is important: ggplot2 is the most popular drawing tool in R. It is flexible and powerful.

Simple explanation: Imagine you are building a house with LEGO blocks. ggplot2 gives you blocks (layers) to build your graph: data, axes, bars, colours, titles.

Real-life example: An artist uses layers of paint to make a painting. ggplot2 uses layers to make graphs.

School example: In art class, you draw a background, then add a tree, then add birds. ggplot2 adds one layer at a time.

Home example: When you make a sandwich, you add bread, then cheese, then ham, then bread. Layers!

Nigerian example: A Nigerian data analyst uses ggplot2 to show the yield of cassava in different states.

Illustration:

   DATA  +  AESTHETICS (x, y)  +  GEOM (bars, lines)  +  THEME  =  GRAPH
   (layer 1)     (layer 2)           (layer 3)         (layer 4)

Mini summary: ggplot2 builds graphs in layers like you build a LEGO castle.


Lesson 3: Installing and Loading ggplot2

Definition: To use ggplot2, we first need to install it (only once) and then load it (every time we start R).

Why it is important: You cannot use a tool until you bring it to your workspace.

Simple explanation: It is like buying a new board game (install) and then taking it out of the box to play (load).

Real-life example: You install a game app on your phone, then you open it to play.

School example: Your teacher gives you a textbook (install) and you open it to read (load).

Home example: Mum buys flour (install) and then opens the bag to bake (load).

Nigerian example: A student in Lagos installs R packages once, then loads them for each project.

Code:

   # Install (do this once)
   install.packages("ggplot2")

   # Load (do this every time)
   library(ggplot2)

Mini summary: Install once, load every session.


Lesson 4: Your First Bar Chart with ggplot2

Definition: A bar chart uses bars to show counts or values for different categories.

Why it is important: It is the easiest way to compare categories.

Simple explanation: The taller the bar, the bigger the number.

Real-life example: A supermarket uses a bar chart to show which fruit sells most.

School example: A bar chart of favourite sports: football, basketball, tennis.

Home example: A bar chart of how many hours you read each day.

Nigerian example: A bar chart showing the number of pupils in each class in a school in Abuja.

Code & Illustration:

   # Sample data: favourite foods
   foods <- data.frame(
       Food = c("Jollof", "Plantain", "Yam", "Beans"),
       Votes = c(45, 30, 20, 5)
   )

   # ggplot bar chart
   ggplot(foods, aes(x = Food, y = Votes)) +
       geom_bar(stat = "identity") +
       labs(title = "Favourite School Lunch", x = "Food", y = "Votes")
   Bar chart (concept):
   Votes
   50 |    β– 
   40 |    β– 
   30 |    β–    β– 
   20 |    β–    β–    β– 
   10 |    β–    β–    β–    β– 
    0 |____β– ___β– ___β– ___β– ____
         J   P   Y   B

Mini summary: geom_bar() draws the bars. Use stat="identity" when you have values.


Lesson 5: Line Charts – Showing Trends Over Time

Definition: A line chart connects data points with lines to show how something changes over time.

Why it is important: It helps us see trends: going up (increase), going down (decrease), or staying flat.

Simple explanation: It is like connecting the dots in a dot-to-dot puzzle, but the dots are data points.

Real-life example: A line chart of temperature every day of the week.

School example: A line chart of your test scores over the term.

Home example: A line chart of your piggy bank savings each month.

Nigerian example: A line chart showing the price of tomatoes in Lagos market over the year.

Code:

   # Sample data: monthly rainfall (mm)
   rain <- data.frame(
       Month = c("Jan", "Feb", "Mar", "Apr"),
       Rainfall = c(15, 20, 35, 40)
   )

   ggplot(rain, aes(x = Month, y = Rainfall, group = 1)) +
       geom_line() +
       geom_point() +    # adds dots
       labs(title = "Monthly Rainfall", x = "Month", y = "Rainfall (mm)")
   Line chart concept:
   Rainfall
   40 |        ●
   35 |      ●
   30 |
   25 |
   20 |   ●
   15 | ●
    0 |___●___●___●___●____
        J   F   M   A

Mini summary: Use geom_line() for trends, add geom_point() to show data points.


Lesson 6: Pie Charts – Showing Parts of a Whole

Definition: A pie chart is a circle divided into slices. Each slice represents a part of the total (100%).

Why it is important: It shows proportions (what fraction of the whole each category is).

Simple explanation: Imagine a pizza cut into slices. Bigger slice = bigger share.

Real-life example: A pie chart of how you spend your 24 hours: sleep, school, play, chores.

School example: A pie chart showing the percentage of pupils who like each subject.

Home example: A pie chart of your weekly allowance spending.

Nigerian example: A pie chart of Nigeria's exports: oil, agriculture, etc.

Code (using ggplot2 with coord_polar):

   # We use geom_bar() + coord_polar()
   foods <- data.frame(
       Food = c("Jollof", "Plantain", "Yam", "Beans"),
       Votes = c(45, 30, 20, 5)
   )

   ggplot(foods, aes(x = "", y = Votes, fill = Food)) +
       geom_bar(stat = "identity", width = 1) +
       coord_polar("y", start = 0) +
       labs(title = "Pie Chart of Favourite Foods")
   Pie chart concept:
        ________
       / Jollof \
      |  45%     |
      |   (big)  |
       \________/
       / Plantain \
      |  30%      |
       \________/
      / Yam 20% \
     / Beans 5%  \

Mini summary: Pie charts show parts of a whole. Use coord_polar() to turn a bar chart into a pie chart.


Lesson 7: Adding Titles and Labels

Definition: Titles and labels are words that tell the reader what the graph is about.

Why it is important: Without labels, your graph is just a picture. Labels make it meaningful.

Simple explanation: Labels are like name tags for your graph – they tell what each part means.

Real-life example: A map without place names is useless. Labels tell you the names of cities.

School example: A graph of test scores needs a title "Test Scores" and labels for subjects.

Home example: A bar chart of chores needs labels: "Wash dishes", "Sweep", etc.

Nigerian example: A graph of state populations must label each state.

Code (using labs()):

   ggplot(foods, aes(x = Food, y = Votes)) +
       geom_bar(stat = "identity") +
       labs(
           title = "Favourite School Lunch",
           subtitle = "Sunshine Primary School Election",
           x = "Food Items",
           y = "Number of Votes",
           caption = "Data from the school election"
       )
   [Title] Favourite School Lunch
   [Subtitle] Sunshine Primary School Election
   Votes |
   45   |   [bar]
   30   |   [bar] [bar]
   20   |   [bar] [bar] [bar]
    ...
   [x-axis label] Food Items
   [y-axis label] Number of Votes
   [caption] Data from the school election

Mini summary: Always use labs() to add title, subtitle, x, y, and caption.


Lesson 8: Colours and Fill

Definition: Colours make your graph attractive and help distinguish different groups.

Why it is important: Colours catch the eye and make comparisons easier.

Simple explanation: Colour is like using different crayons to colour different parts of your drawing.

Real-life example: Traffic lights use red, yellow, green to convey meaning.

School example: In a bar chart, you colour each bar differently to show different classes.

Home example: You use a red marker for important events on your calendar.

Nigerian example: Using green-white-green for Nigerian data highlights.

Code (using fill and scale_fill_manual):

   ggplot(foods, aes(x = Food, y = Votes, fill = Food)) +
       geom_bar(stat = "identity") +
       scale_fill_manual(values = c("red", "yellow", "brown", "green"))
   [Red bar for Jollof]  [Yellow bar for Plantain]  [Brown bar for Yam]  [Green bar for Beans]

Mini summary: Use fill inside aes() to colour by category, and scale_fill_manual() to choose colours.


Lesson 9: Themes – Making Graphs Look Good

Definition: A theme controls the non-data elements of a graph: background, grid lines, font size, etc.

Why it is important: A clean, professional look helps people focus on the data.

Simple explanation: Think of themes as different outfits for your graph – you can dress it up or keep it simple.

Real-life example: A formal report uses a clean theme; a fun poster uses a colourful theme.

School example: Your teacher might prefer a plain background for clarity.

Home example: You can choose a "cartoon" theme for a fun family graph.

Nigerian example: Use theme_minimal() for a clean, professional look in business reports.

Code:

   ggplot(foods, aes(x = Food, y = Votes, fill = Food)) +
       geom_bar(stat = "identity") +
       theme_minimal()   # or theme_classic(), theme_bw()
   [Graph with a clean white background, no gridlines, simple fonts]

Mini summary: Themes change the appearance. theme_minimal() is a good default.


Lesson 10: Saving Your Graphs

Definition: Saving a graph means exporting it as an image file (PNG, JPEG, PDF) so you can use it in reports or presentations.

Why it is important: You often need to share your graphs with others who don't use R.

Simple explanation: It is like taking a screenshot of your graph, but much better quality.

Real-life example: You save a photo to show your friends.

School example: You save your graph to paste into your project document.

Home example: You save a graph of your savings to show your family.

Nigerian example: A researcher saves a graph to include in a report for the government.

Code:

   # Create a graph and save it
   p <- ggplot(foods, aes(x = Food, y = Votes)) + geom_bar(stat = "identity")
   ggsave("favourite_food.png", plot = p, width = 6, height = 4)

Mini summary: Use ggsave() to save your graph as an image.


Lesson 11: Interactive Graphs with Plotly

Definition: plotly is an R package that creates interactive graphs – you can hover, zoom, and click.

Why it is important: Interactive graphs let users explore data themselves.

Simple explanation: It is like a living graph that responds to your mouse.

Real-life example: On a weather website, you can hover over a map to see temperatures.

School example: An interactive graph in a science project where you can click on bars to see details.

Home example: A graph of family expenses that shows details when you hover.

Nigerian example: An interactive dashboard showing COVID-19 cases in Nigeria.

Code:

   install.packages("plotly")
   library(plotly)

   p <- ggplot(foods, aes(x = Food, y = Votes, fill = Food)) +
        geom_bar(stat = "identity")
   ggplotly(p)   # makes it interactive!
   [Interactive graph: hover over a bar to see "Votes: 45"]

Mini summary: ggplotly() turns your static graph into an interactive one.


Lesson 12: Choosing the Right Graph

Definition: Not every graph works for every type of data. You need to choose wisely.

Why it is important: The wrong graph can confuse people; the right graph makes the story clear.

Simple explanation: It is like choosing the right shoe for the occasion – sandals for the beach, boots for rain.

Real-life example: Use a line chart for trends over time, bar chart for comparing categories.

School example: Bar chart for favourite subjects; line chart for temperature change.

Home example: Pie chart for budget; bar chart for chores done.

Nigerian example: Bar chart for state populations; line chart for GDP growth.

Decision table:

PurposeBest Graph
Compare categoriesBar chart
Show trends over timeLine chart
Show parts of a wholePie chart
Show distributionHistogram
Show relationshipScatter plot

Mini summary: Match the graph to your question: compare? use bar. trend? use line. part of whole? use pie.


Lesson 13: Common Mistakes and How to Fix Them

Definition: Mistakes are things we do wrong that make our graph hard to understand.

Why it is important: Avoiding mistakes makes our graphs clear and trustworthy.

Simple explanation: It is like baking: if you add too much salt, the cake tastes bad. With graphs, small errors can confuse.

Real-life example: Forgetting to label the y-axis means no one knows what the numbers mean.

School example: Using too many colours makes the graph look messy.

Home example: Making a pie chart with too many slices – hard to read.

Nigerian example: Not including the source of data can make the graph unreliable.

List of common mistakes:

  • Missing labels: Always add x, y, title.
  • Too many colours: Use a simple colour palette.
  • 3D effects: They distort the data; avoid them.
  • Wrong graph type: Don't use a pie chart for trends.
  • Forgetting to save: Use ggsave().

Mini summary: Check labels, colours, graph type – keep it simple and honest.


Lesson 14: Practice with Real Nigerian Data

Let's use data about Nigeria's states and population (example). We will create a bar chart.

   # Sample data (population in millions)
   states <- data.frame(
       State = c("Lagos", "Kano", "Oyo", "Rivers"),
       Population = c(20, 15, 8, 7)
   )

   ggplot(states, aes(x = State, y = Population, fill = State)) +
       geom_bar(stat = "identity") +
       labs(title = "Population of Selected Nigerian States",
            x = "State", y = "Population (millions)") +
       theme_minimal()
   [Bar chart with Lagos tallest, then Kano, Oyo, Rivers]

Mini summary: You can apply the same R code to any data, including Nigerian data.


Lesson 15: Summary of Graphs

We have learned about bar charts, line charts, pie charts, colours, labels, themes, saving, and interactive graphs.

Remember the golden rule: Keep it clear, keep it simple, and always label your axes!


Key Vocabulary

Data Visualisation
Making pictures from data to help us understand.
Bar Chart
A graph with rectangular bars that show numbers.
Line Chart
A graph that connects points with lines to show change over time.
Pie Chart
A circle divided into slices to show parts of a whole.
ggplot2
An R package for drawing graphs layer by layer.
Plotly
An R package for interactive graphs.
Theme
The style or look of a graph (background, colours, font).
Label
Words that explain parts of a graph (title, axis names).

Important Concepts

  • Layering: In ggplot2, you add layers to build a graph.
  • Aesthetics (aes): Mapping data to visual elements like x, y, colour.
  • Geoms: The geometric shapes (bars, lines, points).
  • Scales: Control how data is mapped to colours, sizes, etc.
  • Faceting: Split your graph into small multiples (we will explore later).

Step-by-Step Explanations

How to create a bar chart in R:

  1. Prepare your data: Make a data frame with categories and values.
  2. Load ggplot2: library(ggplot2)
  3. Start the plot: ggplot(data, aes(x, y))
  4. Add bars: + geom_bar(stat = "identity")
  5. Add labels: + labs(title = "...", x = "...", y = "...")
  6. Add theme: + theme_minimal()
  7. Save: ggsave("filename.png")

Real-life Examples

  • A doctor uses a line chart to track a patient's temperature over days.
  • A shop owner uses a bar chart to compare sales of different products.
  • A teacher uses a pie chart to show how pupils spend their time.

Nigerian Examples

  • Nigerian farmers can use bar charts to compare yam output in different states.
  • A line chart can show the price of petrol over the last five years.
  • A pie chart can show Nigeria's revenue from oil, taxes, and other sources.

Fun Examples Children Relate To

  • Bar chart of favourite cartoons: SpongeBob, Paw Patrol, Peppa Pig.
  • Line chart of your height as you grow up.
  • Pie chart of your day: school, play, eating, sleeping.

Everyday Examples

  • Weather forecast on TV uses maps and graphs.
  • Fitness apps show your steps in a line chart.
  • Restaurant menus might show popularity using bars.

Teacher Notes

  • Encourage pupils to think about what story their data tells.
  • Use real data from the school (e.g., class attendance) for practice.
  • Emphasise that a graph without labels is incomplete.
  • Allow pupils to experiment with colours and themes.

Parent Tips

  • Help your child collect simple data at home (e.g., favourite meals).
  • Ask them to draw graphs on paper first, then in R.
  • Discuss what the graph shows – what is the "story"?
  • Encourage curiosity: "What would happen if we changed the data?"

Interesting Facts

  • The first bar chart was created in the 1700s by William Playfair.
  • Pie charts got their name because they look like a baked pie.
  • R is used by data scientists all over the world, including in Nigeria.
  • There are over 10,000 packages available for R!

Did You Know?

  • You can make graphs move like animations using gganimate.
  • Some graphs can be 3D, but they are often hard to read.
  • There is a whole science of making graphs beautiful – it's called "data art".

Remember This

  • Always label your axes and title.
  • Keep your graph simple and clean.
  • Choose the right graph for your data.
  • Save your graphs using ggsave().

Common Mistakes

  • Forgetting to load the package (library(ggplot2)).
  • Using geom_bar() without stat="identity" when you have pre-summarised data.
  • Putting the wrong variable on x and y axes.
  • Overcrowding the graph with too much information.

Best Practices

  • Start with a question: "What do I want to show?"
  • Choose a graph that answers that question.
  • Use consistent colours.
  • Add a source note if the data is from somewhere.
  • Test your graph on a friend – do they understand it?

ASCII Illustrations

Bar chart concept

   Votes
   50 |    β– 
   40 |    β– 
   30 |    β–    β– 
   20 |    β–    β–    β– 
   10 |    β–    β–    β–    β– 
    0 |____β– ___β– ___β– ___β– ____
         J   P   Y   B

Line chart concept

   40 |        ●
   35 |      ●
   30 |
   25 |
   20 |   ●
   15 | ●
    0 |___●___●___●___●____
        J   F   M   A

Pie chart concept

        ________
       / Jollof \
      |  45%     |
       \________/
       / Plantain \
      |  30%      |
       \________/
      / Yam 20% \
     / Beans 5%  \

Comparison Tables

Graph Types and Their Uses
Graph TypeBest Used ForExample
Bar ChartComparing categoriesFavourite foods
Line ChartShowing trends over timeTemperature changes
Pie ChartShowing parts of a wholeBudget allocation

End-of-Module Summary

In this module, we learned that data visualisation is the art of turning numbers into pictures. We discovered the ggplot2 package, which builds graphs layer by layer. We created bar charts to compare categories, line charts to show trends, and pie charts to show parts of a whole. We added titles, labels, and colours to make our graphs clear and beautiful. We also learned to save our graphs and make them interactive with plotly. Remember: a good graph tells a story without confusion. Keep practising with data from your school, home, and country!


Frequently Asked Questions

1. What is ggplot2?
A package in R for creating beautiful graphs.
2. How do I install ggplot2?
Use install.packages("ggplot2").
3. What is a bar chart?
A graph with bars that show counts or values.
4. What is a line chart?
A graph that connects points with lines to show change over time.
5. When should I use a pie chart?
When you want to show parts of a whole (percentages).
6. How do I add a title to my graph?
Use labs(title = "My Title").
7. Can I change the colours of my bars?
Yes, use scale_fill_manual().
8. What is a theme in ggplot2?
The style of the graph (background, grid, font).
9. How do I make my graph interactive?
Use the plotly package and ggplotly().
10. Why are graphs important?
Because they help us see patterns and stories in data quickly.

Review Questions

  1. What is data visualisation?
  2. Name three types of graphs.
  3. What package do we use to create graphs in R?
  4. What does geom_bar() do?
  5. What does geom_line() do?
  6. How do you add a title to a ggplot graph?
  7. What is a pie chart used for?
  8. What is the difference between a bar chart and a line chart?
  9. How do you save a graph in R?
  10. What does the plotly package do?
  11. Why do we add labels to axes?
  12. What is a theme in ggplot2?
  13. Give a Nigerian example of where a bar chart could be used.
  14. What is the grammar of graphics?
  15. Why is it important to choose the right graph?

Fill-in-the-Blank Exercises

  1. We use ______ to turn numbers into pictures. (visualisation)
  2. ______ is a package for creating graphs in R. (ggplot2)
  3. A ______ chart uses bars to show data. (bar)
  4. A ______ chart connects points with lines. (line)
  5. A ______ chart shows parts of a whole. (pie)
  6. Use ______ to add titles and labels. (labs)
  7. ______ makes graphs interactive. (plotly)
  8. ______ saves your graph as an image. (ggsave)
  9. The ______ of a graph controls its style. (theme)
  10. Colours are added using the ______ aesthetic. (fill)

True or False Exercises

  1. Data visualisation is only for adults. (False)
  2. ggplot2 is a package in R. (True)
  3. Pie charts are best for showing trends. (False)
  4. Line charts show change over time. (True)
  5. You don't need to label your axes. (False)
  6. Colours make graphs easier to read. (True)
  7. You can save graphs with ggsave(). (True)
  8. Plotly makes static graphs. (False)
  9. Themes change the data in the graph. (False)
  10. Bar charts compare categories. (True)

Multiple Choice Questions

  1. Which package is used for creating graphs in R?
    A. dplyr B. ggplot2 C. tidyr D. readr
    Answer: B
  2. Which chart is best for comparing categories?
    A. Pie chart B. Line chart C. Bar chart D. Scatter plot
    Answer: C
  3. Which chart is best for showing trends over time?
    A. Pie chart B. Line chart C. Bar chart D. Histogram
    Answer: B
  4. What does geom_bar(stat = "identity") do?
    A. Draws a line B. Draws bars with pre-summarised data C. Draws points D. Draws a pie
    Answer: B
  5. How do you add a title to a ggplot?
    A. title() B. ggtitle() C. labs(title = "...") D. add_title()
    Answer: C
  6. Which function saves a graph?
    A. save() B. ggsave() C. export() D. write()
    Answer: B
  7. Which package makes graphs interactive?
    A. ggplot2 B. dplyr C. plotly D. shiny
    Answer: C
  8. What is a theme in ggplot2?
    A. The data B. The colours C. The style of the graph D. The title
    Answer: C
  9. Which aesthetic is used to colour bars?
    A. colour B. fill C. size D. shape
    Answer: B
  10. What does the x-axis typically represent in a bar chart?
    A. Values B. Categories C. Time D. Percentages
    Answer: B
  11. Which graph is best for showing parts of a whole?
    A. Bar chart B. Line chart C. Pie chart D. Histogram
    Answer: C
  12. What is the grammar of graphics?
    A. A set of rules for building graphs B. A type of chart C. A package D. A dataset
    Answer: A
  13. Which function converts a ggplot to interactive?
    A. ggplotly() B. interactive() C. plotly() D. as.plotly()
    Answer: A
  14. Why should we avoid 3D effects in graphs?
    A. They are ugly B. They distort the data C. They are slow D. They are not allowed
    Answer: B
  15. What is the first thing to consider when making a graph?
    A. Colours B. Title C. The question you want to answer D. Font size
    Answer: C

Matching Exercises

TermDefinition
1. Bar chartA. Shows parts of a whole
2. Line chartB. Compares categories with bars
3. Pie chartC. Shows trends over time
4. ggplot2D. Interactive graphs package
5. plotlyE. R package for layering graphs

Answers: 1-B, 2-C, 3-A, 4-E, 5-D


Short Answer Questions

  1. What is the main purpose of data visualisation?
  2. Explain the difference between a bar chart and a line chart.
  3. How do you install ggplot2?
  4. Why are labels important on graphs?
  5. What does ggsave() do?

Scenario-based Exercises

  1. Scenario: You have data on the number of books read by pupils in different classes. Which graph would you use and why?
  2. Scenario: You want to show the change in temperature in Lagos from January to December. Which graph is best?
  3. Scenario: Your family wants to see how much of the monthly budget goes to food, rent, and school fees. Which graph would you use?

Group Activity

In groups of four, collect data on the favourite fruits of your classmates (e.g., mango, orange, banana, apple). Create a bar chart using R. Present your graph to the class and explain what it shows.


Individual Activity

Collect data on how you spend your time in a day (sleep, school, play, eating, homework). Create a pie chart using R. Write one sentence about what the pie chart tells you.


Classroom Discussion Questions

  1. Why do we need to label our graphs?
  2. What could go wrong if you use the wrong graph?
  3. How can graphs help Nigerian farmers?
  4. What is your favourite type of graph and why?

Mini Project

Title: "Our School Data Story"
Collect data on the number of pupils in each class in your school (or use sample data). Create a bar chart and a line chart (if you have data over time). Write a short report (2-3 paragraphs) explaining what the graphs show and why it is important.


Practical Assignment

Using R, create the following graphs from the mtcars dataset (built-in):

  1. A bar chart of the number of cars with different numbers of cylinders (cyl).
  2. A line chart of mpg (miles per gallon) vs wt (weight) – use geom_line().
  3. Add a title, x and y labels, and a theme.
  4. Save the graphs as PNG files.

Challenge Exercise

Find a real dataset about Nigeria (e.g., population, agriculture, weather). Create at least three different types of graphs (bar, line, pie) and arrange them in a report using R Markdown (optional). Explain the story behind each graph.


Quiz Answers

Fill-in-the-Blank: 1. visualisation, 2. ggplot2, 3. bar, 4. line, 5. pie, 6. labs, 7. plotly, 8. ggsave, 9. theme, 10. fill.

True/False: 1F, 2T, 3F, 4T, 5F, 6T, 7T, 8F, 9F, 10T.

Multiple Choice: 1B, 2C, 3B, 4B, 5C, 6B, 7C, 8C, 9B, 10B, 11C, 12A, 13A, 14B, 15C.


Key Takeaways

  • Visualisation makes data easy to understand.
  • ggplot2 is the main R package for graphs.
  • Bar charts compare categories, line charts show trends, pie charts show parts of a whole.
  • Always label your graphs.
  • Use themes to make your graphs clean.
  • Save your graphs with ggsave().
  • Interactive graphs can be made with plotly.
  • Choose the right graph for your data story.

Preparation for the Next Module

In Module 8, we will learn about data wrangling – cleaning and preparing data for analysis. We will use the dplyr package to filter, sort, and summarise data. Before the next module, try to practise loading dplyr with library(dplyr) and explore the iris dataset. Great job completing Module 7! You are now a data storyteller.


9

Module Eight

Module 8 Β· R for Data Analysis Β· Data Wrangling with dplyr

Module 8 Β· Data Wrangling – Cleaning and Preparing Data

Hello, young data explorer! In Module 7, we learned how to tell stories with data using pictures and graphs. But before we can make beautiful graphs, we need to make sure our data is clean and ready. Imagine you want to bake a cake – you need to wash the fruits, measure the flour, and mix everything properly. Data is just the same!

In this module, we will learn how to wrangle data. Data wrangling means cleaning, organising, and transforming data so that it is easy to work with. We will use a powerful R package called dplyr. It gives us simple "verbs" – like filter(), select(), mutate(), summarise(), and arrange() – that help us change our data in many useful ways.

By the end of this module, you will be able to take messy data, clean it up, and get it ready for analysis and graphs. You will be a data detective!


Learning Objectives

  • Understand what data wrangling means.
  • Install and load the dplyr package.
  • Use filter() to choose rows that meet a condition.
  • Use select() to choose specific columns.
  • Use mutate() to create new columns.
  • Use summarise() to get summary statistics (like mean, sum).
  • Use arrange() to sort data.
  • Use the pipe operator %>% to chain commands.
  • Apply these tools to real data, including Nigerian examples.
  • Recognise dirty data and know how to fix common issues.

Warm-up Story: The Messy Market List

Chidi loves helping his mother at the market. One day, his mother gave him a messy list of items to buy:

  • tomato, 5
  • onion, 3
  • yam, 2
  • tomato, 4 (wait, she added tomatoes again!)
  • pepper, 1
  • onion, 2 (another onion?)

Chidi was confused. The list had duplicates, the items were not sorted, and he couldn't tell how many of each item he really needed. He decided to clean the list:

  1. He combined the tomatoes: 5 + 4 = 9.
  2. He combined the onions: 3 + 2 = 5.
  3. He sorted the items alphabetically.

Now the list was clear:

  • onion – 5
  • pepper – 1
  • tomato – 9
  • yam – 2

Chidi felt proud. He had just wrangled the data! In R, we do the same thing with our datasets – we clean them, combine them, and sort them so we can understand them better.


Main Lessons

Lesson 1: What is Data Wrangling?

Definition: Data wrangling is the process of cleaning, transforming, and organising data to make it useful.

Why it is important: Real-world data is often messy – it has mistakes, missing values, or is in the wrong format. Wrangling fixes these issues.

Simple explanation: It's like tidying your room before you can find your toys easily.

Real-life example: A librarian organises books by category and author.

School example: Your teacher sorts test scores from highest to lowest.

Home example: You arrange your clothes by colour in the wardrobe.

Nigerian example: A farmer sorts his harvest by crop type and quality.

Illustration:

   Messy Data  -->  Clean Data  -->  Ready for Analysis
   (duplicates,    (no duplicates,
    errors,         correct format,
    unsorted)       sorted)

Mini summary: Data wrangling makes data clean and tidy so we can work with it easily.


Lesson 2: Meet dplyr – Your Data Wrangling Toolkit

Definition: dplyr is a popular R package that provides simple functions for data manipulation.

Why it is important: It makes data wrangling fast and easy, with intuitive "verbs" (words that describe actions).

Simple explanation: Think of dplyr as a Swiss Army knife for data – it has many tools in one.

Real-life example: A chef has different knives for cutting, peeling, and slicing – dplyr has different verbs for different tasks.

School example: A pencil case contains a pencil, an eraser, and a ruler – each does a specific job.

Home example: A toolbox has a hammer, a screwdriver, and a wrench.

Nigerian example: A mechanic uses different spanners for different bolts.

Code to install and load:

   install.packages("dplyr")   # do this once
   library(dplyr)              # do this every session

Mini summary: dplyr is a toolbox of verbs for cleaning and changing data.


Lesson 3: The Pipe Operator %>% – Connecting Steps

Definition: The pipe operator %>% (pronounced "then") lets you chain multiple operations together in a sequence.

Why it is important: It makes your code easier to read and write, like a recipe with steps.

Simple explanation: It means "and then do this". Instead of writing many nested functions, you write a clear list of actions.

Real-life example: A recipe: "Take eggs, then break them, then whisk them."

School example: "Open your book, then read page 10, then answer the questions."

Home example: "Take the laundry, then put it in the washing machine, then turn it on."

Nigerian example: "Get the yam, then peel it, then cut it, then boil it."

Illustration:

   data %>%
     filter(...) %>%
     select(...) %>%
     arrange(...)
   # This means: take data, THEN filter, THEN select, THEN arrange.

Mini summary: The pipe %>% helps us chain commands in a logical order.


Lesson 4: filter() – Picking Rows

Definition: filter() picks rows that meet a condition (like "age > 10" or "city == 'Lagos'").

Why it is important: Often we only need a part of the data – e.g., only pupils from a certain class.

Simple explanation: It's like using a sieve to separate big pieces from small pieces.

Real-life example: A shopkeeper filters out expired products.

School example: Your teacher filters the list to only show pupils who scored above 80%.

Home example: You filter your toys to only show the ones you want to give away.

Nigerian example: A researcher filters survey data to only include respondents from Kano.

Code:

   library(dplyr)

   # Sample data: pupils
   pupils <- data.frame(
       name = c("Ade", "Bola", "Chidi", "Dami"),
       age = c(10, 11, 9, 12),
       score = c(85, 70, 90, 65)
   )

   # Filter to only those with score > 80
   pupils %>%
       filter(score > 80)
   # Result: Ade (85) and Chidi (90)
   Before filter:
   name   age  score
   Ade    10   85
   Bola   11   70
   Chidi  9    90
   Dami   12   65

   After filter(score > 80):
   name   age  score
   Ade    10   85
   Chidi  9    90

Mini summary: filter() keeps only rows that satisfy a condition.


Lesson 5: select() – Picking Columns

Definition: select() picks specific columns (variables) from your data.

Why it is important: Sometimes we have too many columns and only need a few.

Simple explanation: It's like highlighting important words in a paragraph.

Real-life example: When you read a long article, you focus on the main points.

School example: A teacher might only look at the "name" and "score" columns.

Home example: When you check your phone, you only look at messages, not all apps.

Nigerian example: An economist selects only "State" and "GDP" columns from a big table.

Code:

   # Select name and score
   pupils %>%
       select(name, score)
   # Result: only those two columns.
   Before select:
   name   age  score
   Ade    10   85
   Bola   11   70

   After select(name, score):
   name   score
   Ade    85
   Bola   70

Mini summary: select() keeps only the columns you choose.


Lesson 6: mutate() – Creating New Columns

Definition: mutate() creates a new column (variable) from existing columns, or modifies an existing one.

Why it is important: Often we need to calculate new values, like total or percentage.

Simple explanation: It's like adding a new row to your table with new information.

Real-life example: You calculate your total pocket money by adding your allowance and gifts.

School example: You create a column for "percentage" from test scores.

Home example: You convert minutes into hours and minutes.

Nigerian example: A farmer calculates total harvest in kg by adding different crops.

Code:

   # Add a new column: score in percentage (out of 100)
   pupils %>%
       mutate(percentage = score / 100 * 100)  # just score itself, but let's add 10 bonus points
       # Actually, let's add bonus: total = score + 5
       mutate(total_score = score + 5)
   Before mutate:
   name   age  score
   Ade    10   85

   After mutate(total_score = score + 5):
   name   age  score  total_score
   Ade    10   85     90

Mini summary: mutate() adds new columns by performing calculations.


Lesson 7: summarise() – Getting Summaries

Definition: summarise() (or summarize()) collapses your data into a single summary row – like mean, sum, count.

Why it is important: It gives you the big picture (e.g., average score, total votes).

Simple explanation: It's like asking "What is the total?" or "What is the average?"

Real-life example: You calculate the total money spent on your shopping trip.

School example: You find the average test score of your class.

Home example: You sum up the minutes you spent on homework each day.

Nigerian example: You summarise the total population of all states.

Code:

   # Get average age and average score
   pupils %>%
       summarise(avg_age = mean(age), avg_score = mean(score))
   # Result: one row with the averages.
   Before summarise:
   name   age  score
   Ade    10   85
   Bola   11   70
   Chidi  9    90
   Dami   12   65

   After summarise(avg_age = mean(age), avg_score = mean(score)):
   avg_age  avg_score
   10.5     77.5

Mini summary: summarise() gives you one number or a few numbers that describe your data.


Lesson 8: group_by() + summarise() – Grouped Summaries

Definition: group_by() splits your data into groups, and then you apply summarise() to each group.

Why it is important: It lets you compare groups (e.g., average score by class).

Simple explanation: It's like sorting your toys by colour and then counting how many of each colour you have.

Real-life example: A shop counts total sales per product category.

School example: Find the average score for each class.

Home example: Count how many hours you spend on each activity (play, study, sleep).

Nigerian example: Calculate the average income per state.

Code:

   # Group by age, then get average score per age group
   pupils %>%
       group_by(age) %>%
       summarise(avg_score = mean(score))
   Data grouped by age:
   age  avg_score
   9    90
   10   85
   11   70
   12   65

Mini summary: group_by() + summarise() gives summaries for each group.


Lesson 9: arrange() – Sorting Data

Definition: arrange() sorts your rows by one or more columns, in ascending (default) or descending order.

Why it is important: It helps you see the highest, lowest, or alphabetical order.

Simple explanation: It's like putting your books on a shelf from A to Z.

Real-life example: A teacher sorts pupils' names alphabetically.

School example: You sort test scores from highest to lowest.

Home example: You arrange your clothes from biggest to smallest.

Nigerian example: A business owner sorts sales by highest revenue.

Code:

   # Sort by score ascending (lowest to highest)
   pupils %>%
       arrange(score)

   # Sort by score descending (highest first)
   pupils %>%
       arrange(desc(score))
   Original:
   name   score
   Ade    85
   Bola   70
   Chidi  90
   Dami   65

   arrange(desc(score)):
   name   score
   Chidi  90
   Ade    85
   Bola   70
   Dami   65

Mini summary: arrange() sorts rows by column values.


Lesson 10: Combining Verbs – The Power of Pipes

We can combine filter(), select(), mutate(), arrange(), and summarise() in a single pipeline.

Example: We want to:

  1. Filter to only pupils with score > 70.
  2. Select name and score.
  3. Add a column "bonus" = score + 5.
  4. Sort by score descending.
   pupils %>%
       filter(score > 70) %>%
       select(name, score) %>%
       mutate(bonus = score + 5) %>%
       arrange(desc(score))
   Result:
   name   score  bonus
   Chidi  90     95
   Ade    85     90

Mini summary: Pipes let you chain multiple verbs in one clear sequence.


Lesson 11: Handling Missing Data – NA Values

Definition: Missing data are empty values, shown as NA in R.

Why it is important: Missing values can break calculations if not handled.

Simple explanation: It's like a page missing from your book – you need to decide what to do.

Real-life example: A survey form where someone didn't answer a question.

School example: A pupil was absent on the day of the test.

Home example: You forgot to record the time you spent reading.

Nigerian example: A farmer didn't record the harvest for one month.

How to handle: Use filter() to remove NAs, or use na.rm = TRUE inside functions like mean().

   # Remove rows with NA in score
   pupils %>% filter(!is.na(score))

   # Get mean ignoring NAs
   summarise(avg = mean(score, na.rm = TRUE))

Mini summary: Always check for NA values and decide to remove them or ignore them in calculations.


Lesson 12: Renaming Columns – rename()

Definition: rename() changes the names of columns.

Why it is important: Column names might be too long, misspelled, or unclear.

Simple explanation: It's like giving your pet a new nickname.

Real-life example: You change a file name to make it more descriptive.

School example: Your teacher renames "Test_1" to "Math_Test".

Home example: You rename your folders from "Stuff" to "Schoolwork".

Nigerian example: An analyst renames a column from "Pop" to "Population".

   pupils %>%
       rename(Score = score, Name = name)

Mini summary: rename() changes column names to be clearer.


Lesson 13: Distinct Rows – distinct()

Definition: distinct() removes duplicate rows from your data.

Why it is important: Duplicates can overcount and give wrong results.

Simple explanation: It's like removing double entries in your list.

Real-life example: You delete duplicate contacts on your phone.

School example: You remove duplicate names from the class register.

Home example: You make sure you don't have two of the same toy listed.

Nigerian example: A voter register removes duplicate names.

   pupils %>% distinct()   # removes exact duplicate rows
   pupils %>% distinct(name, .keep_all = TRUE)  # keep only unique names

Mini summary: distinct() removes duplicate rows.


Lesson 14: Applying dplyr to Real Nigerian Data

Let's use a small dataset about Nigerian states and their populations (example). We will wrangle it.

   # Sample data (population in millions)
   states <- data.frame(
       State = c("Lagos", "Kano", "Oyo", "Lagos", "Rivers", "Kano"),
       Population = c(20, 15, 8, 20, 7, 15),
       Region = c("SW", "NW", "SW", "SW", "SS", "NW")
   )
   # We have duplicates! Let's remove them.
   states_unique <- states %>% distinct(State, .keep_all = TRUE)

   # Filter to only SW states
   sw_states <- states_unique %>% filter(Region == "SW")

   # Select State and Population
   sw_pop <- sw_states %>% select(State, Population)

   # Arrange by population descending
   sw_pop %>% arrange(desc(Population))
   Result:
   State   Population
   Lagos   20
   Oyo     8

Mini summary: You can use all dplyr verbs on Nigerian data to clean and explore it.


Lesson 15: Summary of dplyr Verbs

We have learned six main verbs:

  • filter() – pick rows.
  • select() – pick columns.
  • mutate() – create new columns.
  • summarise() – get summaries (mean, sum, etc.).
  • arrange() – sort rows.
  • group_by() – split data into groups.

And the pipe %>% connects them.

Remember: These verbs make data wrangling fun and easy!


Key Vocabulary

Data Wrangling
Cleaning and organising data to make it ready for analysis.
dplyr
An R package with verbs for data manipulation.
Pipe (%>%)
An operator that chains commands together.
filter()
Selects rows that meet a condition.
select()
Selects columns by name.
mutate()
Adds new columns.
summarise()
Calculates summary statistics.
arrange()
Sorts rows.
group_by()
Splits data into groups for grouped operations.
NA
Missing value in R.

Important Concepts

  • Tidy Data: Each variable is a column, each observation is a row.
  • Piping: The %>% operator makes code readable.
  • Non-destructive: dplyr functions do not change the original data unless you assign the result.
  • Vectorized Operations: Operations are applied to whole columns efficiently.

Step-by-Step Explanations

How to clean a dataset with dplyr:

  1. Load dplyr: library(dplyr)
  2. Remove duplicates: data %>% distinct()
  3. Filter rows: data %>% filter(condition)
  4. Select columns: data %>% select(col1, col2)
  5. Mutate (add columns): data %>% mutate(new = old * 2)
  6. Summarise: data %>% summarise(mean = mean(col))
  7. Arrange: data %>% arrange(desc(col))

Real-life Examples

  • A shop uses filter() to find products with low stock.
  • A teacher uses arrange() to sort students by grades.
  • A hospital uses summarise() to get average patient age.

Nigerian Examples

  • A government agency uses filter() to get data for only Lagos State.
  • An economist uses group_by() and summarise() to calculate average income by region.
  • A school uses mutate() to add a "pass/fail" column based on scores.

Fun Examples Children Relate To

  • filter(): Select only your favourite toys from a list.
  • select(): Choose only the name and colour of your toys.
  • mutate(): Add a column for "toy value" = price + age.
  • arrange(): Sort your toys by size from smallest to largest.

Everyday Examples

  • You use a shopping list and cross out items you already bought (filter).
  • You write a list of your friends' phone numbers and sort them alphabetically.
  • You calculate your total pocket money for the month.

Teacher Notes

  • Emphasise that real data is messy and wrangling is a necessary step.
  • Use small datasets for practice so pupils can see the changes.
  • Encourage pupils to experiment with different verbs.
  • Relate the verbs to everyday actions (filter = sieve, select = choose, etc.).

Parent Tips

  • Help your child collect simple data at home (e.g., daily temperatures).
  • Ask them to clean the data (remove duplicates, sort).
  • Discuss how cleaning data helps in real life (e.g., organising a party).

Interesting Facts

  • The dplyr package was created by Hadley Wickham, a famous data scientist.
  • Data scientists spend about 80% of their time on data wrangling!
  • dplyr is used by companies like Facebook and Google.
  • R has over 15,000 packages, and dplyr is one of the most popular.

Did You Know?

  • dplyr can work with large datasets very quickly because it uses C++ behind the scenes.
  • You can use dplyr with databases without loading all data into memory.
  • There is a package called dtplyr that uses dplyr syntax on data.table for even more speed.

Remember This

  • Always load dplyr with library(dplyr).
  • Use the pipe %>% to chain commands.
  • filter() picks rows, select() picks columns.
  • mutate() creates new columns, summarise() gives summaries.
  • arrange() sorts, group_by() splits into groups.

Common Mistakes

  • Forgetting to load dplyr before using its functions.
  • Using filter() with the wrong condition (e.g., == instead of =).
  • Not using na.rm = TRUE when calculating means with missing values.
  • Forgetting to assign the result to a new variable; dplyr does not modify in place.
  • Using the pipe without proper indentation (though it still works, it's harder to read).

Best Practices

  • Start with a clear question: "What do I want to do with this data?"
  • Use descriptive variable names.
  • Chain operations with pipes for readability.
  • Check for NA values early.
  • Use glimpse() or str() to understand your data before wrangling.

ASCII Illustrations

Piping concept

   data
     |
     V
   filter()   -->  (keeps rows)
     |
     V
   select()   -->  (keeps columns)
     |
     V
   mutate()   -->  (adds columns)
     |
     V
   arrange()  -->  (sorts rows)

Filtering rows

   Before filter:
   +------+-----+-------+
   | name | age | score |
   +------+-----+-------+
   | Ade  | 10  | 85    |
   | Bola | 11  | 70    |
   | Chidi| 9   | 90    |
   +------+-----+-------+

   filter(score > 80):
   +------+-----+-------+
   | name | age | score |
   +------+-----+-------+
   | Ade  | 10  | 85    |
   | Chidi| 9   | 90    |
   +------+-----+-------+

Selecting columns

   Before select:
   +------+-----+-------+
   | name | age | score |
   +------+-----+-------+
   | Ade  | 10  | 85    |
   +------+-----+-------+

   select(name, score):
   +------+-------+
   | name | score |
   +------+-------+
   | Ade  | 85    |
   +------+-------+

Comparison Tables

dplyr Verbs and Their Uses
VerbWhat it doesExample
filter()Picks rows based on conditionfilter(score > 80)
select()Picks columns by nameselect(name, score)
mutate()Creates new columnsmutate(total = score + 5)
summarise()Calculates summary statssummarise(avg = mean(score))
arrange()Sorts rowsarrange(desc(score))
group_by()Splits data into groupsgroup_by(age)

End-of-Module Summary

In this module, we learned how to wrangle data – that is, clean and prepare it for analysis. We used the dplyr package, which provides simple verbs: filter(), select(), mutate(), summarise(), arrange(), and group_by(). We also learned the pipe %>% to chain commands. We handled missing values (NA), removed duplicates, and renamed columns. Data wrangling is an essential step before any analysis or visualisation. Remember: clean data leads to clear insights!


Frequently Asked Questions

1. What is data wrangling?
Cleaning and organising data to make it useful.
2. What is dplyr?
An R package for data manipulation.
3. How do I install dplyr?
install.packages("dplyr")
4. What does filter() do?
Selects rows that meet a condition.
5. What is the pipe %>%?
It chains commands together.
6. How do I remove duplicates?
Use distinct().
7. What is mutate() used for?
To create new columns.
8. How do I sort data?
Use arrange().
9. What is NA?
A missing value in R.
10. Can I use dplyr with other packages?
Yes, it works well with ggplot2, tidyr, and others.

Review Questions

  1. What is data wrangling?
  2. What package do we use for data wrangling in R?
  3. What does the pipe %>% do?
  4. Which verb selects rows based on a condition?
  5. Which verb selects columns?
  6. Which verb creates new columns?
  7. Which verb calculates summary statistics?
  8. Which verb sorts data?
  9. How do you remove duplicate rows?
  10. What does NA represent?
  11. How do you handle missing values in summarise()?
  12. What does group_by() do?
  13. Give an example of using filter() in real life.
  14. What is the benefit of using pipes?
  15. Why is data wrangling important before making graphs?

Fill-in-the-Blank Exercises

  1. Data ______ means cleaning and organising data. (wrangling)
  2. ______ is an R package for data manipulation. (dplyr)
  3. The pipe operator is ______. (%>%)
  4. ______ selects rows that meet a condition. (filter)
  5. ______ selects columns by name. (select)
  6. ______ creates new columns. (mutate)
  7. ______ calculates summary statistics. (summarise)
  8. ______ sorts rows. (arrange)
  9. ______ splits data into groups. (group_by)
  10. Missing values in R are shown as ______. (NA)

True or False Exercises

  1. Data wrangling is only for experts. (False)
  2. dplyr is a package in R. (True)
  3. The pipe operator is written as <-. (False)
  4. filter() picks columns. (False)
  5. select() picks rows. (False)
  6. mutate() adds new columns. (True)
  7. summarise() gives one row of summaries. (True)
  8. arrange() sorts columns. (False)
  9. group_by() is used with summarise(). (True)
  10. NA means missing value. (True)

Multiple Choice Questions

  1. Which package is used for data wrangling?
    A. ggplot2 B. dplyr C. tidyr D. readr
    Answer: B
  2. Which verb selects rows?
    A. select() B. filter() C. mutate() D. arrange()
    Answer: B
  3. Which verb creates new columns?
    A. summarise() B. mutate() C. filter() D. arrange()
    Answer: B
  4. What does the pipe %>% do?
    A. Sorts data B. Chains commands C. Filters rows D. Selects columns
    Answer: B
  5. Which verb gives summary statistics?
    A. mutate() B. summarise() C. filter() D. select()
    Answer: B
  6. How do you sort data in descending order?
    A. arrange(desc(col)) B. arrange(col) C. sort(desc(col)) D. order(col)
    Answer: A
  7. What does NA stand for?
    A. Not Available B. No Answer C. New Attribute D. None
    Answer: A
  8. Which verb removes duplicate rows?
    A. unique() B. distinct() C. filter() D. arrange()
    Answer: B
  9. Which verb splits data into groups?
    A. group_by() B. split() C. arrange() D. filter()
    Answer: A
  10. What is the first step to use dplyr?
    A. install.packages("dplyr") B. library(dplyr) C. Both A and B D. None
    Answer: C
  11. Which verb renames columns?
    A. rename() B. select() C. mutate() D. summarise()
    Answer: A
  12. Which function shows structure of data?
    A. str() B. glimpse() C. Both A and B D. view()
    Answer: C
  13. What is the result of summarise()?
    A. A new column B. A summary table C. Filtered rows D. Sorted data
    Answer: B
  14. Can dplyr be used with databases?
    A. Yes B. No C. Only with small data D. Only with CSV
    Answer: A
  15. Which of these is NOT a dplyr verb?
    A. filter B. select C. plot D. mutate
    Answer: C

Matching Exercises

VerbDescription
1. filterA. Adds new columns
2. selectB. Picks rows that meet condition
3. mutateC. Chooses columns
4. summariseD. Sorts rows
5. arrangeE. Calculates summary statistics

Answers: 1-B, 2-C, 3-A, 4-E, 5-D


Short Answer Questions

  1. What is data wrangling and why is it important?
  2. Explain the difference between filter() and select().
  3. How do you use the pipe operator? Give an example.
  4. What does mutate() do? Provide a simple example.
  5. How do you handle missing values in dplyr?

Scenario-based Exercises

  1. Scenario: You have a dataset of pupils with columns: name, age, class, score. You want to get the average score for each class. Which verbs would you use?
  2. Scenario: Your data has duplicate rows and missing scores. How would you clean it?
  3. Scenario: You want to create a new column "grade" based on score (A for >80, B for 60-80, etc.). Which verb do you use?

Group Activity

In groups of four, create a small dataset (10 rows) about your favourite foods, with columns: food, cost, rating. Then use dplyr to:

  • Filter to only foods with rating > 4.
  • Select food and cost.
  • Add a new column "total_cost" = cost * 2.
  • Arrange by cost descending.

Present your code and output to the class.


Individual Activity

Use the iris dataset (built-in) and dplyr to:

  1. Filter to only setosa species.
  2. Select Sepal.Length and Petal.Length.
  3. Add a new column "ratio" = Sepal.Length / Petal.Length.
  4. Arrange by ratio descending.

Classroom Discussion Questions

  1. Why is data wrangling an important skill?
  2. What happens if you don't clean your data before analysis?
  3. How can data wrangling help Nigerian businesses?
  4. Which dplyr verb do you find most useful and why?

Mini Project

Title: "Clean the School Data"
You are given a messy dataset of school attendance (columns: name, class, days_present, days_absent) with duplicates, missing values, and errors. Use dplyr to:

  1. Remove duplicates.
  2. Filter out rows with missing values.
  3. Add a column "total_days" = days_present + days_absent.
  4. Calculate average attendance per class.
  5. Sort by class.

Practical Assignment

Using the mtcars dataset, perform the following using dplyr:

  1. Filter to cars with more than 4 cylinders.
  2. Select mpg, cyl, and hp.
  3. Add a new column "hp_per_cyl" = hp / cyl.
  4. Group by cyl and calculate average mpg.
  5. Arrange by average mpg descending.

Challenge Exercise

Find a real, messy dataset online (e.g., from Kaggle or Nigerian open data). Use dplyr to perform a complete cleaning and summarisation. Write a short report (2 paragraphs) describing your steps and what you discovered.


Quiz Answers

Fill-in-the-Blank: 1. wrangling, 2. dplyr, 3. %>%, 4. filter, 5. select, 6. mutate, 7. summarise, 8. arrange, 9. group_by, 10. NA.

True/False: 1F, 2T, 3F, 4F, 5F, 6T, 7T, 8F, 9T, 10T.

Multiple Choice: 1B, 2B, 3B, 4B, 5B, 6A, 7A, 8B, 9A, 10C, 11A, 12C, 13B, 14A, 15C.


Key Takeaways

  • Data wrangling is cleaning and organising data.
  • dplyr provides verbs: filter(), select(), mutate(), summarise(), arrange(), group_by().
  • The pipe %>% chains commands.
  • Always check for NA values.
  • Use distinct() to remove duplicates.
  • Wrangling makes data ready for visualisation and analysis.

Preparation for the Next Module

In Module 9, we will learn about joining datasets – combining information from different tables, like merging pupil data with class data. We will use dplyr joins: left_join(), inner_join(), and more. Before that, practise the verbs you learned in this module on different datasets. Well done! You are becoming a data wrangler!


10

Module Nine

Module 9 Β· R for Data Analysis Β· Joining Tables with dplyr

Module 9 Β· Joining Tables – Combining Data from Multiple Sources

Hello, young data explorer! In Module 8, we learned how to clean and prepare our data using dplyr. We used verbs like filter(), select(), and mutate(). But what if our data is spread across two or more tables? For example, one table has pupils' names and ages, and another table has their test scores. How do we bring them together?

In this module, we will learn how to join tables. Joining means combining rows or columns from different tables based on a common key (like an ID or name). It's like putting puzzle pieces together to get a complete picture.

We will use the dplyr package again, because it has special functions for joining: inner_join(), left_join(), right_join(), and full_join(). By the end of this module, you will be able to merge different datasets and unlock new insights.


Learning Objectives

  • Understand why we need to join tables.
  • Learn what a key is and how it is used to match rows.
  • Use inner_join() to keep only rows that match in both tables.
  • Use left_join() to keep all rows from the left table and matching rows from the right.
  • Use right_join() to keep all rows from the right table and matching rows from the left.
  • Use full_join() to keep all rows from both tables.
  • Understand the difference between the four join types.
  • Apply joins to real data, including Nigerian examples.
  • Handle duplicate keys and missing values after joins.
  • Choose the right join for your analysis.

Warm-up Story: The School Records Puzzle

At Sunshine Primary School, the principal, Mrs. Ade, kept two lists. The first list had pupils' names and their class:

  • Ade – Class 5
  • Bola – Class 6
  • Chidi – Class 5
  • Dami – Class 4

The second list had pupils' names and their favourite subject:

  • Ade – Maths
  • Bola – English
  • Chidi – Science
  • Efe – Art (Efe was new, not in the first list)

Mrs. Ade wanted a single list with each pupil's class and favourite subject. She needed to join the two lists using the pupils' names as the common key. She used a left_join() to keep all pupils from the first list and add their favourite subjects. For Efe, who wasn't in the first list, she had no class – so it showed NA.

Now, the combined list looked like this:

  • Ade – Class 5, Favourite: Maths
  • Bola – Class 6, Favourite: English
  • Chidi – Class 5, Favourite: Science
  • Dami – Class 4, Favourite: NA

Mrs. Ade was happy because she could see everything in one place. That is the power of joining tables!


Main Lessons

Lesson 1: What is a Join?

Definition: A join is a way to combine two tables by matching rows based on a common column (called a key).

Why it is important: Data is often stored in separate tables to avoid duplication. Joins let us bring them back together for analysis.

Simple explanation: It's like joining two pieces of a puzzle that have matching edges.

Real-life example: A shop has a list of products (product ID, name) and a list of sales (product ID, quantity). Joining them gives each product's sales.

School example: A teacher has a table of pupils' names and a table of their test scores. Joining gives a complete view.

Home example: You have a list of family members and a list of their birthdays. Joining them gives everyone's birthday.

Nigerian example: A government agency has a table of states and a table of governors. Joining gives the governor for each state.

Illustration:

   Table A (Pupils)         Table B (Favourite Subject)
   +------+-------+         +------+----------+
   | name | class |         | name | subject  |
   +------+-------+         +------+----------+
   | Ade  | 5     |         | Ade  | Maths    |
   | Bola | 6     |         | Bola | English  |
   +------+-------+         +------+----------+

   Join on name:
   +------+-------+----------+
   | name | class | subject  |
   +------+-------+----------+
   | Ade  | 5     | Maths    |
   | Bola | 6     | English  |
   +------+-------+----------+

Mini summary: Joins combine tables using a common column (key).


Lesson 2: The Key – The Matching Column

Definition: A key is a column that appears in both tables and is used to match rows.

Why it is important: Without a key, we cannot know which rows belong together.

Simple explanation: The key is like the "glue" that sticks the two tables together.

Real-life example: In school, your admission number is a key that links your name to your grades.

School example: The "student ID" column is a key.

Home example: Your phone number can be a key to find your address in a phone book.

Nigerian example: The "state code" is a key to join state names with their capital cities.

Illustration:

   Table A           Table B
   +-------+------+  +-------+--------+
   | id    | name |  | id    | score  |
   +-------+------+  +-------+--------+
   | 1     | Ade  |  | 1     | 85     |
   | 2     | Bola |  | 2     | 70     |
   +-------+------+  +-------+--------+
   The key is "id". It appears in both tables.

Mini summary: A key is the column that matches rows between tables.


Lesson 3: Inner Join – Only Matching Rows

Definition: inner_join() keeps only rows that have a match in both tables.

Why it is important: It gives you only the complete cases where information is available from both sides.

Simple explanation: It's like inviting only the friends who RSVP'd "yes" to your party.

Real-life example: A company wants to see only employees who have a department assigned.

School example: You want to see only pupils who have both a name and a score.

Home example: You want to see only family members who have both a birthday and a favourite meal.

Nigerian example: You want to see only states that have both a governor and a population figure.

Code:

   library(dplyr)

   # Table A: pupils
   pupils <- data.frame(
       name = c("Ade", "Bola", "Chidi"),
       class = c(5, 6, 5)
   )

   # Table B: subjects
   subjects <- data.frame(
       name = c("Ade", "Bola", "Dami"),
       subject = c("Maths", "English", "Art")
   )

   # Inner join
   pupils %>%
       inner_join(subjects, by = "name")
   # Result: only Ade and Bola (common names)
   Result:
   name   class  subject
   Ade    5      Maths
   Bola   6      English

Mini summary: inner_join() keeps only rows that match in both tables.


Lesson 4: Left Join – All Left, Matching Right

Definition: left_join() keeps all rows from the left table and adds matching rows from the right. If no match, it fills with NA.

Why it is important: Often we want to keep all our main data and just "decorate" it with extra information.

Simple explanation: It's like starting with your full guest list and only adding RSVP notes for those who replied.

Real-life example: You have a list of all products and you want to add their current stock from another table.

School example: You have a list of all pupils and you want to add their scores (even if some scores are missing).

Home example: You have a list of all your chores and you want to add the time each took (some may be empty).

Nigerian example: You have a list of all states and you want to add their populations (some may be missing).

Code:

   # Left join
   pupils %>%
       left_join(subjects, by = "name")
   # Result: all pupils, with subjects added.
   Result:
   name   class  subject
   Ade    5      Maths
   Bola   6      English
   Chidi  5      NA       (no match in subjects)

Mini summary: left_join() keeps all rows from the left table.


Lesson 5: Right Join – All Right, Matching Left

Definition: right_join() keeps all rows from the right table and adds matching rows from the left. If no match, it fills with NA.

Why it is important: It's useful when the right table is the main one you care about.

Simple explanation: It's like starting with the RSVP list and adding guest details from the main list.

Real-life example: You have a list of orders (right) and a product list (left). You want to see all orders with product details.

School example: You have a list of scores (right) and want to add pupil names (left).

Home example: You have a list of expenses (right) and want to add category descriptions (left).

Nigerian example: You have a list of governors (right) and want to add state names (left).

Code:

   # Right join
   pupils %>%
       right_join(subjects, by = "name")
   # Result: all subjects, with pupils added.
   Result:
   name   class  subject
   Ade    5      Maths
   Bola   6      English
   Dami   NA     Art      (no match in pupils)

Mini summary: right_join() keeps all rows from the right table.


Lesson 6: Full Join – All Rows from Both Tables

Definition: full_join() keeps all rows from both tables, matching where possible and filling with NA where no match.

Why it is important: It gives you the complete union of both datasets.

Simple explanation: It's like inviting everyone from two guest lists and merging them.

Real-life example: You have two customer lists and you want to combine them into one master list.

School example: You have two class lists and want to combine them.

Home example: You have a list of your toys and a list of your sibling's toys – combine them.

Nigerian example: You have a list of states and a list of capital cities – combine them.

Code:

   # Full join
   pupils %>%
       full_join(subjects, by = "name")
   # Result: all names from both tables.
   Result:
   name   class  subject
   Ade    5      Maths
   Bola   6      English
   Chidi  5      NA
   Dami   NA     Art

Mini summary: full_join() keeps all rows from both tables.


Lesson 7: Joining with Different Column Names

Sometimes the key columns have different names in each table. You can specify by = c("col1" = "col2").

Example:

   # Table A has column "student_id", Table B has column "id"
   # We want to join on student_id = id
   tableA %>%
       left_join(tableB, by = c("student_id" = "id"))

Mini summary: Use by = c("name_in_A" = "name_in_B") when key names differ.


Lesson 8: Handling Duplicate Keys

If there are duplicate keys in one table, the join will produce all combinations (cartesian product). This can lead to more rows than expected.

Why it is important: Duplicates can cause your results to be overcounted. Always check for duplicates before joining.

Solution: Use distinct() to remove duplicates from the key column.

   tableA %>% distinct(key, .keep_all = TRUE)

Mini summary: Check for duplicate keys; remove them with distinct() if needed.


Lesson 9: Joining with Multiple Keys

You can join on two or more columns by providing a vector of column names: by = c("col1", "col2").

Example: Join on "first_name" and "last_name" together to avoid confusion.

   tableA %>%
       left_join(tableB, by = c("first_name", "last_name"))

Mini summary: Use multiple columns as a composite key when one column is not enough.


Lesson 10: Real Nigerian Example – States and Governors

Let's join a table of states with a table of governors.

   states <- data.frame(
       state = c("Lagos", "Kano", "Oyo", "Rivers"),
       region = c("SW", "NW", "SW", "SS")
   )

   governors <- data.frame(
       state = c("Lagos", "Kano", "Rivers", "Abia"),
       governor = c("Sanwo-Olu", "Yusuf", "Fubara", "Ottu")
   )

   # Left join to keep all states
   states %>%
       left_join(governors, by = "state")
   # Result: Lagos, Kano, Oyo (NA governor), Rivers
   Result:
   state    region  governor
   Lagos    SW      Sanwo-Olu
   Kano     NW      Yusuf
   Oyo      SW      NA
   Rivers   SS      Fubara

Mini summary: Joins work perfectly with Nigerian data too.


Lesson 11: Visualising Joins – Venn Diagrams

Joins can be visualised as Venn diagrams:

   inner_join:  only the overlapping part.
   left_join:   all of left circle + overlap.
   right_join:  all of right circle + overlap.
   full_join:   entire union of both circles.
   Venn diagram concept:
   +----------+    +----------+
   |  left    |    |  right   |
   |  table   |    |  table   |
   +----------+    +----------+
   Overlap = matching rows

Mini summary: Joins are like Venn diagrams for data.


Lesson 12: Choosing the Right Join

How to choose:

  • Use inner_join when you only want rows that exist in both tables.
  • Use left_join when you want all rows from the main table and add extra info.
  • Use right_join when you want all rows from the secondary table.
  • Use full_join when you want to combine all rows.

Mini summary: Choose the join that keeps the rows you need for your analysis.


Lesson 13: Joining More Than Two Tables

You can chain multiple joins using pipes.

   tableA %>%
       left_join(tableB, by = "key1") %>%
       left_join(tableC, by = "key2")

Mini summary: You can join many tables one after another.


Lesson 14: Checking Your Join Result

After joining, always check:

  • Number of rows (did you get what you expected?).
  • Any NA values – are they okay?
  • Any duplicate rows?

Use str() or glimpse() to see the structure.

   result %>% glimpse()

Mini summary: Always inspect your joined data to make sure it's correct.


Lesson 15: Summary of Joins

We have learned four main joins:

  • inner_join() – only matches.
  • left_join() – all left, matches right.
  • right_join() – all right, matches left.
  • full_join() – all rows from both.

And we learned about keys, handling duplicates, and checking results.


Key Vocabulary

Join
Combining two tables based on a key.
Key
A column used to match rows between tables.
inner_join
Keeps only rows that match in both tables.
left_join
Keeps all rows from the left table.
right_join
Keeps all rows from the right table.
full_join
Keeps all rows from both tables.
Duplicate
When a key appears more than once in a table.
NA
Missing value after a join when there is no match.

Important Concepts

  • Key: The link between tables.
  • Match vs. No Match: Matches produce values; no matches produce NA.
  • Duplicate Keys: Can cause row multiplication.
  • Composite Key: Using more than one column as a key.

Step-by-Step Explanations

How to perform a left join:

  1. Load dplyr: library(dplyr)
  2. Identify the left table (main) and right table (secondary).
  3. Identify the key column (same name or different).
  4. Use left_join(left_table, right_table, by = "key").
  5. Check the result for NA values and row count.

Real-life Examples

  • A library joins book records with borrower records to see who has which book.
  • A hospital joins patient records with treatment records.
  • A shop joins product list with sales data to calculate total revenue.

Nigerian Examples

  • Join state population data with state governor data.
  • Join school attendance data with exam scores.
  • Join agricultural yield data with rainfall data.

Fun Examples Children Relate To

  • Join a list of your friends with their favourite ice cream flavours.
  • Join a list of your toys with their colours.
  • Join a list of your books with their authors.

Everyday Examples

  • You have a contact list and a birthday list – join them to see birthdays.
  • You have a shopping list and a price list – join to get total cost.
  • You have a class list and a list of group assignments – join to assign groups.

Teacher Notes

  • Use small, relatable datasets to illustrate joins.
  • Draw Venn diagrams on the board to explain each join type.
  • Encourage students to predict the output before running the code.
  • Emphasise the importance of checking for duplicates.

Parent Tips

  • Help your child find examples of joining in everyday life (e.g., phone book with addresses).
  • Ask them to think about what happens when information is missing.
  • Encourage them to practice with simple lists at home.

Interesting Facts

  • Joins are one of the most common operations in SQL (a language for databases).
  • The concept of joins was introduced in the 1970s.
  • Data scientists often use joins to combine data from multiple sources.

Did You Know?

  • There are other types of joins like semi_join() and anti_join() that filter rather than combine.
  • You can join tables from different databases using R.
  • Joins can be used to merge data from Excel files, CSVs, and databases.

Remember This

  • Always identify the key column(s) before joining.
  • left_join() is the most commonly used join.
  • Check for NA values after a join.
  • Use distinct() to remove duplicate keys if necessary.
  • Inspect your joined data with glimpse() or head().

Common Mistakes

  • Forgetting to specify the by argument when key names are different.
  • Joining on columns that have duplicate values, causing row multiplication.
  • Not realising that NA values mean no match.
  • Using the wrong join type for the analysis.
  • Joining without checking the structure of the tables first.

Best Practices

  • Always start with a small subset of data to test your join.
  • Use by explicitly, even if the key names are the same.
  • Remove duplicates from keys using distinct() before joining.
  • Use left_join() unless you have a specific reason for another join.
  • Document your joins – explain why you chose a particular type.

ASCII Illustrations

Join Types (Venn diagrams)

   inner_join:       [overlap only]
      +------+------+
      |      |      |
      |  A   |  B   |
      |      |      |
      +------+------+

   left_join:        [all of A + overlap]
      +------+------+
      |      |      |
      |  A   |  B   |
      |      |      |
      +------+------+

   right_join:       [all of B + overlap]
      +------+------+
      |      |      |
      |  A   |  B   |
      |      |      |
      +------+------+

   full_join:        [all of A and all of B]
      +------+------+
      |      |      |
      |  A   |  B   |
      |      |      |
      +------+------+

Join process

   Left Table          Right Table
   +----+------+       +----+--------+
   | id | name |       | id | score  |
   +----+------+       +----+--------+
   | 1  | Ade  |       | 1  | 85     |
   | 2  | Bola |       | 3  | 90     |
   +----+------+       +----+--------+

   left_join on id:
   +----+------+--------+
   | id | name | score  |
   +----+------+--------+
   | 1  | Ade  | 85     |
   | 2  | Bola | NA     |
   +----+------+--------+

Comparison Tables

Join Types Compared
JoinRows KeptWhen to Use
inner_joinOnly matchesComplete cases only
left_joinAll left + matchesKeep main table
right_joinAll right + matchesKeep secondary table
full_joinAll rows from bothUnion of data

End-of-Module Summary

In this module, we learned how to join tables using dplyr. We discovered that a key is a column that links two tables. We explored four types of joins: inner_join() (only matching rows), left_join() (all rows from the left table), right_join() (all rows from the right table), and full_join() (all rows from both tables). We also learned how to handle duplicate keys, join on different column names, and check our results. Joins are a powerful way to combine information from multiple sources, which is essential in data analysis.


Frequently Asked Questions

1. What is a join in R?
It combines two tables based on a common column.
2. What is a key?
A column used to match rows between tables.
3. What does inner_join do?
Keeps only rows that match in both tables.
4. What does left_join do?
Keeps all rows from the left table.
5. When should I use right_join?
When you want to keep all rows from the right table.
6. What is full_join?
Keeps all rows from both tables.
7. What if the key columns have different names?
Use by = c("name_in_A" = "name_in_B").
8. What are duplicate keys?
When a key appears more than once in a table.
9. How do I remove duplicates before joining?
Use distinct().
10. Can I join more than two tables?
Yes, by chaining joins with pipes.

Review Questions

  1. What is a join?
  2. What is a key?
  3. Which join keeps only matching rows?
  4. Which join keeps all rows from the left table?
  5. Which join keeps all rows from the right table?
  6. Which join keeps all rows from both tables?
  7. What happens if there are duplicate keys?
  8. How do you specify different key names in a join?
  9. What does NA mean after a join?
  10. Why should you check the result after a join?
  11. What is a composite key?
  12. Give an example of a left join in real life.
  13. How do you remove duplicate keys?
  14. What is the most commonly used join?
  15. Can you join tables from different datasets?

Fill-in-the-Blank Exercises

  1. A ______ combines two tables based on a common column. (join)
  2. The common column used to match rows is called a ______. (key)
  3. ______ keeps only rows that match in both tables. (inner_join)
  4. ______ keeps all rows from the left table. (left_join)
  5. ______ keeps all rows from the right table. (right_join)
  6. ______ keeps all rows from both tables. (full_join)
  7. If a key appears more than once, it is called a ______. (duplicate)
  8. When there is no match, the join fills with ______. (NA)
  9. Use ______ to remove duplicate rows. (distinct)
  10. To join on different column names, use the ______ argument. (by)

True or False Exercises

  1. A join combines two tables. (True)
  2. A key is not needed for a join. (False)
  3. inner_join keeps all rows from both tables. (False)
  4. left_join keeps all rows from the left table. (True)
  5. right_join keeps all rows from the right table. (True)
  6. full_join keeps only matching rows. (False)
  7. Duplicate keys can cause more rows in the result. (True)
  8. NA means that a match was found. (False)
  9. You can join more than two tables. (True)
  10. The by argument is used to specify the key. (True)

Multiple Choice Questions

  1. Which join keeps only rows that match in both tables?
    A. left_join B. right_join C. inner_join D. full_join
    Answer: C
  2. Which join keeps all rows from the left table?
    A. left_join B. inner_join C. right_join D. full_join
    Answer: A
  3. What is the common column used to match rows called?
    A. index B. key C. value D. row
    Answer: B
  4. What does full_join do?
    A. Keeps only matches B. Keeps all left and matches right C. Keeps all rows from both D. Keeps all right and matches left
    Answer: C
  5. If key names differ, which argument do you use?
    A. by B. on C. key D. match
    Answer: A
  6. What is a duplicate key?
    A. A key that appears once B. A key that appears multiple times C. A missing key D. A different key
    Answer: B
  7. What does NA mean in a join result?
    A. Match found B. No match found C. Error D. Duplicate
    Answer: B
  8. Which function removes duplicate rows?
    A. unique() B. distinct() C. filter() D. select()
    Answer: B
  9. Can you join tables from different sources?
    A. Yes B. No C. Only if they are CSV D. Only if they are in R
    Answer: A
  10. Which join is most commonly used?
    A. inner_join B. left_join C. right_join D. full_join
    Answer: B
  11. What is a composite key?
    A. A single column key B. A key with multiple columns C. A key with different names D. A missing key
    Answer: B
  12. What should you do after a join?
    A. Nothing B. Check the result C. Delete the tables D. Rename columns
    Answer: B
  13. Which join keeps all rows from the right table?
    A. left_join B. right_join C. inner_join D. full_join
    Answer: B
  14. How do you chain multiple joins?
    A. Using pipes B. Using commas C. Using semicolons D. Using brackets
    Answer: A
  15. What is the purpose of a join?
    A. To combine tables B. To delete tables C. To sort tables D. To filter tables
    Answer: A

Matching Exercises

JoinDescription
1. inner_joinA. All rows from left + matches
2. left_joinB. All rows from both
3. right_joinC. Only matching rows
4. full_joinD. All rows from right + matches

Answers: 1-C, 2-A, 3-D, 4-B


Short Answer Questions

  1. What is the purpose of a join?
  2. Explain the difference between inner_join and left_join.
  3. What is a key and why is it important?
  4. How do you handle duplicate keys before joining?
  5. What does NA represent in a joined table?

Scenario-based Exercises

  1. Scenario: You have a table of students (id, name) and a table of test scores (id, score). You want to create a table with all students and their scores, but some students may not have scores. Which join do you use?
  2. Scenario: You have a table of products and a table of sales. You want to know which products have been sold. Which join would you use to see only sold products?
  3. Scenario: You have a table of employees and a table of departments. You want to see all employees, even if they are not assigned to a department. Which join do you use?

Group Activity

In groups, create two small datasets (e.g., pupils and their classes, and pupils and their favourite subjects). Use all four types of joins and discuss the differences in the results. Present your findings to the class.


Individual Activity

Using the mtcars dataset and a small custom dataset (e.g., car names and their country of origin), perform a left join and a full join. Write down the differences in the resulting tables.


Classroom Discussion Questions

  1. Why are joins important in data analysis?
  2. What could go wrong if you join on the wrong column?
  3. How can joins help in Nigerian government data analysis?
  4. Which join do you think you will use most often and why?

Mini Project

Title: "Combining School Data"
You have two tables: students (id, name, class) and grades (id, subject, score). Perform joins to create a complete dataset that shows each student's name, class, subject, and score. Then, summarise the average score by class using group_by() and summarise().


Practical Assignment

Using the nycflights13 package (flights data), join the flights table with the airlines table to get the airline names for each flight. Then, join with the airports table to get the destination airport names. Write the code and show the first few rows.


Challenge Exercise

Find two real datasets online (e.g., population and GDP for Nigerian states). Perform a join to combine them, then create a visualisation (bar chart) showing the relationship. Write a short report explaining your steps.


Quiz Answers

Fill-in-the-Blank: 1. join, 2. key, 3. inner_join, 4. left_join, 5. right_join, 6. full_join, 7. duplicate, 8. NA, 9. distinct, 10. by.

True/False: 1T, 2F, 3F, 4T, 5T, 6F, 7T, 8F, 9T, 10T.

Multiple Choice: 1C, 2A, 3B, 4C, 5A, 6B, 7B, 8B, 9A, 10B, 11B, 12B, 13B, 14A, 15A.


Key Takeaways

  • Joins combine tables using a key.
  • inner_join keeps only matching rows.
  • left_join keeps all rows from the left table.
  • right_join keeps all rows from the right table.
  • full_join keeps all rows from both tables.
  • Always check for duplicate keys and NA values.
  • Joins are essential for combining data from multiple sources.

Preparation for the Next Module

In Module 10, we will learn about tidy data and the tidyr package. We will reshape data, convert between wide and long formats, and use functions like pivot_longer() and pivot_wider(). This will help us prepare data for analysis and visualisation. Practise joining tables and reflect on how combining data makes it more powerful.


11

Module Ten

Module 10 Β· R for Data Analysis Β· Tidy Data with tidyr

Module 10 Β· Tidy Data – Reshaping and Organising with tidyr

Hello, young data explorer! In Module 9, we learned how to join tables to bring data together. Now, we will learn how to reshape data so it is in the best format for analysis. Sometimes data is wide (many columns) and sometimes it is long (many rows). We need to choose the right shape for our work.

This is called tidy data. Tidy data has a simple rule: each variable is a column, each observation is a row, and each value is a cell. But real data is often messy. We will use the tidyr package to clean and reshape data.

By the end of this module, you will be able to convert data between wide and long formats, split columns, and handle missing values. You will be a data shaper!


Learning Objectives

  • Understand what tidy data means.
  • Install and load the tidyr package.
  • Use pivot_longer() to convert wide data to long data.
  • Use pivot_wider() to convert long data to wide data.
  • Use separate() to split one column into multiple columns.
  • Use unite() to combine multiple columns into one.
  • Use drop_na() to remove rows with missing values.
  • Use fill() to fill missing values with the previous value.
  • Apply these tools to real data, including Nigerian examples.
  • Choose the appropriate shape for your analysis.

Warm-up Story: The Messy Score Sheet

At Sunshine Primary School, the teachers recorded test scores in a wide format:

   name    Math    English    Science
   Ade     85      90         78
   Bola    70      65         80
   Chidi   95      88         92

This is easy to read, but it is hard to analyse because subjects are in columns. If we want to calculate the average score per subject, we have to do extra work. The data is not tidy because the subjects should be in a column, not spread across columns.

So, the teacher used pivot_longer() to reshape it into a long format:

   name    subject   score
   Ade     Math      85
   Ade     English   90
   Ade     Science   78
   Bola    Math      70
   ...

Now, each row is one observation (a pupil and a subject). It is easy to calculate averages, make graphs, and analyse. The teacher was happy because the data was now tidy!


Main Lessons

Lesson 1: What is Tidy Data?

Definition: Tidy data is data that follows three rules: (1) each variable is a column, (2) each observation is a row, and (3) each value is a cell.

Why it is important: Tidy data is easy to analyse, visualise, and model. Most R tools expect data in a tidy format.

Simple explanation: It's like organising your bookshelf: every book has a title (variable), and each book is on its own (observation).

Real-life example: A grocery list where each item is a row and columns are item, quantity, and price.

School example: A class register where each pupil is a row and columns are name, age, and grade.

Home example: A list of chores where each chore is a row and columns are chore, person, and time.

Nigerian example: A government dataset where each state is a row and columns are state, population, and region.

Illustration:

   Tidy data:
   +------+-------+--------+
   | name | class | score  |
   +------+-------+--------+
   | Ade  | 5     | 85     |
   | Bola | 6     | 70     |
   +------+-------+--------+
   Each variable (name, class, score) is a column.
   Each observation (pupil) is a row.

Mini summary: Tidy data has one variable per column, one observation per row.


Lesson 2: Wide vs Long Data

Definition: Wide data has many columns, where each category is a separate column. Long data has fewer columns, but more rows, with one column for category names and one for values.

Why it is important: Different analyses require different shapes. Graphs often need long data, while reports often need wide data.

Simple explanation: Wide is like a spreadsheet with many columns; long is like a database with many rows.

Real-life example: Wide: monthly sales for each month as a column. Long: a column for month and a column for sales.

School example: Wide: scores for each subject as columns. Long: subject and score columns.

Home example: Wide: each family member has a column for each chore. Long: columns for person, chore, and done?

Nigerian example: Wide: population for each state as separate columns. Long: state and population columns.

Illustration:

   Wide (not tidy):
   +------+------+----------+
   | name | Math | English  |
   +------+------+----------+
   | Ade  | 85   | 90       |
   +------+------+----------+

   Long (tidy):
   +------+---------+-------+
   | name | subject | score |
   +------+---------+-------+
   | Ade  | Math    | 85    |
   | Ade  | English | 90    |
   +------+---------+-------+

Mini summary: Wide has many columns; long has many rows. Tidy data is often long.


Lesson 3: Meet tidyr – The Reshaping Package

Definition: tidyr is an R package that helps you reshape and organise data to make it tidy.

Why it is important: It provides functions like pivot_longer(), pivot_wider(), separate(), and unite().

Simple explanation: Think of tidyr as a clothes folder – it folds your data into the right shape.

Real-life example: A tailor reshapes cloth to make a shirt – tidyr reshapes data for analysis.

School example: You rearrange your school bag to fit everything neatly.

Home example: You fold clothes to fit in the wardrobe.

Nigerian example: A farmer arranges yams in rows and columns for storage.

   install.packages("tidyr")
   library(tidyr)

Mini summary: tidyr helps you change the shape of your data.


Lesson 4: pivot_longer() – From Wide to Long

Definition: pivot_longer() collapses multiple columns into two columns: one for the column names (keys) and one for the values.

Why it is important: It makes data tidy by turning column names into a variable.

Simple explanation: It's like taking a table with subjects as columns and stacking them into a single column.

Real-life example: You have sales data for each month in separate columns. You pivot to a month column and sales column.

School example: You have test scores for Math, English, Science as columns. You pivot to subject and score columns.

Home example: You have a list of expenses for each day as columns. You pivot to day and expense columns.

Nigerian example: You have population data for each state as columns. You pivot to state and population.

Code:

   # Wide data
   scores_wide <- data.frame(
       name = c("Ade", "Bola"),
       Math = c(85, 70),
       English = c(90, 65)
   )

   # Pivot to long
   scores_long <- scores_wide %>%
       pivot_longer(cols = c(Math, English),
                    names_to = "subject",
                    values_to = "score")
   # Result:
   # name  subject  score
   # Ade   Math     85
   # Ade   English  90
   # Bola  Math     70
   # Bola  English  65
   Wide:
   +------+------+---------+
   | name | Math | English |
   +------+------+---------+
   | Ade  | 85   | 90      |
   +------+------+---------+

   Long (after pivot_longer):
   +------+---------+-------+
   | name | subject | score |
   +------+---------+-------+
   | Ade  | Math    | 85    |
   | Ade  | English | 90    |
   +------+---------+-------+

Mini summary: pivot_longer() makes data longer by stacking columns.


Lesson 5: pivot_wider() – From Long to Wide

Definition: pivot_wider() spreads a column into multiple columns, making data wider.

Why it is important: Sometimes we need data in a wider format for reporting or presentation.

Simple explanation: It's like taking a column of categories and making them into separate columns.

Real-life example: You have a column for months and a column for sales. You want months as columns.

School example: You have subject and score columns. You want subjects as columns for a report card.

Home example: You have chore and time columns. You want each chore as a column.

Nigerian example: You have state and population columns. You want each state as a column.

Code:

   # Long data
   scores_long <- data.frame(
       name = c("Ade", "Ade", "Bola", "Bola"),
       subject = c("Math", "English", "Math", "English"),
       score = c(85, 90, 70, 65)
   )

   # Pivot to wide
   scores_wide <- scores_long %>%
       pivot_wider(names_from = subject,
                   values_from = score)
   # Result: name, Math, English columns
   Long:
   +------+---------+-------+
   | name | subject | score |
   +------+---------+-------+
   | Ade  | Math    | 85    |
   | Ade  | English | 90    |
   +------+---------+-------+

   Wide (after pivot_wider):
   +------+------+---------+
   | name | Math | English |
   +------+------+---------+
   | Ade  | 85   | 90      |
   +------+------+---------+

Mini summary: pivot_wider() makes data wider by spreading a column.


Lesson 6: separate() – Splitting a Column

Definition: separate() splits a single column into multiple columns based on a separator (like a comma or space).

Why it is important: Sometimes a column contains combined information (e.g., "first_name last_name").

Simple explanation: It's like cutting a sandwich into two halves.

Real-life example: Splitting full name into first name and last name.

School example: Splitting "class_roll" into class and roll number.

Home example: Splitting "meal_time" into meal and time.

Nigerian example: Splitting "state_capital" into state and capital.

Code:

   # Column: full_name = "Ade Bola"
   data <- data.frame(full_name = c("Ade Bola", "Chidi Ngozi"))
   data %>%
       separate(full_name, into = c("first", "last"), sep = " ")
   # Result: first = "Ade", last = "Bola"
   Before:
   +------------+
   | full_name  |
   +------------+
   | Ade Bola   |
   +------------+

   After separate:
   +-------+------+
   | first | last |
   +-------+------+
   | Ade   | Bola |
   +-------+------+

Mini summary: separate() splits a column into two or more columns.


Lesson 7: unite() – Combining Columns

Definition: unite() combines multiple columns into one column.

Why it is important: Sometimes we need to merge information for a unique identifier.

Simple explanation: It's like gluing two pieces of paper together.

Real-life example: Combining first name and last name into full name.

School example: Combining class and roll number into "class_roll".

Home example: Combining meal and time into "meal_time".

Nigerian example: Combining state and region into "state_region".

Code:

   data <- data.frame(first = c("Ade", "Chidi"),
                      last = c("Bola", "Ngozi"))
   data %>%
       unite(full_name, first, last, sep = " ")
   # Result: full_name = "Ade Bola"
   Before:
   +-------+-------+
   | first | last  |
   +-------+-------+
   | Ade   | Bola  |
   +-------+-------+

   After unite:
   +------------+
   | full_name  |
   +------------+
   | Ade Bola   |
   +------------+

Mini summary: unite() combines columns into one.


Lesson 8: drop_na() – Removing Missing Values

Definition: drop_na() removes rows that contain any missing values (NA).

Why it is important: Missing values can cause errors in analysis. Removing them is a quick fix.

Simple explanation: It's like taking out the rotten apples from a basket.

Real-life example: A survey respondent didn't answer a question – you drop that row.

School example: A pupil was absent, so you drop that score row.

Home example: You forgot to record a chore, so you drop that row.

Nigerian example: A state's data is missing, so you drop that row.

Code:

   data <- data.frame(name = c("Ade", "Bola", "Chidi"),
                      score = c(85, NA, 90))
   data %>% drop_na()
   # Result: only Ade and Chidi (Bola removed)
   Before:
   +-------+-------+
   | name  | score |
   +-------+-------+
   | Ade   | 85    |
   | Bola  | NA    |
   | Chidi | 90    |
   +-------+-------+

   After drop_na():
   +-------+-------+
   | name  | score |
   +-------+-------+
   | Ade   | 85    |
   | Chidi | 90    |
   +-------+-------+

Mini summary: drop_na() removes rows with missing values.


Lesson 9: fill() – Filling Missing Values

Definition: fill() fills missing values (NA) with the previous or next value in the column.

Why it is important: Sometimes missing values are just gaps that can be filled logically.

Simple explanation: It's like copying the last known value to fill a blank.

Real-life example: A time series where missing days can be filled with the previous day's value.

School example: If a pupil's class is missing, you can fill with the class from the previous row.

Home example: If you forget to record your expense for a day, you can use the previous day.

Nigerian example: If a state's region is missing, fill with the region from a nearby state.

Code:

   data <- data.frame(
       day = c(1, 2, 3, 4),
       sales = c(100, NA, NA, 150)
   )
   data %>% fill(sales, .direction = "down")
   # Result: 100, 100, 100, 150
   Before:
   +-----+-------+
   | day | sales |
   +-----+-------+
   | 1   | 100   |
   | 2   | NA    |
   | 3   | NA    |
   | 4   | 150   |
   +-----+-------+

   After fill(down):
   +-----+-------+
   | day | sales |
   +-----+-------+
   | 1   | 100   |
   | 2   | 100   |
   | 3   | 100   |
   | 4   | 150   |
   +-----+-------+

Mini summary: fill() replaces missing values with nearby values.


Lesson 10: Real Nigerian Example – Reshaping Population Data

Let's reshape Nigerian population data from wide to long.

   # Wide data: columns for each year
   pop_wide <- data.frame(
       state = c("Lagos", "Kano", "Oyo"),
       X2020 = c(20, 15, 8),
       X2021 = c(21, 16, 9)
   )

   # Pivot to long
   pop_long <- pop_wide %>%
       pivot_longer(cols = c(X2020, X2021),
                    names_to = "year",
                    values_to = "population")
   # Result: state, year, population
   Before (wide):
   +-------+------+------+
   | state | 2020 | 2021 |
   +-------+------+------+
   | Lagos | 20   | 21   |
   +-------+------+------+

   After (long):
   +-------+------+------------+
   | state | year | population |
   +-------+------+------------+
   | Lagos | 2020 | 20         |
   | Lagos | 2021 | 21         |
   +-------+------+------------+

Mini summary: pivot_longer() works on any data, including Nigerian data.


Lesson 11: Combining tidyr with dplyr

You can chain tidyr functions with dplyr verbs using pipes.

   data %>%
       pivot_longer(cols = c(Math, English), names_to = "subject", values_to = "score") %>%
       filter(score > 70) %>%
       arrange(desc(score))

Mini summary: tidyr and dplyr work together beautifully.


Lesson 12: Choosing Between Wide and Long

When to use wide: For human-readable tables, reports, and some visualisations.

When to use long: For analysis, modelling, and many ggplot2 graphs.

Mini summary: Use wide for presentation, long for analysis.


Lesson 13: Handling Duplicate Rows in Long Data

After pivoting, you might get duplicate rows. Use distinct() to remove them.

   data %>% distinct()

Mini summary: Check for duplicates after reshaping.


Lesson 14: Summary of tidyr Functions

  • pivot_longer() – wide to long
  • pivot_wider() – long to wide
  • separate() – split one column
  • unite() – combine columns
  • drop_na() – remove missing rows
  • fill() – fill missing values

Lesson 15: Final Thoughts on Tidy Data

Tidy data is the foundation of data analysis in R. Once your data is tidy, everything else becomes easier.


Key Vocabulary

Tidy Data
Data where each variable is a column, each observation is a row.
Wide Data
Data with many columns, each representing a variable.
Long Data
Data with fewer columns but more rows, with a key-value structure.
pivot_longer()
Function to make data longer.
pivot_wider()
Function to make data wider.
separate()
Function to split a column.
unite()
Function to combine columns.
drop_na()
Function to remove rows with missing values.
fill()
Function to fill missing values.

Important Concepts

  • Tidy data is the standard for R analysis.
  • Wide is good for humans; long is good for machines.
  • Pivoting changes the shape of data.
  • Missing values can be removed or filled.
  • Chaining with pipes makes code clear.

Step-by-Step Explanations

How to convert wide to long:

  1. Load tidyr and dplyr.
  2. Identify the columns you want to pivot.
  3. Use pivot_longer(cols = c(col1, col2), names_to = "key", values_to = "value").
  4. Check the result with head().

How to handle missing values:

  1. Use drop_na() to remove rows.
  2. Use fill() to propagate values.
  3. Decide based on your analysis.

Real-life Examples

  • A business converts monthly sales from columns to rows for trend analysis.
  • A hospital reshapes patient data to long format for modelling.
  • A school converts test scores from wide to long for summary statistics.

Nigerian Examples

  • Reshape state population by year from wide to long.
  • Split full names of Nigerian cities into city and state.
  • Combine local government and state into one column.

Fun Examples Children Relate To

  • Pivot a list of your favourite toys (wide: each toy as a column) to long (toy and rating).
  • Separate your full name into first and last.
  • Fill missing values in your daily reading time.

Everyday Examples

  • You have a table of meals for each day (wide) and you pivot to meal and day (long).
  • You have a list of chores and you unite person and chore into one column.
  • You drop rows where you forgot to record time.

Teacher Notes

  • Use simple datasets that students can relate to.
  • Draw wide and long formats on the board.
  • Emphasise that tidy data is the key to using R effectively.
  • Show how pivot_longer helps with ggplot2.

Parent Tips

  • Help your child collect data in both wide and long formats.
  • Ask them to convert between formats by hand to understand the concept.
  • Discuss how shape affects what you can see in the data.

Interesting Facts

  • The concept of tidy data was introduced by Hadley Wickham (who also created ggplot2 and dplyr).
  • Most data in the world is not tidy – that's why data scientists are in demand!
  • tidyr is used by thousands of companies worldwide.

Did You Know?

  • You can use pivot_longer() with multiple columns at once.
  • There is a pivot_longer_spec() for advanced pivoting.
  • tidyr also has complete() to fill missing combinations.

Remember This

  • Tidy data is the gold standard.
  • pivot_longer() makes data longer; pivot_wider() makes it wider.
  • Use separate() and unite() for columns.
  • Handle NA with drop_na() or fill().
  • Always check your reshaped data.

Common Mistakes

  • Forgetting to load tidyr.
  • Specifying the wrong columns in pivot_longer().
  • Not using names_to and values_to correctly.
  • Using pivot_wider() when pivot_longer() is needed.
  • Not checking for duplicates after pivoting.

Best Practices

  • Start with tidy data for all analysis.
  • Use clear names for the new columns (names_to, values_to).
  • Chain operations with pipes for readability.
  • Use glimpse() to inspect your data after reshaping.
  • Document why you chose a particular shape.

ASCII Illustrations

Wide to Long

   Wide:
   +------+------+------+
   | name | Math | Eng  |
   +------+------+------+
   | Ade  | 85   | 90   |
   +------+------+------+

   pivot_longer:
   +------+---------+-------+
   | name | subject | score |
   +------+---------+-------+
   | Ade  | Math    | 85    |
   | Ade  | English | 90    |
   +------+---------+-------+

Long to Wide

   Long:
   +------+---------+-------+
   | name | subject | score |
   +------+---------+-------+
   | Ade  | Math    | 85    |
   | Ade  | English | 90    |
   +------+---------+-------+

   pivot_wider:
   +------+------+---------+
   | name | Math | English |
   +------+------+---------+
   | Ade  | 85   | 90      |
   +------+------+---------+

Comparison Tables

Wide vs Long
FeatureWideLong
Number of columnsManyFew
Number of rowsFewMany
Best forReportingAnalysis
ExampleSubjects as columnsSubject column

End-of-Module Summary

In this module, we learned about tidy data and how to reshape data using the tidyr package. We can convert data from wide to long with pivot_longer() and from long to wide with pivot_wider(). We also learned to split columns with separate(), combine columns with unite(), and handle missing values with drop_na() and fill(). Tidy data is essential for analysis and visualisation in R. Remember, the shape of your data matters!


Frequently Asked Questions

1. What is tidy data?
Data where each variable is a column, each observation is a row.
2. What is the difference between wide and long data?
Wide has many columns; long has many rows.
3. What does pivot_longer do?
Converts wide data to long data.
4. What does pivot_wider do?
Converts long data to wide data.
5. When should I use wide data?
For reports and human-readable tables.
6. When should I use long data?
For analysis and visualisation.
7. What is separate used for?
Splitting a column into multiple columns.
8. What is unite used for?
Combining multiple columns into one.
9. How do I remove rows with missing values?
Use drop_na().
10. How do I fill missing values?
Use fill().

Review Questions

  1. What is tidy data?
  2. What is the difference between wide and long data?
  3. Which function converts wide to long?
  4. Which function converts long to wide?
  5. What does separate do?
  6. What does unite do?
  7. How do you remove rows with missing values?
  8. How do you fill missing values?
  9. When should you use long data?
  10. When should you use wide data?
  11. What is the tidyr package used for?
  12. How do you pivot multiple columns?
  13. What is the names_to argument in pivot_longer?
  14. What is the values_to argument in pivot_longer?
  15. Why is tidy data important?

Fill-in-the-Blank Exercises

  1. ______ data has each variable as a column and each observation as a row. (Tidy)
  2. ______ data has many columns. (Wide)
  3. ______ data has many rows. (Long)
  4. ______ converts wide to long. (pivot_longer)
  5. ______ converts long to wide. (pivot_wider)
  6. ______ splits a column. (separate)
  7. ______ combines columns. (unite)
  8. ______ removes rows with missing values. (drop_na)
  9. ______ fills missing values. (fill)
  10. The package for reshaping data is ______. (tidyr)

True or False Exercises

  1. Tidy data has one variable per column. (True)
  2. Wide data is always better than long data. (False)
  3. pivot_longer makes data wider. (False)
  4. pivot_wider makes data longer. (False)
  5. separate splits a column into multiple columns. (True)
  6. unite combines columns into one. (True)
  7. drop_na removes rows with missing values. (True)
  8. fill replaces missing values with the previous value. (True)
  9. tidyr is used for data visualisation. (False)
  10. Long data is usually better for analysis. (True)

Multiple Choice Questions

  1. Which package is used for reshaping data?
    A. ggplot2 B. dplyr C. tidyr D. readr
    Answer: C
  2. Which function converts wide to long?
    A. pivot_wider B. pivot_longer C. separate D. unite
    Answer: B
  3. Which function converts long to wide?
    A. pivot_wider B. pivot_longer C. separate D. unite
    Answer: A
  4. What does separate do?
    A. Combines columns B. Splits a column C. Removes rows D. Fills missing values
    Answer: B
  5. What does unite do?
    A. Splits a column B. Combines columns C. Removes rows D. Fills missing values
    Answer: B
  6. Which function removes rows with missing values?
    A. drop_na B. fill C. separate D. unite
    Answer: A
  7. Which function fills missing values?
    A. drop_na B. fill C. separate D. unite
    Answer: B
  8. What is the names_to argument used for?
    A. Values B. Column names to be stacked C. New column name for keys D. New column name for values
    Answer: C
  9. What is the values_to argument used for?
    A. Values B. Column names to be stacked C. New column name for keys D. New column name for values
    Answer: D
  10. When should you use long data?
    A. For reports B. For analysis C. For printing D. For storage
    Answer: B
  11. When should you use wide data?
    A. For analysis B. For reporting C. For modelling D. For graphing
    Answer: B
  12. What is tidy data?
    A. Each variable is a column B. Each observation is a row C. Each value is a cell D. All of the above
    Answer: D
  13. Which function can handle missing values?
    A. drop_na B. fill C. Both A and B D. Neither
    Answer: C
  14. Can you chain tidyr functions with pipes?
    A. Yes B. No C. Only with dplyr D. Only with ggplot2
    Answer: A
  15. What does the sep argument in separate do?
    A. Specifies separator B. Specifies columns C. Specifies values D. Specifies names
    Answer: A

Matching Exercises

FunctionAction
1. pivot_longerA. Combines columns
2. pivot_widerB. Splits a column
3. separateC. Wide to long
4. uniteD. Long to wide
5. drop_naE. Removes missing rows

Answers: 1-C, 2-D, 3-B, 4-A, 5-E


Short Answer Questions

  1. What is tidy data?
  2. Explain the difference between wide and long data.
  3. What does pivot_longer do? Give an example.
  4. How do you handle missing values in tidyr?
  5. Why is tidy data important for analysis?

Scenario-based Exercises

  1. Scenario: You have a dataset with columns: name, subject1_score, subject2_score, subject3_score. You want to analyse scores by subject. What do you do?
  2. Scenario: You have a long dataset with name, subject, score. You want to create a report card with subjects as columns. What do you do?
  3. Scenario: Your data has a column "full_name" with first and last names together. You need to separate them. Which function do you use?

Group Activity

In groups, create a wide dataset (e.g., favourite foods with ratings for each family member). Use pivot_longer() to make it long, then pivot_wider() to make it wide again. Discuss the differences.


Individual Activity

Use the iris dataset, but modify it to have wide format (one column per species for Sepal.Length). Then use pivot_longer() to restore it to long format.


Classroom Discussion Questions

  1. Why is tidy data important for data science?
  2. When would you choose wide over long, and vice versa?
  3. How does tidyr help in preparing data for ggplot2?
  4. What are the challenges of reshaping data?

Mini Project

Title: "Reshaping Nigerian Education Data"
Create a wide dataset of test scores for pupils in different subjects. Use pivot_longer() to tidy it. Then, use group_by() and summarise() to find the average score per subject. Present your findings.


Practical Assignment

Using the economics dataset (built-in), which is in long format, use pivot_wider() to make it wide with variables as columns. Then use pivot_longer() to return to long.


Challenge Exercise

Find a messy dataset online (e.g., from Kaggle). Use tidyr and dplyr to clean and reshape it into a tidy format. Write a short report on your steps.


Quiz Answers

Fill-in-the-Blank: 1. Tidy, 2. Wide, 3. Long, 4. pivot_longer, 5. pivot_wider, 6. separate, 7. unite, 8. drop_na, 9. fill, 10. tidyr.

True/False: 1T, 2F, 3F, 4F, 5T, 6T, 7T, 8T, 9F, 10T.

Multiple Choice: 1C, 2B, 3A, 4B, 5B, 6A, 7B, 8C, 9D, 10B, 11B, 12D, 13C, 14A, 15A.


Key Takeaways

  • Tidy data is the foundation of R analysis.
  • pivot_longer() and pivot_wider() reshape data.
  • separate() and unite() handle columns.
  • drop_na() and fill() manage missing values.
  • Choose the shape that fits your analysis.

Preparation for the Next Module

In Module 11, we will learn about data import – reading data from different file types like CSV, Excel, and databases. We will use packages like readr, readxl, and haven. Practise reshaping data with tidyr and reflect on how tidy data makes analysis easier.


12

Module Eleven

Module 11 Β· R for Data Analysis Β· Data Import and Export

Module 11 Β· Data Import and Export – Reading and Writing Files

Hello, young data explorer! In Module 10, we learned how to reshape data using tidyr. But where does data come from? Usually, data is stored in files on your computer. These files can be in many formats: CSV (comma-separated values), Excel, or even databases.

In this module, we will learn how to import (read) data into R from different file types, and export (write) data from R to files. This is like opening a book (reading) and writing your own book (saving). We will use packages like readr, readxl, and haven to handle different formats.

By the end of this module, you will be able to load any dataset into R and save your results for sharing. You will be a data librarian!


Learning Objectives

  • Understand why we need to import and export data.
  • Identify common file formats (CSV, Excel, text, etc.).
  • Install and load the readr package.
  • Use read_csv() to read CSV files.
  • Use write_csv() to write CSV files.
  • Install and load the readxl package.
  • Use read_excel() to read Excel files.
  • Learn about other formats: read_delim(), read_table().
  • Understand the working directory and file paths.
  • Apply these skills to real data, including Nigerian examples.

Warm-up Story: The Lost Data File

Chidi was doing a school project on Nigerian states. He had collected data in a CSV file on his computer. He opened R and typed read_csv("states.csv"), but R gave an error: "File not found." Chidi was confused. He realised he needed to tell R exactly where the file was on his computer – the path. He set his working directory to the folder containing the file, and it worked!

Later, he analysed the data and created a summary. He wanted to share his results with his teacher, so he used write_csv() to save the summary as a new CSV file. He emailed it to his teacher. Chidi learned that importing and exporting data is like opening and saving files – it's essential for working with real data.


Main Lessons

Lesson 1: What is Data Import and Export?

Definition: Import means reading data from a file into R. Export means writing data from R to a file.

Why it is important: Data lives in files (CSV, Excel, etc.). To analyse it, we must bring it into R. After analysis, we save results for sharing.

Simple explanation: It's like reading a book (import) and writing your own book (export).

Real-life example: A shopkeeper reads a price list from a spreadsheet (import) and writes a sales report (export).

School example: Your teacher reads a class list from a file (import) and saves grades (export).

Home example: You open a recipe from a file (import) and save your grocery list (export).

Nigerian example: A researcher reads population data from a CSV (import) and saves a summary report (export).

Illustration:

   File (on computer)  ----Import--->  R (data frame)
   R (data frame)       ----Export--->  File (saved)

Mini summary: Import = read into R; Export = save from R.


Lesson 2: Common File Formats

Definition: File formats are ways of storing data. Common ones are:

  • CSV (Comma-Separated Values): text file with commas.
  • Excel (.xlsx, .xls): spreadsheet files.
  • TSV (Tab-Separated Values): text file with tabs.
  • Text (.txt): plain text files.
  • RDS (.rds): R's own format.

Mini summary: CSV and Excel are most common for beginners.


Lesson 3: The Working Directory – Where R Looks for Files

Definition: The working directory is the folder where R looks for files.

Why it is important: If your file is not in the working directory, R won't find it.

Simple explanation: It's like the "home base" of your R session.

Real-life example: You keep your school books in a specific bag – you know where to find them.

School example: Your teacher tells you to open a file from the "Class Data" folder.

Home example: You keep your toys in a box – you know where to look.

Nigerian example: A researcher keeps all data files in a "Projects" folder.

Code:

   getwd()   # shows current working directory
   setwd("path/to/folder")  # changes it

Mini summary: Set your working directory to the folder containing your files.


Lesson 4: Reading CSV Files with readr

Definition: read_csv() is a function from the readr package that reads CSV files.

Why it is important: CSV is the most common data format.

Simple explanation: It's like opening a book – R reads the text and turns it into a table.

Real-life example: A scientist reads a CSV of weather data.

School example: A teacher reads a CSV of pupil scores.

Home example: You read a CSV of your family expenses.

Nigerian example: An economist reads a CSV of Nigeria's GDP data.

Code:

   install.packages("readr")
   library(readr)

   data <- read_csv("file.csv")
   # Or with a full path:
   data <- read_csv("C:/Users/YourName/Downloads/file.csv")

Illustration:

   CSV file (text):
   name,age,score
   Ade,10,85
   Bola,11,70

   read_csv() -> R data frame:
   +------+-----+-------+
   | name | age | score |
   +------+-----+-------+
   | Ade  | 10  | 85    |
   | Bola | 11  | 70    |
   +------+-----+-------+

Mini summary: read_csv() reads CSV files into R.


Lesson 5: Writing CSV Files with readr

Definition: write_csv() saves a data frame as a CSV file.

Why it is important: You can save your results and share them.

Simple explanation: It's like writing a story and saving it as a document.

Real-life example: A shop saves a sales summary as CSV.

School example: A teacher saves grades as CSV.

Home example: You save your chore chart as CSV.

Nigerian example: A researcher saves analysed data as CSV.

Code:

   write_csv(data, "output.csv")
   # Saves to working directory.

Mini summary: write_csv() saves data frames as CSV.


Lesson 6: Reading Excel Files with readxl

Definition: read_excel() from the readxl package reads Excel (.xlsx, .xls) files.

Why it is important: Excel is widely used in business and schools.

Simple explanation: It's like opening a spreadsheet in R.

Real-life example: A business reads an Excel budget file.

School example: A teacher reads an Excel attendance sheet.

Home example: You read an Excel list of family members.

Nigerian example: A government official reads an Excel file of state budgets.

Code:

   install.packages("readxl")
   library(readxl)

   data <- read_excel("file.xlsx")
   # Specify sheet if needed:
   data <- read_excel("file.xlsx", sheet = "Sheet1")

Mini summary: read_excel() reads Excel files.


Lesson 7: Writing Excel Files

To write Excel files, you can use the writexl package or openxlsx.

   install.packages("writexl")
   library(writexl)
   write_xlsx(data, "output.xlsx")

Mini summary: Use write_xlsx() to save as Excel.


Lesson 8: Reading Other Text Formats

Definition: read_delim() reads files with any separator (e.g., tabs, semicolons).

Real-life example: A tab-separated file.

School example: A file with pipe (|) separators.

Code:

   data <- read_delim("file.txt", delim = "\t")  # tab-separated
   data <- read_delim("file.txt", delim = ";")   # semicolon

Mini summary: read_delim() handles any delimiter.


Lesson 9: Handling File Paths

Definition: A file path tells R exactly where the file is located.

Why it is important: If you don't give the full path, R looks in the working directory.

Simple explanation: It's like giving an address to find a house.

Types:

  • Absolute path: full address (e.g., "C:/Users/Chidi/Data/file.csv")
  • Relative path: relative to working directory (e.g., "Data/file.csv")

Mini summary: Use full paths to avoid errors.


Lesson 10: Importing Data from the Web

You can read data directly from a URL.

   data <- read_csv("https://example.com/data.csv")

Mini summary: read_csv() works with web URLs too.


Lesson 11: Exporting Data for Sharing

After analysis, you might want to save your data in different formats.

   write_csv(data, "clean_data.csv")
   write_xlsx(data, "clean_data.xlsx")
   write_rds(data, "clean_data.rds")  # R's native format

Mini summary: Save your data in the format your audience needs.


Lesson 12: Reading Data from Databases

Definition: Databases are organised collections of data. R can connect to them using packages like DBI and odbc.

For beginners, we focus on files.

Mini summary: Databases are more advanced – we'll start with files.


Lesson 13: Common Errors and Solutions

  • File not found: Check the working directory or use full path.
  • Wrong separator: Use read_delim() with correct delim.
  • Header issues: Use col_names = TRUE or skip.
  • Encoding problems: Use locale() to specify encoding.

Mini summary: Most errors are due to file location or format.


Lesson 14: Real Nigerian Example – Importing Data

Let's say we have a CSV file "nigeria_population.csv" with columns: State, Population, Region.

   pop <- read_csv("nigeria_population.csv")
   head(pop)

Mini summary: Import Nigerian data easily with read_csv().


Lesson 15: Summary of Import/Export

We learned to read and write CSV, Excel, and other formats. We also learned about working directories and file paths.


Key Vocabulary

Import
Reading data into R from a file.
Export
Writing data from R to a file.
CSV
Comma-Separated Values – a text file format.
Excel
A spreadsheet file format (.xlsx).
Working Directory
The folder where R looks for files.
File Path
The location of a file on your computer.
readr
Package for reading text files.
readxl
Package for reading Excel files.

Important Concepts

  • Working directory is the default folder for files.
  • CSV is the most common data exchange format.
  • read_csv() is faster and better than base R's read.csv().
  • Relative paths are easier to use in projects.
  • Always check that your data imported correctly with head() or glimpse().

Step-by-Step Explanations

How to import a CSV file:

  1. Install readr: install.packages("readr").
  2. Load it: library(readr).
  3. Set working directory or use full path.
  4. Use read_csv("filename.csv").
  5. Assign to a variable: data <- read_csv("filename.csv").
  6. Check with head(data).

How to export a CSV:

  1. Prepare your data frame.
  2. Use write_csv(data, "output.csv").
  3. Check your working directory for the file.

Real-life Examples

  • A scientist reads weather data from a CSV file.
  • A teacher exports grades to Excel for the principal.
  • A shop imports inventory from a CSV.

Nigerian Examples

  • A researcher reads population data from NBS (National Bureau of Statistics) CSV files.
  • A student imports educational data from an Excel file.
  • A government analyst exports budget data to CSV for sharing.

Fun Examples Children Relate To

  • Read a CSV of your favourite cartoon characters and their ratings.
  • Export a list of your friends' birthdays to Excel.
  • Import a text file with your homework assignments.

Everyday Examples

  • You read a recipe from a text file.
  • You save a list of your chores to a CSV.
  • You open a spreadsheet of your savings.

Teacher Notes

  • Emphasise the importance of the working directory.
  • Show how to use relative paths with R projects.
  • Provide sample CSV files for practice.
  • Encourage students to explore file structures.

Parent Tips

  • Help your child create a folder for R projects.
  • Practice importing and exporting simple files together.
  • Discuss the importance of file organisation.

Interesting Facts

  • The CSV format has been around since the 1970s.
  • Excel files can contain multiple sheets.
  • R can read data from ZIP files directly.

Did You Know?

  • You can read data from Google Sheets into R using the googlesheets4 package.
  • R can read JSON and XML files too.
  • readr is part of the tidyverse.

Remember This

  • Always set your working directory or use full paths.
  • Use read_csv() for CSV files, read_excel() for Excel.
  • Check your imported data with head().
  • Save results with write_csv() or write_xlsx().

Common Mistakes

  • Forgetting to install/load the required package.
  • Using read.csv() instead of read_csv() (both work, but read_csv is faster).
  • Not specifying the full path when the file is not in the working directory.
  • Using the wrong separator (e.g., read_csv() for tab-separated files).
  • Not checking if the import was successful.

Best Practices

  • Use R projects to manage working directories automatically.
  • Store data files in a subfolder (e.g., "data/") to keep projects tidy.
  • Use readr functions for better performance.
  • Always specify the col_types argument for large files.
  • Document your data sources.

ASCII Illustrations

Import process

   File on disk  ----read_csv---->  R data frame
   (data.csv)                        (data)

Export process

   R data frame  ----write_csv---->  File on disk
   (data)                            (output.csv)

Comparison Tables

Import Functions
File TypeFunctionPackage
CSVread_csv()readr
Excelread_excel()readxl
Tab-separatedread_delim(file, delim = "\t")readr
RDSreadRDS()base R

End-of-Module Summary

In this module, we learned how to import (read) data from files into R and export (write) data from R to files. We used read_csv() and write_csv() for CSV files, and read_excel() for Excel files. We also learned about the working directory, file paths, and common errors. These skills are essential because most data comes from files, and we need to share our results.


Frequently Asked Questions

1. What is the difference between import and export?
Import reads data into R; export writes data from R.
2. What is a CSV file?
A text file with comma-separated values.
3. What package reads CSV files?
readr.
4. What package reads Excel files?
readxl.
5. What is the working directory?
The folder where R looks for files.
6. How do I check my working directory?
getwd().
7. How do I set my working directory?
setwd("path").
8. What is a file path?
The location of a file on your computer.
9. Can I read data from the internet?
Yes, read_csv("http://example.com/data.csv").
10. How do I save a data frame as CSV?
write_csv(data, "filename.csv").

Review Questions

  1. What is the difference between import and export?
  2. What is a CSV file?
  3. Which function reads CSV files?
  4. Which package reads Excel files?
  5. What is the working directory?
  6. How do you set the working directory?
  7. How do you save a data frame as CSV?
  8. What is a file path?
  9. What is the difference between absolute and relative paths?
  10. How do you check if a file was imported correctly?
  11. Can you read Excel files with readr?
  12. What is the delimiter in a CSV file?
  13. What does read_delim() do?
  14. What is an RDS file?
  15. Why is it important to set the working directory?

Fill-in-the-Blank Exercises

  1. ______ means reading data into R. (Import)
  2. ______ means writing data from R. (Export)
  3. ______ stands for Comma-Separated Values. (CSV)
  4. ______ is a package for reading CSV files. (readr)
  5. ______ is a package for reading Excel files. (readxl)
  6. The ______ is the folder where R looks for files. (working directory)
  7. Use ______ to set the working directory. (setwd)
  8. Use ______ to save a data frame as CSV. (write_csv)
  9. Use ______ to read an Excel file. (read_excel)
  10. A ______ is the location of a file on your computer. (file path)

True or False Exercises

  1. Import means writing data from R. (False)
  2. CSV files are text files. (True)
  3. read_csv() is from the readxl package. (False)
  4. read_excel() reads Excel files. (True)
  5. The working directory is the folder where R looks for files. (True)
  6. write_csv() saves data as Excel. (False)
  7. You can read files from the web. (True)
  8. RDS is an R-specific file format. (True)
  9. You don't need to check imported data. (False)
  10. A file path tells R where a file is. (True)

Multiple Choice Questions

  1. Which function reads CSV files?
    A. read_excel B. read_csv C. read_delim D. read_table
    Answer: B
  2. Which package is used for reading CSV files?
    A. readxl B. readr C. dplyr D. tidyr
    Answer: B
  3. Which package is used for reading Excel files?
    A. readr B. readxl C. dplyr D. tidyr
    Answer: B
  4. What is the separator in a CSV file?
    A. Tab B. Comma C. Semicolon D. Space
    Answer: B
  5. How do you set the working directory?
    A. getwd() B. setwd() C. dir() D. path()
    Answer: B
  6. How do you check the working directory?
    A. getwd() B. setwd() C. dir() D. path()
    Answer: A
  7. Which function writes a CSV file?
    A. write_excel B. write_csv C. write_delim D. write_table
    Answer: B
  8. What is a file path?
    A. The name of a file B. The location of a file C. The content of a file D. The size of a file
    Answer: B
  9. Can you read a file from a URL?
    A. Yes B. No C. Only with Excel D. Only with CSV
    Answer: A
  10. What is an RDS file?
    A. An Excel file B. A CSV file C. An R-specific file D. A text file
    Answer: C
  11. Which function reads tab-separated files?
    A. read_csv B. read_excel C. read_delim D. read_table
    Answer: C
  12. What does head() do?
    A. Shows the last rows B. Shows the first rows C. Shows the structure D. Shows the summary
    Answer: B
  13. Why is the working directory important?
    A. It saves files B. It tells R where to look for files C. It deletes files D. It changes file formats
    Answer: B
  14. What is the extension for CSV files?
    A. .txt B. .csv C. .xlsx D. .rds
    Answer: B
  15. Which package is part of the tidyverse?
    A. readxl B. readr C. both D. neither
    Answer: C

Matching Exercises

FunctionFile Type
1. read_csvA. Excel
2. read_excelB. CSV
3. read_delimC. Any delimiter
4. write_csvD. Save as CSV

Answers: 1-B, 2-A, 3-C, 4-D


Short Answer Questions

  1. What is the purpose of importing data?
  2. Explain the difference between CSV and Excel files.
  3. How do you set the working directory in R?
  4. What is the difference between absolute and relative paths?
  5. Why should you check your imported data?

Scenario-based Exercises

  1. Scenario: You have a CSV file "grades.csv" in your Downloads folder. You try to read it with read_csv("grades.csv") but get an error. What might be the problem?
  2. Scenario: You want to share your analysis results with a teacher who uses Excel. How would you save your results?
  3. Scenario: You have a tab-separated file. Which function would you use to read it?

Group Activity

In groups, create a small dataset (e.g., favourite foods with ratings). Save it as a CSV file. Then, each group member imports it into R and creates a summary. Export the summary as a new CSV file and share it with the group.


Individual Activity

Find a CSV file online (e.g., from a data portal). Import it into R, explore it with glimpse() and head(), and then save a filtered version as a new CSV file.


Classroom Discussion Questions

  1. Why is it important to organise files and folders?
  2. What are the advantages of using CSV over Excel?
  3. How would you import data from a website?
  4. What would you do if your file import fails?

Mini Project

Title: "Importing and Analysing Nigerian Data"
Find a CSV or Excel file about Nigeria (e.g., population, education, or weather). Import it into R, perform some basic analysis (summary, filtering), and export the results as a new file. Write a short report on what you found.


Practical Assignment

Using the iris dataset, save it as a CSV file using write_csv(). Then, import it back into R and verify it matches the original. Do the same with an Excel file using write_xlsx().


Challenge Exercise

Find a dataset in a format other than CSV (e.g., JSON, XML) and use appropriate R packages (e.g., jsonlite) to import it. Write a brief explanation of the process.


Quiz Answers

Fill-in-the-Blank: 1. Import, 2. Export, 3. CSV, 4. readr, 5. readxl, 6. working directory, 7. setwd, 8. write_csv, 9. read_excel, 10. file path.

True/False: 1F, 2T, 3F, 4T, 5T, 6F, 7T, 8T, 9F, 10T.

Multiple Choice: 1B, 2B, 3B, 4B, 5B, 6A, 7B, 8B, 9A, 10C, 11C, 12B, 13B, 14B, 15C.


Key Takeaways

  • Import = read data into R; Export = save data from R.
  • CSV is the most common format; use read_csv().
  • Excel files use read_excel().
  • Always set your working directory or use full paths.
  • Check your data with head() after import.
  • Use write_csv() to save your results.

Preparation for the Next Module

In Module 12, we will learn about data cleaning – handling missing values, outliers, and inconsistencies in detail. We will combine skills from Modules 8-11 to create a complete data cleaning workflow. Practise importing and exporting data, and think about how you would clean messy data.


13

Module Twelve

Module 12 Β· R for Data Analysis Β· Data Cleaning

Module 12 Β· Data Cleaning – Making Data Sparkling Clean

Hello, young data explorer! In Module 11, we learned how to import data from files. But what if the data we import is messy? It might have missing values, wrong data types, duplicates, or inconsistent spellings. We need to clean it before we can analyse it.

Data cleaning is like washing fruits before eating them – you want them to be fresh and ready. In this module, we will learn how to handle missing values, fix data types, remove duplicates, and standardise text. We will use tools from dplyr, tidyr, and base R.

By the end of this module, you will be able to take any messy dataset and make it clean and ready for analysis. You will be a data cleaner!


Learning Objectives

  • Understand why data cleaning is important.
  • Identify common data quality issues.
  • Handle missing values with na.omit(), drop_na(), and replace_na().
  • Fix data types using as.numeric(), as.character(), etc.
  • Remove duplicate rows with distinct().
  • Standardise text (lowercase, uppercase, trim spaces).
  • Use mutate() to create corrected columns.
  • Apply cleaning to real data, including Nigerian examples.
  • Create a reproducible cleaning pipeline.

Warm-up Story: The Dirty Data Dilemma

Chidi was excited to analyse data from his school's sports day. He imported the data, but it was a mess! Some scores were missing, some names had extra spaces, and one pupil's name was spelled as "Ade" in some rows and "AdE" in others. The ages were stored as text, not numbers. He couldn't make a graph because R didn't understand the data.

Chidi decided to clean the data. He fixed the spellings, converted ages to numbers, filled missing scores with the average, and removed duplicates. Now, the data was perfect. He could analyse it and create beautiful graphs. Chidi learned that cleaning data is the first and most important step in any analysis.


Main Lessons

Lesson 1: What is Data Cleaning?

Definition: Data cleaning is the process of fixing errors, handling missing values, and standardising data to make it ready for analysis.

Why it is important: Dirty data leads to wrong conclusions. Cleaning ensures accuracy.

Simple explanation: It's like tidying your room – you put everything in its place.

Real-life example: A shopkeeper corrects typos in product names.

School example: Your teacher checks that all names are spelled correctly.

Home example: You sort your clothes and remove those with holes.

Nigerian example: A researcher cleans census data to remove errors.

Illustration:

   Dirty Data --> Clean Data --> Ready for Analysis
   (errors,      (fixed,
    missing,      complete,
    duplicates)   consistent)

Mini summary: Data cleaning fixes problems so your analysis is correct.


Lesson 2: Common Data Problems

Here are some common problems:

  • Missing values (NA)
  • Wrong data types (numbers as text)
  • Duplicates (same row twice)
  • Inconsistent text (e.g., "lagos" vs "Lagos")
  • Extra spaces ("Ade " vs "Ade")
  • Outliers (extreme values)

Mini summary: We will learn how to fix each of these.


Lesson 3: Handling Missing Values – drop_na()

Definition: drop_na() removes rows that have any missing values.

Why it is important: It's a quick way to get complete cases.

Simple explanation: It's like throwing away rotten apples.

Real-life example: A survey where some questions were not answered – you remove those responses.

School example: A pupil was absent – you remove that row.

Home example: You forgot to record a chore – you remove that entry.

Nigerian example: A state's data is missing – you remove that state.

   library(tidyr)
   data %>% drop_na()  # removes rows with any NA
   data %>% drop_na(score)  # removes rows where score is NA

Mini summary: drop_na() removes rows with missing values.


Lesson 4: Handling Missing Values – replace_na()

Definition: replace_na() replaces missing values with a specified value.

Why it is important: Sometimes we want to fill in missing values instead of removing rows.

Simple explanation: It's like drawing a smile on a blank face.

Real-life example: You replace a missing score with the average.

School example: You fill absent marks with 0.

Home example: You replace a missing chore time with "0 minutes".

Nigerian example: You replace missing population with the national average.

   data %>% replace_na(list(score = 0, age = 10))
   # Replaces NA in score with 0, in age with 10.

Mini summary: replace_na() fills missing values.


Lesson 5: Handling Missing Values – na.omit()

Definition: na.omit() is a base R function that removes rows with NA.

Why it is important: It's a quick alternative to drop_na().

   clean_data <- na.omit(data)

Mini summary: na.omit() removes rows with missing values.


Lesson 6: Fixing Data Types – as.numeric()

Definition: as.numeric() converts text to numbers.

Why it is important: Numbers stored as text cannot be used in calculations.

Simple explanation: It's like changing a word "five" to the number 5.

Real-life example: A column "age" might be stored as "10" (text) – we need it as number.

School example: Scores stored as "85" (text) – we need numbers to calculate averages.

Home example: Time stored as "30" (text) – we need numbers to sum.

Nigerian example: Population stored as "20 million" – we need to extract numbers.

   data %>% mutate(score = as.numeric(score))
   # Converts score column to numeric.

Mini summary: as.numeric() turns text into numbers.


Lesson 7: Fixing Data Types – as.character() and as.factor()

Definition: as.character() converts to text; as.factor() converts to categorical data.

Why it is important: Sometimes we need text for labels, or factors for grouping.

Real-life example: "Male"/"Female" should be a factor.

School example: "Class" (5,6) should be a factor, not a number.

Home example: "Chore" names should be text.

Nigerian example: "State" should be a factor.

   data %>% mutate(class = as.factor(class))

Mini summary: Use as.character() and as.factor() to fix types.


Lesson 8: Removing Duplicates – distinct()

Definition: distinct() removes duplicate rows.

Why it is important: Duplicates can overcount and skew results.

Simple explanation: It's like deleting repeated entries in a list.

Real-life example: A customer appears twice in a mailing list – remove one.

School example: A pupil's name appears twice – remove duplicate.

Home example: You wrote the same chore twice – remove one.

Nigerian example: A state appears twice in a dataset – remove duplicate.

   data %>% distinct()  # removes all duplicate rows
   data %>% distinct(name, .keep_all = TRUE)  # keep first occurrence

Mini summary: distinct() removes duplicate rows.


Lesson 9: Standardising Text – tolower(), toupper(), trimws()

Definition: tolower() converts text to lowercase; toupper() to uppercase; trimws() removes extra spaces.

Why it is important: Inconsistent text prevents matching and grouping.

Simple explanation: It's like making sure everyone writes their name the same way.

Real-life example: "Lagos" and "lagos" should be the same – convert to "lagos" or "Lagos".

School example: "Ade" and "aDe" should be "Ade".

Home example: "Chore" and "chore " – remove spaces.

Nigerian example: "Abuja" and "abuja" – standardise.

   data %>% mutate(name = tolower(name))  # all lowercase
   data %>% mutate(name = trimws(name))   # remove leading/trailing spaces

Mini summary: Use tolower(), toupper(), and trimws() to clean text.


Lesson 10: Handling Outliers

Definition: Outliers are values that are very different from the rest.

Why it is important: They can skew averages and graphs.

Simple explanation: A very tall person in a class is an outlier.

How to handle: You can remove them or cap them at a threshold.

   # Remove values above 100 or below 0
   data %>% filter(score >= 0 & score <= 100)
   # Or use quantiles
   data %>% filter(score > quantile(score, 0.25) - 1.5*IQR(score) &
                   score < quantile(score, 0.75) + 1.5*IQR(score))

Mini summary: Outliers are extreme values – handle with care.


Lesson 11: Creating a Cleaning Pipeline

We can chain all cleaning steps together:

   clean_data <- raw_data %>%
       drop_na() %>%
       distinct() %>%
       mutate(name = tolower(trimws(name))) %>%
       mutate(age = as.numeric(age)) %>%
       filter(age > 0 & age < 120)  # reasonable age range

Mini summary: A pipeline makes cleaning reproducible and clear.


Lesson 12: Real Nigerian Example – Cleaning Population Data

Let's clean a messy dataset of Nigerian states.

   # Messy data
   messy <- data.frame(
       state = c("Lagos", "Lagos ", "lagos", "Kano", "Kano"),
       population = c("20", "21", NA, "15", "15"),
       region = c("SW", "SW", "SW", "NW", "NW")
   )

   clean <- messy %>%
       distinct() %>%
       mutate(state = tolower(trimws(state))) %>%
       mutate(population = as.numeric(population)) %>%
       drop_na()
   # Result: one row per state, with numeric population, no NAs.
   Before:
   state    population region
   Lagos    20         SW
   Lagos    21         SW
   lagos    NA         SW
   Kano     15         NW
   Kano     15         NW

   After cleaning:
   state    population region
   lagos    20         SW
   lagos    21         SW   (if distinct on all columns, duplicates removed)
   kano     15         NW

Mini summary: Cleaning makes Nigerian data usable.


Lesson 13: Using janitor for Quick Cleaning

Definition: The janitor package has helpers for cleaning column names and more.

   install.packages("janitor")
   library(janitor)
   data %>% clean_names()  # makes column names consistent

Mini summary: janitor is a helpful addition.


Lesson 14: Validating Data After Cleaning

Always check:

  • No more NA values (if you intended to remove them).
  • Data types are correct.
  • No duplicates.
  • Text is consistent.

Use glimpse(), summary(), and head().

Mini summary: Validate to ensure cleaning worked.


Lesson 15: Summary of Cleaning

We learned to handle missing values, fix types, remove duplicates, standardise text, and manage outliers. Cleaning is the foundation of good analysis.


Key Vocabulary

Data Cleaning
Fixing errors and inconsistencies in data.
Missing Values
Empty cells represented as NA.
Outlier
A value that is very different from others.
Duplicate
An identical row that appears more than once.
Data Type
Whether data is numeric, text, factor, etc.
Standardise
Making data consistent (e.g., all lowercase).
drop_na()
Removes rows with missing values.
replace_na()
Fills missing values.
distinct()
Removes duplicate rows.
trimws()
Removes extra spaces.

Important Concepts

  • Garbage in, garbage out: If your data is dirty, your analysis will be wrong.
  • Reproducibility: Cleaning steps should be recorded in code.
  • Validation: Always check your cleaned data.
  • Consistency: Use the same cleaning rules for all data.

Step-by-Step Explanations

How to clean a dataset:

  1. Import data.
  2. Check for missing values (is.na()).
  3. Remove or fill missing values.
  4. Check data types (str()).
  5. Fix data types.
  6. Remove duplicates (distinct()).
  7. Standardise text.
  8. Handle outliers.
  9. Validate with glimpse().

Real-life Examples

  • A hospital cleans patient records to have consistent names and dates.
  • A bank cleans transaction data to remove duplicates.
  • A school cleans attendance data to fill missing marks.

Nigerian Examples

  • Cleaning population data from different sources to match formats.
  • Standardising state names (e.g., "Lagos" vs "LAGOS").
  • Handling missing values in agricultural surveys.

Fun Examples Children Relate To

  • Cleaning a list of your friends' names (fixing spelling).
  • Removing duplicate entries from your toy list.
  • Filling missing scores in your games.

Everyday Examples

  • You correct spelling errors in your homework.
  • You remove duplicate items from your shopping list.
  • You convert "10" (text) to 10 (number) for calculations.

Teacher Notes

  • Emphasise that cleaning is often the most time-consuming part of data analysis.
  • Provide dirty datasets for practice.
  • Encourage students to write cleaning pipelines.
  • Discuss the consequences of not cleaning data.

Parent Tips

  • Help your child collect data and point out errors.
  • Discuss how cleaning is like organising a room.
  • Encourage them to check their work after cleaning.

Interesting Facts

  • Data scientists spend 60-80% of their time cleaning data.
  • Poor data quality costs the US economy trillions of dollars each year.
  • R has many packages dedicated to data cleaning.

Did You Know?

  • The janitor package has a function get_dupes() to find duplicates.
  • You can use skimr to get a quick summary of data quality.
  • Some packages automatically clean data during import.

Remember This

  • Always clean your data before analysis.
  • Handle missing values carefully.
  • Check data types and fix them.
  • Remove duplicates.
  • Standardise text.
  • Validate your cleaned data.

Common Mistakes

  • Forgetting to check for NA values.
  • Removing too many rows with drop_na().
  • Converting factors to numbers incorrectly.
  • Not removing duplicates before analysis.
  • Ignoring text inconsistencies.

Best Practices

  • Write a cleaning script that you can reuse.
  • Use pipelines to make cleaning steps clear.
  • Always keep a copy of the raw data.
  • Document your cleaning decisions.
  • Test your cleaning with small data first.

ASCII Illustrations

Cleaning pipeline

   Raw Data
       |
       V
   drop_na()  ---->  remove missing
       |
       V
   distinct() ---->  remove duplicates
       |
       V
   mutate()   ---->  fix types & text
       |
       V
   Clean Data

Comparison Tables

Handling Missing Values
FunctionActionWhen to Use
drop_na()Removes rowsIf few missing values
replace_na()Fills with valueTo keep all rows
na.omit()Removes rowsBase R alternative

End-of-Module Summary

In this module, we learned how to clean dirty data. We handled missing values with drop_na() and replace_na(), fixed data types with as.numeric(), removed duplicates with distinct(), and standardised text with tolower() and trimws(). We also learned about outliers and validation. Clean data is essential for accurate analysis and visualisation.


Frequently Asked Questions

1. What is data cleaning?
Fixing errors and inconsistencies in data.
2. Why is data cleaning important?
It ensures accurate analysis.
3. What is a missing value?
An empty cell represented as NA.
4. How do you remove missing values?
Use drop_na().
5. How do you fill missing values?
Use replace_na().
6. What is a duplicate?
An identical row appearing more than once.
7. How do you remove duplicates?
Use distinct().
8. What is an outlier?
A value very different from others.
9. How do you fix data types?
Use as.numeric(), as.character(), etc.
10. What is validation?
Checking that cleaning worked correctly.

Review Questions

  1. What is data cleaning?
  2. Why is cleaning important?
  3. What is a missing value?
  4. How do you remove rows with missing values?
  5. How do you fill missing values?
  6. What is a duplicate row?
  7. How do you remove duplicates?
  8. What is an outlier?
  9. How do you fix data types?
  10. What does trimws() do?
  11. What does tolower() do?
  12. What is a cleaning pipeline?
  13. Why should you validate cleaned data?
  14. Give an example of dirty text.
  15. What is the janitor package?

Fill-in-the-Blank Exercises

  1. ______ means fixing errors in data. (Cleaning)
  2. ______ removes rows with missing values. (drop_na)
  3. ______ fills missing values. (replace_na)
  4. ______ removes duplicate rows. (distinct)
  5. ______ converts text to numbers. (as.numeric)
  6. ______ removes extra spaces. (trimws)
  7. ______ converts text to lowercase. (tolower)
  8. An ______ is a very different value. (outlier)
  9. ______ is checking if cleaning worked. (Validation)
  10. A ______ is an identical row that appears more than once. (duplicate)

True or False Exercises

  1. Data cleaning is not necessary. (False)
  2. drop_na() removes rows with missing values. (True)
  3. replace_na() replaces missing values. (True)
  4. distinct() removes duplicate rows. (True)
  5. as.numeric() converts numbers to text. (False)
  6. trimws() removes spaces. (True)
  7. Outliers are always bad. (False)
  8. You should always validate cleaned data. (True)
  9. Cleaning pipelines are not useful. (False)
  10. janitor is a cleaning package. (True)

Multiple Choice Questions

  1. Which function removes rows with missing values?
    A. replace_na B. drop_na C. na.fill D. na.omit
    Answer: B
  2. Which function fills missing values?
    A. drop_na B. replace_na C. na.omit D. distinct
    Answer: B
  3. Which function removes duplicate rows?
    A. distinct B. unique C. duplicate D. Both A and B
    Answer: D
  4. Which function converts text to numbers?
    A. as.character B. as.numeric C. as.factor D. as.text
    Answer: B
  5. What does trimws() do?
    A. Removes spaces B. Adds spaces C. Converts to lowercase D. Converts to uppercase
    Answer: A
  6. What is an outlier?
    A. A missing value B. A duplicate C. An extreme value D. A text value
    Answer: C
  7. What is validation?
    A. Removing outliers B. Checking cleaning worked C. Importing data D. Exporting data
    Answer: B
  8. Which package helps with cleaning?
    A. janitor B. readr C. dplyr D. All of the above
    Answer: D
  9. What does tolower() do?
    A. Converts to uppercase B. Converts to lowercase C. Removes spaces D. Converts to numbers
    Answer: B
  10. Why is cleaning important?
    A. It saves time B. It ensures accuracy C. It is fun D. It is required
    Answer: B
  11. What is a duplicate?
    A. A missing value B. An extreme value C. An identical row D. A different row
    Answer: C
  12. How do you fix data types?
    A. as.numeric B. as.character C. as.factor D. All of the above
    Answer: D
  13. What is a cleaning pipeline?
    A. A series of steps B. A single function C. A package D. A file
    Answer: A
  14. What does na.omit() do?
    A. Replaces NA B. Removes NA C. Counts NA D. Ignores NA
    Answer: B
  15. Which function removes duplicates in dplyr?
    A. unique B. distinct C. remove_dups D. dedupe
    Answer: B

Matching Exercises

ProblemSolution
1. Missing valuesA. distinct()
2. DuplicatesB. as.numeric()
3. Text as numbersC. drop_na()
4. Extra spacesD. trimws()
5. Inconsistent caseE. tolower()

Answers: 1-C, 2-A, 3-B, 4-D, 5-E


Short Answer Questions

  1. What is data cleaning and why is it important?
  2. Explain the difference between drop_na() and replace_na().
  3. How do you remove duplicates from a dataset?
  4. What is an outlier and how might you handle it?
  5. Why is validation important after cleaning?

Scenario-based Exercises

  1. Scenario: You have a dataset with pupil scores. Some scores are missing, and some are stored as text. How would you clean it?
  2. Scenario: You have a list of Nigerian states with inconsistent spellings and extra spaces. How would you standardise them?
  3. Scenario: Your data has duplicate rows due to data entry errors. How would you fix it?

Group Activity

In groups, create a dirty dataset (with missing values, duplicates, inconsistent text). Exchange datasets with another group and clean them using R. Compare your cleaning approaches.


Individual Activity

Use the airquality dataset (built-in). It has missing values. Clean it by removing rows with NA and converting month to a factor. Save the cleaned dataset.


Classroom Discussion Questions

  1. Why is data cleaning often the most time-consuming part of data science?
  2. What are the risks of not cleaning data?
  3. How can cleaning pipelines help?
  4. What would you do if you found an outlier that seems real?

Mini Project

Title: "Clean and Analyse Nigerian School Data"
You are given a messy dataset of Nigerian school attendance with missing values, duplicates, and inconsistent state names. Clean the data and then summarise attendance by state. Write a short report.


Practical Assignment

Using the diamonds dataset (built-in), perform cleaning: check for missing values, duplicates, and outliers (price, carat). Then, create a summary table of cleaned data.


Challenge Exercise

Find a real, messy dataset online (e.g., from Kaggle). Apply a complete cleaning pipeline using dplyr and tidyr. Document each step and explain why you made each choice.


Quiz Answers

Fill-in-the-Blank: 1. Cleaning, 2. drop_na, 3. replace_na, 4. distinct, 5. as.numeric, 6. trimws, 7. tolower, 8. outlier, 9. Validation, 10. duplicate.

True/False: 1F, 2T, 3T, 4T, 5F, 6T, 7F, 8T, 9F, 10T.

Multiple Choice: 1B, 2B, 3D, 4B, 5A, 6C, 7B, 8D, 9B, 10B, 11C, 12D, 13A, 14B, 15B.


Key Takeaways

  • Data cleaning is essential for accurate analysis.
  • Handle missing values with drop_na() or replace_na().
  • Fix data types with as.numeric() etc.
  • Remove duplicates with distinct().
  • Standardise text with tolower() and trimws().
  • Use pipelines for reproducible cleaning.
  • Always validate your cleaned data.

Preparation for the Next Module

In Module 13, we will learn about exploratory data analysis (EDA) – using summary statistics and visualisations to understand your data. We will combine all our skills: importing, cleaning, wrangling, and visualising. Practise cleaning different datasets to prepare.


14

Module Thirteen

Module 13 Β· R for Data Analysis Β· Exploratory Data Analysis

Module 13 Β· Exploratory Data Analysis – Getting to Know Your Data

Hello, young data explorer! In Module 12, we learned how to clean our data. Now that our data is clean, it's time to explore it. Exploratory Data Analysis (EDA) means looking at the data from many angles to understand its patterns, relationships, and surprises.

EDA is like being a detective. You ask questions, make plots, calculate summaries, and look for clues. The goal is to understand the data before you do any formal analysis. In this module, we will use summary statistics, visualisations, and grouping to explore data.

By the end of this module, you will be able to explore any dataset and tell its story. You will be a data detective!


Learning Objectives

  • Understand what Exploratory Data Analysis (EDA) is.
  • Use summary statistics (mean, median, mode, range, etc.).
  • Visualise distributions with histograms and boxplots.
  • Explore relationships with scatter plots.
  • Use group_by() and summarise() to compare groups.
  • Use ggplot2 to create exploratory plots.
  • Ask questions and answer them with data.
  • Identify patterns, trends, and outliers.
  • Apply EDA to real data, including Nigerian examples.
  • Document your findings in a report.

Warm-up Story: The Mystery of the Missing Snacks

Chidi noticed that snacks in his school canteen were disappearing quickly. He wanted to find out which snacks were most popular and when they were bought. He collected data on snack sales: snack name, price, quantity sold, and time of day.

He imported the data, cleaned it, and then started to explore. He calculated the average quantity sold per snack, made a bar chart to compare popularity, and used a line chart to see sales over time. He discovered that meat pies were the most popular, and sales peaked during lunchtime. He presented his findings to the canteen manager, who used the information to stock more meat pies.

Chidi had performed Exploratory Data Analysis!


Main Lessons

Lesson 1: What is Exploratory Data Analysis (EDA)?

Definition: EDA is the process of exploring data to understand its main characteristics, patterns, and relationships.

Why it is important: EDA helps you know your data before you do any formal analysis. It guides your next steps.

Simple explanation: It's like exploring a new playground – you check out all the equipment, find the best spots, and see what's fun.

Real-life example: A detective gathers clues at a crime scene.

School example: A teacher looks at test scores to see which topics need more review.

Home example: You look at your piggy bank to see how much money you have and where it came from.

Nigerian example: A researcher explores census data to understand population patterns.

Illustration:

   Data --> Ask Questions --> Plot/Summarise --> Find Patterns --> Tell Story

Mini summary: EDA helps you understand your data and find interesting insights.


Lesson 2: First Look – glimpse() and head()

Definition: glimpse() shows a summary of the data structure; head() shows the first few rows.

Why it is important: They give you a quick overview of the data.

Simple explanation: It's like peeking into a box to see what's inside.

Real-life example: You open a book and read the first page.

School example: You look at the class register to see who is in your class.

Home example: You check the pantry to see what food you have.

Nigerian example: You view the first rows of a dataset of Nigerian states.

   library(dplyr)
   glimpse(data)
   head(data, 10)  # first 10 rows

Mini summary: Always start with glimpse() and head().


Lesson 3: Summary Statistics – summary()

Definition: summary() gives basic statistics for each column: min, max, mean, median, quartiles.

Why it is important: It tells you about the center and spread of your data.

Simple explanation: It's like a report card for each column.

Real-life example: A teacher summarises test scores: average, highest, lowest.

School example: You summarise your weekly allowance.

Home example: You summarise the time you spend on homework each day.

Nigerian example: You summarise population data to see average state population.

   summary(data)
   # For a specific column:
   summary(data$score)
   Example output for score:
      Min. 1st Qu.  Median    Mean 3rd Qu.    Max.
     65.00   70.00   85.00   82.75   90.00   95.00

Mini summary: summary() gives a numeric snapshot of your data.


Lesson 4: Mean, Median, and Mode

Definition:

  • Mean: The average (sum divided by count).
  • Median: The middle value when sorted.
  • Mode: The most frequent value.

Why they are important: They describe the center of the data.

Simple explanation: Mean is like sharing candies equally; median is the middle candy; mode is the most common candy.

Real-life example: A shop uses mean to find average sales.

School example: Teacher calculates mean test score.

Home example: You find the median age in your family.

Nigerian example: You find the mean population of states.

   mean(data$score, na.rm = TRUE)
   median(data$score, na.rm = TRUE)
   # For mode, we can use table():
   table(data$favourite_food)  # shows frequency

Mini summary: Mean, median, and mode tell you about the typical value.


Lesson 5: Range, Variance, and Standard Deviation

Definition:

  • Range: Max - Min.
  • Variance: Average of squared differences from the mean.
  • Standard Deviation: Square root of variance.

Why they are important: They measure spread – how much data varies.

Simple explanation: Range is the distance between the smallest and largest; standard deviation is how much scores typically differ from the average.

Real-life example: A teacher sees if scores are close together or spread out.

School example: You compare the spread of test scores in two classes.

Home example: You check how much your daily reading time varies.

Nigerian example: You check the variability of rainfall across states.

   range(data$score)
   var(data$score, na.rm = TRUE)
   sd(data$score, na.rm = TRUE)

Mini summary: Range, variance, and SD measure how spread out data is.


Lesson 6: Histograms – Seeing Distribution

Definition: A histogram is a bar chart that shows the frequency of values in bins (intervals).

Why it is important: It shows the shape of the distribution (e.g., bell-shaped, skewed).

Simple explanation: It's like sorting your toys into boxes by size.

Real-life example: A teacher looks at a histogram of test scores to see how many students got each grade range.

School example: You see how many friends live in each neighbourhood.

Home example: You make a histogram of the time you spend on different activities.

Nigerian example: A histogram of ages in a community.

   library(ggplot2)
   ggplot(data, aes(x = score)) +
       geom_histogram(binwidth = 5, fill = "blue", color = "black") +
       labs(title = "Histogram of Scores")
   Histogram concept:
   Frequency
   8 |   ###
   6 |   ###   ###
   4 |   ###   ###   ###
   2 |   ###   ###   ###   ###
   0 |___###___###___###___###____
       60-64 65-69 70-74 75-79

Mini summary: Histograms show the distribution of a single variable.


Lesson 7: Boxplots – Seeing Outliers and Spread

Definition: A boxplot shows the median, quartiles, and outliers of a variable.

Why it is important: It quickly shows the spread and identifies outliers.

Simple explanation: It's like a picture of a box with whiskers.

Real-life example: A scientist uses a boxplot to show temperature variation.

School example: You compare test scores across classes with boxplots.

Home example: You compare the amount of pocket money you and your friend get.

Nigerian example: A boxplot of salaries in different states.

   ggplot(data, aes(x = "", y = score)) +
       geom_boxplot(fill = "orange") +
       labs(title = "Boxplot of Scores", y = "Score")
   Boxplot concept:
   +-----+   outlier (o)
   |     |
   |  +--+--+  (box = IQR)
   |  |  |  |
   |  +--+--+
   |     |
   +-----+
   (whiskers extend to min/max within 1.5*IQR)

Mini summary: Boxplots show spread and outliers.


Lesson 8: Scatter Plots – Relationships

Definition: A scatter plot shows the relationship between two numeric variables.

Why it is important: It reveals correlations (positive, negative, or none).

Simple explanation: It's like plotting points on a map to see if they form a pattern.

Real-life example: A doctor plots height vs weight to see if they are related.

School example: You plot study time vs test score to see if more study leads to higher scores.

Home example: You plot age vs screen time.

Nigerian example: A scatter plot of education level vs income.

   ggplot(data, aes(x = study_time, y = score)) +
       geom_point() +
       labs(title = "Study Time vs Score")
   Scatter plot concept:
   Score
   100 |        .
    80 |      .   .
    60 |    .       .
    40 |  .
    20 |.
     0 |___._.___.___.___.___.___
        0   20  40  60  80 100
        Study Time

Mini summary: Scatter plots show relationships between variables.


Lesson 9: Grouped Summaries – compare groups

Definition: Use group_by() and summarise() to get statistics for each group.

Why it is important: It lets you compare groups (e.g., boys vs girls).

Simple explanation: It's like sorting your toys and then counting how many in each group.

Real-life example: A store compares sales by product category.

School example: A teacher compares average scores by class.

Home example: You compare the time you spend on school vs play.

Nigerian example: You compare average income by region.

   data %>%
       group_by(class) %>%
       summarise(avg_score = mean(score),
                 median_score = median(score),
                 n = n())

Mini summary: Grouped summaries reveal differences between groups.


Lesson 10: Faceting – Multiple Plots

Definition: Faceting creates separate plots for each group in a single figure.

Why it is important: It makes comparisons easy.

Simple explanation: It's like having a separate page for each group.

   ggplot(data, aes(x = score)) +
       geom_histogram() +
       facet_wrap(~ class)

Mini summary: Faceting helps compare groups visually.


Lesson 11: Correlation – How Variables Move Together

Definition: Correlation measures the strength and direction of a relationship between two variables.

Why it is important: It tells you if variables are related.

Simple explanation: It's like seeing if two friends always walk together.

Real-life example: Height and weight have a positive correlation.

School example: Study time and test scores often have a positive correlation.

Home example: The amount of exercise and health might have a positive correlation.

Nigerian example: Education and income have a positive correlation.

   cor(data$study_time, data$score, use = "complete.obs")

Mini summary: Correlation shows if variables change together.


Lesson 12: EDA with Nigerian Data

Let's explore a dataset of Nigerian states.

   # Example: population and region
   data %>%
       group_by(region) %>%
       summarise(avg_pop = mean(population),
                 sd_pop = sd(population)) %>%
       ggplot(aes(x = region, y = avg_pop)) +
           geom_bar(stat = "identity")

Mini summary: Apply EDA to any Nigerian data.


Lesson 13: Asking Questions – The Key to EDA

Always ask questions like:

  • What is the average?
  • What is the range?
  • Are there outliers?
  • Are groups different?
  • Are variables related?

Mini summary: Good questions lead to good explorations.


Lesson 14: Documenting Your EDA

Keep a record of your findings. Use an R Markdown or a script with comments.

Mini summary: Document your EDA so others can follow.


Lesson 15: Summary of EDA

EDA is a journey of discovery. You use summary stats, plots, and grouping to understand your data. It's the foundation of all analysis.


Key Vocabulary

Exploratory Data Analysis (EDA)
Exploring data to understand it.
Mean
Average.
Median
Middle value.
Mode
Most frequent value.
Range
Max - Min.
Standard Deviation
How spread out data is.
Histogram
Chart showing distribution.
Boxplot
Chart showing spread and outliers.
Scatter Plot
Chart showing relationship.
Correlation
How variables move together.

Important Concepts

  • Summary statistics describe data.
  • Visualisation reveals patterns.
  • Grouping compares categories.
  • Correlation shows relationships.
  • Ask questions to guide exploration.

Step-by-Step Explanations

How to perform EDA:

  1. Import and clean data.
  2. Use glimpse() and head().
  3. Get summary statistics with summary().
  4. Create histograms for numeric variables.
  5. Create boxplots to check for outliers.
  6. Create scatter plots for pairs of variables.
  7. Group data and compare groups.
  8. Document your findings.

Real-life Examples

  • A business explores sales data to find best-selling products.
  • A doctor explores patient data to understand disease patterns.
  • A teacher explores test scores to identify weak areas.

Nigerian Examples

  • Exploring population data to find most populous states.
  • Exploring agricultural data to find crop yields by region.
  • Exploring education data to find literacy rates.

Fun Examples Children Relate To

  • Exploring your toy collection: which type is most common?
  • Exploring your friends' ages: what is the average?
  • Exploring your screen time: how does it vary by day?

Everyday Examples

  • You explore your pocket money to see how much you spend on snacks.
  • You explore the weather to see the hottest month.
  • You explore your homework time to see which day you work most.

Teacher Notes

  • Encourage students to ask questions.
  • Use real datasets that students find interesting.
  • Show how EDA guides further analysis.
  • Emphasise that EDA is creative and exploratory.

Parent Tips

  • Help your child collect data at home and explore it.
  • Discuss what the data tells you.
  • Ask open-ended questions to guide exploration.

Interesting Facts

  • EDA was popularised by John Tukey in the 1970s.
  • EDA is often the most creative part of data analysis.
  • Many discoveries have come from EDA.

Did You Know?

  • You can use GGally::ggpairs() to create a matrix of plots.
  • EDA is often used in business to find opportunities.
  • EDA can be done with just a few lines of code.

Remember This

  • Always start with EDA.
  • Use summary stats and plots.
  • Look for patterns, outliers, and relationships.
  • Ask questions and answer them.
  • Document your findings.

Common Mistakes

  • Skipping EDA and going straight to modelling.
  • Not looking at the data distribution.
  • Ignoring outliers.
  • Assuming correlation means causation.
  • Not asking enough questions.

Best Practices

  • Start with a clean dataset.
  • Use multiple visualisations.
  • Summarise groups separately.
  • Write notes as you explore.
  • Use EDA to refine your questions.

ASCII Illustrations

EDA process

   Data --> Clean --> Explore --> Summarise --> Visualise --> Insights

Histogram concept

   Frequency
   8 |   ###
   6 |   ###   ###
   4 |   ###   ###   ###
   2 |   ###   ###   ###   ###
   0 |___###___###___###___###____
       60-64 65-69 70-74 75-79

Comparison Tables

Summary Statistics
StatisticWhat it showsFunction
MeanAveragemean()
MedianMiddlemedian()
RangeMin to Maxrange()
SDSpreadsd()

End-of-Module Summary

In this module, we learned how to explore data using Exploratory Data Analysis (EDA). We used summary statistics (mean, median, range, SD), histograms, boxplots, scatter plots, and grouped summaries. We asked questions and looked for patterns, outliers, and relationships. EDA is the first and most important step in any data analysis.


Frequently Asked Questions

1. What is EDA?
Exploratory Data Analysis – exploring data to understand it.
2. What is the mean?
The average.
3. What is the median?
The middle value.
4. What is a histogram?
A chart showing distribution.
5. What is a boxplot?
A chart showing spread and outliers.
6. What is a scatter plot?
A chart showing relationships.
7. What is correlation?
How variables move together.
8. Why is EDA important?
It helps you understand your data.
9. What is a grouped summary?
Statistics for each group.
10. What should I do after EDA?
Use the insights for further analysis.

Review Questions

  1. What is EDA?
  2. What is the mean?
  3. What is the median?
  4. What is a histogram used for?
  5. What is a boxplot used for?
  6. What is a scatter plot used for?
  7. What is correlation?
  8. How do you group data in R?
  9. Why is EDA important?
  10. What is the range?
  11. What is standard deviation?
  12. What does glimpse() do?
  13. What does summary() do?
  14. What is a grouped summary?
  15. Give an example of an EDA question.

Fill-in-the-Blank Exercises

  1. ______ is exploring data to understand it. (EDA)
  2. The ______ is the average. (mean)
  3. The ______ is the middle value. (median)
  4. A ______ shows the distribution of a variable. (histogram)
  5. A ______ shows spread and outliers. (boxplot)
  6. A ______ shows relationships between variables. (scatter plot)
  7. ______ measures how variables move together. (Correlation)
  8. Use ______ to group data. (group_by)
  9. The ______ is the difference between max and min. (range)
  10. ______ shows the first few rows. (head)

True or False Exercises

  1. EDA is not important. (False)
  2. The mean is the same as the median. (False)
  3. A histogram shows frequency. (True)
  4. A boxplot shows outliers. (True)
  5. A scatter plot shows correlation. (True)
  6. Correlation means causation. (False)
  7. glimpse() shows data structure. (True)
  8. summary() gives statistics. (True)
  9. Grouped summaries compare groups. (True)
  10. EDA is the last step. (False)

Multiple Choice Questions

  1. What is EDA?
    A. Exploratory Data Analysis B. Extreme Data Analysis C. Easy Data Analysis D. Extra Data Analysis
    Answer: A
  2. What is the mean?
    A. The middle B. The average C. The most frequent D. The max
    Answer: B
  3. What is the median?
    A. The average B. The middle C. The most frequent D. The min
    Answer: B
  4. Which chart shows distribution?
    A. Scatter plot B. Histogram C. Boxplot D. Bar chart
    Answer: B
  5. Which chart shows outliers?
    A. Scatter plot B. Histogram C. Boxplot D. Pie chart
    Answer: C
  6. Which chart shows relationships?
    A. Scatter plot B. Histogram C. Boxplot D. Bar chart
    Answer: A
  7. What does summary() do?
    A. Shows first rows B. Shows statistics C. Shows structure D. Shows plot
    Answer: B
  8. What does glimpse() do?
    A. Shows statistics B. Shows structure C. Shows plot D. Shows first rows
    Answer: B
  9. What is correlation?
    A. How variables move together B. The average C. The middle D. The spread
    Answer: A
  10. How do you group data in R?
    A. group_by() B. group() C. by_group() D. separate()
    Answer: A
  11. What is the range?
    A. Max - Min B. Mean - Median C. Variance D. SD
    Answer: A
  12. What is standard deviation?
    A. Average B. Spread C. Middle D. Max
    Answer: B
  13. Which function shows first rows?
    A. head() B. tail() C. glimpse() D. summary()
    Answer: A
  14. What is a grouped summary?
    A. Statistics for each group B. Overall statistics C. Plot for each group D. No statistics
    Answer: A
  15. Why is EDA important?
    A. It is required B. It helps understand data C. It is easy D. It is fast
    Answer: B

Matching Exercises

ConceptDescription
1. MeanA. Middle value
2. MedianB. Average
3. HistogramC. Shows distribution
4. BoxplotD. Shows outliers
5. Scatter plotE. Shows relationship

Answers: 1-B, 2-A, 3-C, 4-D, 5-E


Short Answer Questions

  1. What is Exploratory Data Analysis?
  2. Explain the difference between mean and median.
  3. What is a histogram used for?
  4. How does a boxplot help you find outliers?
  5. Why is it important to explore data before analysis?

Scenario-based Exercises

  1. Scenario: You have data on pupils' test scores. You want to know if there is a difference between boys and girls. What would you do?
  2. Scenario: You have data on daily sales. You notice some days have very high sales. How would you explore this?
  3. Scenario: You have data on height and weight. You want to see if they are related. What plot would you make?

Group Activity

In groups, choose a dataset (e.g., iris). Perform EDA: summary stats, histograms, boxplots, scatter plots, and grouped summaries. Present your findings to the class.


Individual Activity

Use the mtcars dataset. Perform EDA to understand the relationship between horsepower (hp) and fuel efficiency (mpg). Write a short paragraph on your findings.


Classroom Discussion Questions

  1. What are the most interesting patterns you found?
  2. What questions did your EDA raise?
  3. How can EDA help in Nigerian business?
  4. What would you do if you found outliers?

Mini Project

Title: "Explore Nigerian States"
Find a dataset of Nigerian states (population, region, etc.). Perform a complete EDA: summary stats, histograms, boxplots, and grouped summaries. Write a report with your findings and include visualisations.


Practical Assignment

Using the diamonds dataset, perform EDA: explore price, carat, cut, and their relationships. Create at least 5 visualisations and summarise your findings.


Challenge Exercise

Find a real dataset online (e.g., Kaggle). Perform a complete EDA and write a report. Include at least 3 different types of plots and grouped summaries. Present your findings in a clear and engaging way.


Quiz Answers

Fill-in-the-Blank: 1. EDA, 2. mean, 3. median, 4. histogram, 5. boxplot, 6. scatter plot, 7. Correlation, 8. group_by, 9. range, 10. head.

True/False: 1F, 2F, 3T, 4T, 5T, 6F, 7T, 8T, 9T, 10F.

Multiple Choice: 1A, 2B, 3B, 4B, 5C, 6A, 7B, 8B, 9A, 10A, 11A, 12B, 13A, 14A, 15B.


Key Takeaways

  • EDA is the first step in data analysis.
  • Use summary statistics to describe data.
  • Use histograms, boxplots, and scatter plots to visualise.
  • Grouped summaries compare categories.
  • Correlation shows relationships.
  • Always ask questions and document your findings.

Preparation for the Next Module

In Module 14, we will learn about statistical testing – how to make decisions based on data. We will use tests like t-test, ANOVA, and chi-square to see if differences are real or due to chance. Practise EDA on different datasets to get comfortable with exploring.


15

Module Fourteen

Module 14 Β· R for Data Analysis Β· Statistical Testing

Module 14 Β· Statistical Testing – Making Decisions with Data

Hello, young data explorer! In Module 13, we learned how to explore data and find patterns. But sometimes we need to know if a pattern is real or just due to chance. For example, if boys score higher than girls on a test, is that a real difference or just a coincidence?

This is where statistical testing comes in. Statistical tests help us make decisions based on data. They tell us if the differences we see are significant (real) or not. We will learn about the t-test (comparing two groups), ANOVA (comparing more than two groups), and chi-square test (for categorical data).

By the end of this module, you will be able to test your data and make confident decisions. You will be a data decision-maker!


Learning Objectives

  • Understand what statistical testing is.
  • Learn the concept of p-value and significance.
  • Use the t-test to compare two groups.
  • Use ANOVA to compare more than two groups.
  • Use the chi-square test for categorical data.
  • Interpret test results and p-values.
  • Apply tests to real data, including Nigerian examples.
  • Avoid common pitfalls in testing.

Warm-up Story: The Basketball Game

Chidi and his friend Bola argued about who was the better basketball player. They decided to record the number of points they scored in 10 games. Chidi's average was 12 points, and Bola's was 10 points. But was that difference real or just luck? They used a t-test to find out!

The t-test gave a p-value of 0.20. Since this was greater than 0.05, they concluded that the difference was not significant – it could have happened by chance. They decided they were equally good and stopped arguing. Statistical testing helped them settle the debate.


Main Lessons

Lesson 1: What is Statistical Testing?

Definition: Statistical testing is a way to decide if a pattern in data is real or just due to chance.

Why it is important: It helps us make objective decisions, not just guesses.

Simple explanation: It's like a referee who decides if a goal is valid or not.

Real-life example: A doctor tests if a new medicine works better than an old one.

School example: A teacher tests if teaching method A gives better scores than method B.

Home example: You test if you sleep better with or without a nightlight.

Nigerian example: A researcher tests if a new farming method increases crop yield.

Illustration:

   Data --> Hypothesis --> Test --> p-value --> Decision

Mini summary: Statistical testing helps us decide if patterns are real.


Lesson 2: The p-value – The Probability of Chance

Definition: The p-value is the probability of getting the observed result (or more extreme) if there is no real difference (i.e., if the null hypothesis is true).

Why it is important: A small p-value (usually < 0.05) means the result is unlikely to be due to chance, so we call it statistically significant.

Simple explanation: It's like a magic number that tells you how surprising your result is. If it's very small (less than 5%), it's a real effect.

Real-life example: A p-value of 0.03 means there's only a 3% chance the difference is due to luck.

School example: If p < 0.05, the new teaching method is likely better.

Home example: If p < 0.05, the nightlight might really affect your sleep.

Nigerian example: If p < 0.05, the new fertiliser really increases crop yield.

Key rule: p < 0.05 = significant (real effect). p > 0.05 = not significant (could be chance).

Mini summary: p-value tells us if a result is real or due to chance.


Lesson 3: The Null and Alternative Hypotheses

Definition:

  • Null hypothesis (H0): There is no real difference or effect.
  • Alternative hypothesis (H1): There is a real difference or effect.

Why it is important: The test decides whether to reject the null hypothesis.

Simple explanation: Null is like saying "nothing is going on"; alternative is "something is happening".

Real-life example: H0: the medicine doesn't work; H1: the medicine works.

School example: H0: boys and girls have same average score; H1: they differ.

Home example: H0: the nightlight has no effect; H1: it affects sleep.

Nigerian example: H0: fertiliser has no effect; H1: it increases yield.

Mini summary: We test if we can reject H0 in favour of H1.


Lesson 4: The t-test – Comparing Two Groups

Definition: The t-test compares the means of two groups to see if they are significantly different.

Why it is important: It's the most common test for comparing two groups.

Simple explanation: It's like weighing two boxes to see if one is heavier.

Real-life example: Compare the height of boys and girls.

School example: Compare test scores of class A and class B.

Home example: Compare your screen time with your friend's.

Nigerian example: Compare average income of urban vs rural areas.

Code:

   t.test(data$score ~ data$group)  # where group has two categories
   # Example: t.test(scores ~ gender)
   t-test output:
   t = 2.45, df = 18, p-value = 0.025
   alternative hypothesis: true difference in means is not equal to 0
   95% confidence interval: [2.1, 8.9]
   sample estimates:
   mean in group A  mean in group B
   75.0             69.0

Mini summary: t-test compares means of two groups.


Lesson 5: Interpreting t-test Results

Steps:

  1. Look at the p-value.
  2. If p < 0.05, the groups are significantly different.
  3. If p > 0.05, there is no significant difference.
  4. Check the confidence interval – it tells the range of the difference.

Mini summary: p < 0.05 = significant difference.


Lesson 6: Assumptions of t-test

Definition: t-test assumes:

  • Data is numeric.
  • Data is approximately normally distributed (bell-shaped).
  • Variances (spread) are equal (if using standard t-test).

Why it is important: If assumptions are violated, the test may not be valid.

Simple explanation: It's like rules of a game – you need to follow them.

Mini summary: Check assumptions before using t-test.


Lesson 7: ANOVA – Comparing More Than Two Groups

Definition: Analysis of Variance (ANOVA) compares means of three or more groups.

Why it is important: When you have more than two groups, you use ANOVA.

Simple explanation: It's like a t-test for many groups.

Real-life example: Compare test scores of three different schools.

School example: Compare scores of students from different teachers.

Home example: Compare time spent on different activities.

Nigerian example: Compare crop yield across four regions.

Code:

   # One-way ANOVA
   aov_model <- aov(score ~ group, data = data)
   summary(aov_model)
   ANOVA output:
               Df Sum Sq Mean Sq F value Pr(>F)
   group        2  123.4   61.7   4.23 0.023 *
   Residuals   27  393.5   14.6
   Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Mini summary: ANOVA compares means of three or more groups.


Lesson 8: Post-hoc Tests – Finding Which Groups Differ

Definition: After ANOVA, if the result is significant, post-hoc tests tell us which groups are different.

Why it is important: ANOVA tells us there is a difference, but not where it is.

Simple explanation: It's like finding out which specific pair of teams caused the difference.

Code:

   TukeyHSD(aov_model)

Mini summary: Post-hoc tests identify which groups differ.


Lesson 9: Chi-Square Test – For Categorical Data

Definition: The chi-square test is used when data is categorical (counts in categories).

Why it is important: It tests if two categorical variables are related.

Simple explanation: It's like checking if favourite foods are different between boys and girls.

Real-life example: Test if gender and favourite subject are related.

School example: Test if class and preference for school lunch are related.

Home example: Test if age and favourite TV show are related.

Nigerian example: Test if region and voting preference are related.

Code:

   table_data <- table(data$gender, data$favourite_food)
   chisq.test(table_data)
   Chi-square output:
   X-squared = 10.2, df = 2, p-value = 0.006

Mini summary: Chi-square tests relationships between categories.


Lesson 10: Interpreting Chi-Square Results

If p < 0.05, the variables are related (dependent). If p > 0.05, they are not related (independent).

Mini summary: p < 0.05 = relationship exists.


Lesson 11: Real Nigerian Example – Test on Education and Income

Let's say we have data on education level (primary, secondary, tertiary) and income level (low, medium, high). We want to see if education and income are related.

   table_data <- table(education, income)
   chisq.test(table_data)
   # If p < 0.05, education and income are related.

Mini summary: Chi-square is great for survey data.


Lesson 12: Choosing the Right Test

QuestionData TypeTest
Compare two groupsNumerict-test
Compare three+ groupsNumericANOVA
Relationship between categoriesCategoricalChi-square

Mini summary: Choose the test based on your data and question.


Lesson 13: Statistical Significance vs Practical Significance

Definition: Statistical significance means the result is unlikely to be chance. Practical significance means the result is big enough to matter.

Why it is important: A difference can be statistically significant but very small and unimportant.

Simple explanation: A small difference might be real (statistically significant) but not meaningful (e.g., 0.1 kg weight loss).

Mini summary: Always consider if the result is meaningful in real life.


Lesson 14: Common Errors in Testing

  • Using the wrong test.
  • Ignoring assumptions.
  • Only looking at p-value, not effect size.
  • Over-interpreting a significant result.
  • Data dredging (testing many things until something is significant).

Mini summary: Be careful not to misuse tests.


Lesson 15: Summary of Statistical Testing

We learned about p-values, t-tests, ANOVA, and chi-square tests. Testing helps us make decisions based on data.


Key Vocabulary

Statistical Test
A method to decide if a pattern is real.
p-value
Probability that the result is due to chance.
Significant
p < 0.05 – the result is likely real.
Null Hypothesis
The assumption of no effect.
t-test
Test comparing two means.
ANOVA
Test comparing three or more means.
Chi-square
Test for categorical relationships.
Post-hoc
Test to find which groups differ after ANOVA.

Important Concepts

  • p-value is the key decision metric.
  • Null hypothesis is always "no effect".
  • t-test for two groups, ANOVA for three+ groups.
  • Chi-square for categorical data.
  • Always check assumptions.

Step-by-Step Explanations

How to perform a t-test:

  1. Clean and prepare your data.
  2. Check assumptions (normality, equal variance).
  3. Use t.test() with your data.
  4. Look at p-value.
  5. If p < 0.05, conclude a significant difference.

Real-life Examples

  • A company tests if a new marketing strategy increases sales.
  • A doctor tests if a new drug lowers blood pressure.
  • A school tests if a new teaching method improves scores.

Nigerian Examples

  • Test if rural and urban areas differ in average income.
  • Test if fertiliser type affects crop yield (ANOVA).
  • Test if region and preferred political party are related (chi-square).

Fun Examples Children Relate To

  • Test if boys and girls have different average heights (t-test).
  • Test if three different cereals are equally tasty (ANOVA).
  • Test if favourite colour is related to pet preference (chi-square).

Everyday Examples

  • You test if you score better after eating breakfast.
  • You test if your running speed differs by day of the week.
  • You test if your favourite food is chosen equally by your friends.

Teacher Notes

  • Use relatable examples to explain concepts.
  • Emphasise that p-value is a tool, not a magic rule.
  • Show how to use R for testing.
  • Discuss the meaning of "significance" in context.

Parent Tips

  • Help your child understand that testing is like making a decision.
  • Discuss examples from daily life where you make decisions based on evidence.
  • Encourage them to ask questions and test their ideas.

Interesting Facts

  • The t-test was invented by William Gossett in 1908.
  • ANOVA was developed by Ronald Fisher.
  • Chi-square is used in many fields, including medicine and business.

Did You Know?

  • Some researchers are moving away from p-value only and using confidence intervals.
  • There are non-parametric tests when assumptions are not met.
  • R has many built-in functions for statistical testing.

Remember This

  • p < 0.05 is significant.
  • Choose the right test for your data.
  • Check assumptions.
  • Consider practical significance too.

Common Mistakes

  • Using t-test on more than two groups.
  • Ignoring assumptions (e.g., normality).
  • Interpreting non-significant results as "no effect" (they could be due to small sample size).
  • Fishing for significance (doing many tests until one is significant).

Best Practices

  • Plan your analysis before testing.
  • Use the appropriate test.
  • Check assumptions visually (e.g., histograms).
  • Report effect size and confidence intervals.
  • Be honest about limitations.

ASCII Illustrations

t-test concept

   Group A:  ****      Group B:  ***
             (mean = 75)          (mean = 69)
   t-test checks if the gap is real.
   Gap (difference)
   |--------------|
   75             69

ANOVA concept

   Group A:  ****
   Group B:  ***
   Group C:  *****
   ANOVA checks if any group differs.

Chi-square concept

   Table of counts:
             Food A  Food B
   Boys       10       5
   Girls      6        9
   Chi-square checks if food preference is independent of gender.

Comparison Tables

Statistical Tests
TestUseData TypeR Function
t-testCompare two groupsNumerict.test()
ANOVACompare three+ groupsNumericaov()
Chi-squareTest relationshipCategoricalchisq.test()

End-of-Module Summary

In this module, we learned about statistical testing – a way to make decisions based on data. We learned about the p-value and how it helps us decide if a result is significant (p < 0.05). We used the t-test for two groups, ANOVA for three or more groups, and chi-square for categorical data. We also learned about assumptions and choosing the right test. Testing is a powerful tool in data analysis.


Frequently Asked Questions

1. What is a p-value?
The probability that the result is due to chance.
2. What does p < 0.05 mean?
The result is significant (real).
3. What is a t-test?
A test to compare two groups.
4. What is ANOVA?
A test to compare more than two groups.
5. What is chi-square?
A test for categorical data.
6. What is the null hypothesis?
The assumption of no effect.
7. When should I use t-test?
When comparing two groups.
8. When should I use ANOVA?
When comparing three or more groups.
9. What are assumptions?
Conditions that must be met for a valid test.
10. What is practical significance?
Whether the effect is big enough to matter.

Review Questions

  1. What is a p-value?
  2. What does p < 0.05 mean?
  3. What is the null hypothesis?
  4. What is a t-test used for?
  5. What is ANOVA used for?
  6. What is chi-square used for?
  7. What are assumptions in testing?
  8. What is the difference between statistical and practical significance?
  9. How do you choose the right test?
  10. What is a post-hoc test?
  11. What does t.test() do in R?
  12. What does aov() do in R?
  13. What does chisq.test() do in R?
  14. Why is it important to check assumptions?
  15. What is the danger of data dredging?

Fill-in-the-Blank Exercises

  1. A ______ tells us if a result is real or due to chance. (p-value)
  2. If p < ______, the result is significant. (0.05)
  3. The ______ hypothesis assumes no effect. (null)
  4. A ______ compares two groups. (t-test)
  5. ______ compares three or more groups. (ANOVA)
  6. ______ tests relationships between categories. (Chi-square)
  7. ______ tests identify which groups differ after ANOVA. (Post-hoc)
  8. ______ significance means the effect is big enough to matter. (Practical)
  9. ______ are conditions that must be met for a valid test. (Assumptions)
  10. ______ means testing many things to find a significant result. (Data dredging)

True or False Exercises

  1. A p-value of 0.01 means the result is significant. (True)
  2. p < 0.05 means the result is due to chance. (False)
  3. t-test compares more than two groups. (False)
  4. ANOVA compares two groups. (False)
  5. Chi-square is for categorical data. (True)
  6. Assumptions are not important. (False)
  7. Post-hoc tests are used after ANOVA. (True)
  8. Statistical significance always means practical significance. (False)
  9. Data dredging is a good practice. (False)
  10. aov() is used for ANOVA in R. (True)

Multiple Choice Questions

  1. What does a p-value of 0.03 mean?
    A. Significant B. Not significant C. Need more data D. Error
    Answer: A
  2. Which test compares two groups?
    A. t-test B. ANOVA C. Chi-square D. Post-hoc
    Answer: A
  3. Which test compares three or more groups?
    A. t-test B. ANOVA C. Chi-square D. None
    Answer: B
  4. Which test is for categorical data?
    A. t-test B. ANOVA C. Chi-square D. Post-hoc
    Answer: C
  5. What is the null hypothesis?
    A. There is an effect B. There is no effect C. The test is valid D. The data is clean
    Answer: B
  6. What does significant mean?
    A. p < 0.05 B. p > 0.05 C. p = 0.05 D. p is large
    Answer: A
  7. What is a post-hoc test?
    A. Test after ANOVA B. Test before ANOVA C. Test for categorical data D. Test for two groups
    Answer: A
  8. What is practical significance?
    A. The effect is real B. The effect is big enough to matter C. The p-value is small D. The test is valid
    Answer: B
  9. What is data dredging?
    A. Cleaning data B. Testing many things to find significance C. Visualising data D. Importing data
    Answer: B
  10. Which function performs t-test in R?
    A. t.test() B. aov() C. chisq.test() D. TukeyHSD()
    Answer: A
  11. Which function performs ANOVA in R?
    A. t.test() B. aov() C. chisq.test() D. TukeyHSD()
    Answer: B
  12. Which function performs chi-square in R?
    A. t.test() B. aov() C. chisq.test() D. TukeyHSD()
    Answer: C
  13. What is an assumption?
    A. A condition for a valid test B. A conclusion C. A p-value D. A hypothesis
    Answer: A
  14. Why are assumptions important?
    A. They make the test faster B. They ensure the test is valid C. They are optional D. They are not important
    Answer: B
  15. What does TukeyHSD do?
    A. t-test B. ANOVA C. Post-hoc test D. Chi-square test
    Answer: C

Matching Exercises

TestUse
1. t-testA. Categorical data
2. ANOVAB. Compare two groups
3. Chi-squareC. Compare three+ groups

Answers: 1-B, 2-C, 3-A


Short Answer Questions

  1. What is a p-value and what does it tell us?
  2. Explain the difference between t-test and ANOVA.
  3. When would you use a chi-square test?
  4. What are the assumptions of a t-test?
  5. Why is practical significance important?

Scenario-based Exercises

  1. Scenario: You want to compare test scores of students from three different schools. Which test do you use?
  2. Scenario: You want to see if gender and favourite subject are related. Which test do you use?
  3. Scenario: You want to compare the average height of boys and girls. Which test do you use?

Group Activity

In groups, create a dataset with two groups (e.g., boys and girls) and a numeric variable (e.g., test score). Perform a t-test and interpret the results. Then, create a categorical dataset and perform a chi-square test.


Individual Activity

Use the iris dataset. Perform a t-test to compare petal length between two species (e.g., setosa and versicolor). Perform ANOVA to compare petal length among all three species.


Classroom Discussion Questions

  1. What is the meaning of a p-value in your own words?
  2. How would you explain statistical significance to a friend?
  3. Can you think of a situation where a result is statistically significant but not practically important?
  4. Why is it important to check assumptions before testing?

Mini Project

Title: "Testing a Nigerian Hypothesis"
Find a dataset from Nigeria (e.g., education, health, or agriculture). Formulate a hypothesis (e.g., "Rural and urban areas have different access to clean water"). Use an appropriate statistical test to test your hypothesis. Write a report of your findings.


Practical Assignment

Using the mtcars dataset, compare the mpg (miles per gallon) of automatic and manual transmission cars using a t-test. Then, perform ANOVA to compare mpg across different numbers of cylinders (4, 6, 8).


Challenge Exercise

Find a real dataset online. Perform at least three different statistical tests on it (t-test, ANOVA, chi-square). Write a complete report explaining the data, the tests used, and the conclusions.


Quiz Answers

Fill-in-the-Blank: 1. p-value, 2. 0.05, 3. null, 4. t-test, 5. ANOVA, 6. Chi-square, 7. Post-hoc, 8. Practical, 9. Assumptions, 10. Data dredging.

True/False: 1T, 2F, 3F, 4F, 5T, 6F, 7T, 8F, 9F, 10T.

Multiple Choice: 1A, 2A, 3B, 4C, 5B, 6A, 7A, 8B, 9B, 10A, 11B, 12C, 13A, 14B, 15C.


Key Takeaways

  • p-value helps us decide if a result is significant.
  • t-test compares two groups.
  • ANOVA compares three or more groups.
  • Chi-square tests relationships between categories.
  • Always check assumptions.
  • Consider both statistical and practical significance.

Preparation for the Next Module

In Module 15, we will learn about regression – how to predict one variable using another. We will use linear regression to model relationships and make predictions. Practise testing and reflecting on how tests guide decisions.


16

Module Fifteen

Module 15 Β· R for Data Analysis Β· Regression – Predicting the Future

Module 15 Β· Regression – Predicting the Future

Hello, young data explorer! In Module 14, we learned how to test if patterns are real. Now, we will learn how to predict one variable from another. Imagine you know a person's height – can you predict their weight? Or if you know how many hours you study, can you predict your test score? This is what regression does.

Regression helps us understand the relationship between two or more variables and use that to make predictions. The simplest type is linear regression, where we draw a straight line through our data to predict outcomes.

By the end of this module, you will be able to build regression models, interpret them, and make predictions. You will be a data predictor!


Learning Objectives

  • Understand what regression is and why it is useful.
  • Learn the concept of linear regression.
  • Use lm() to build a regression model in R.
  • Interpret the coefficients (slope and intercept).
  • Make predictions using the model.
  • Evaluate the model using R-squared.
  • Understand residuals and how to check them.
  • Use regression on real data, including Nigerian examples.
  • Know the difference between simple and multiple regression.

Warm-up Story: The Ice Cream Sales Prediction

Chidi sells ice cream at school. He noticed that on hot days, he sells more ice cream. He wanted to predict how many ice creams he would sell based on the temperature. He collected data for 10 days: temperature and sales.

He used linear regression to draw a line through the data points. The line showed that for every degree increase in temperature, he sells 2 more ice creams. Now, when the weather forecast says it will be 30Β°C, he predicts he will sell about 60 ice creams. He can prepare enough stock.

Chidi learned that regression helps him make better decisions!


Main Lessons

Lesson 1: What is Regression?

Definition: Regression is a statistical method to model the relationship between a dependent (response) variable and one or more independent (predictor) variables.

Why it is important: It helps us understand how variables are related and make predictions.

Simple explanation: It's like finding a line that best fits the data points.

Real-life example: A real estate agent uses regression to predict house prices based on size and location.

School example: A teacher predicts test scores based on study time.

Home example: You predict your allowance based on your chores.

Nigerian example: A farmer predicts crop yield based on rainfall.

Illustration:

   Data points (dots) --> Fit a line --> Use line to predict new points

Mini summary: Regression models relationships and predicts outcomes.


Lesson 2: Linear Regression – The Straight Line

Definition: Linear regression fits a straight line to the data: y = a + b*x, where:

  • y is the dependent variable (what we predict).
  • x is the independent variable (what we use to predict).
  • a is the intercept (value of y when x = 0).
  • b is the slope (how much y changes when x increases by 1).

Simple explanation: It's like drawing a straight line through the cloud of dots.

Real-life example: y = 2x + 10, where x is temperature, y is ice cream sales.

School example: y = 2*study_time + 50 (predicting score).

Home example: y = 1.5*chores + 5 (predicting allowance).

Nigerian example: y = 0.5*rainfall + 20 (predicting crop yield).

Illustration:

   y
   ^
   |   / (line: y = a + b*x)
   |  /
   | /
   |/__________________ x

Mini summary: Linear regression fits a straight line to data.


Lesson 3: Building a Regression Model in R – lm()

Definition: lm() is the R function for linear regression. It stands for "linear model".

Why it is important: It's the standard way to perform regression in R.

Simple explanation: It's like giving R a recipe to draw the best line.

Code:

   model <- lm(y ~ x, data = data)
   # y ~ x means "y is predicted by x"
   Example:
   model <- lm(score ~ study_time, data = students)

Mini summary: Use lm() to create a regression model.


Lesson 4: Interpreting the Output – summary()

Definition: summary(model) gives detailed output including coefficients, R-squared, and p-values.

Why it is important: It tells you how good the model is and the strength of relationships.

Simple explanation: It's like a report card for your model.

Code:

   summary(model)
   Output example:
   Call:
   lm(formula = score ~ study_time, data = students)

   Residuals:
       Min      1Q  Median      3Q     Max
   -8.345  -3.234   0.123   2.456   7.890

   Coefficients:
               Estimate Std. Error t value Pr(>|t|)
   (Intercept)  50.1234     2.3456  21.376   <2e-16 ***
   study_time    2.3456     0.4567   5.135   0.0002 ***

   Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

   Residual standard error: 4.567 on 28 degrees of freedom
   Multiple R-squared:  0.485,    Adjusted R-squared:  0.467
   F-statistic: 26.38 on 1 and 28 DF,  p-value: 0.0002

Mini summary: summary() evaluates your model.


Lesson 5: Coefficients – The Intercept and Slope

Definition:

  • Intercept: The value of y when x is 0.
  • Slope: The change in y for a 1-unit change in x.

Why they are important: They define the line and tell you the relationship.

Simple explanation: Intercept is where the line starts; slope is how steep it is.

Real-life example: In our ice cream example, intercept = 10 (sales when 0Β°C) and slope = 2 (each degree adds 2 sales).

School example: Intercept = 50, slope = 2 (score = 50 + 2*study_time).

Home example: Intercept = 5, slope = 1.5 (allowance = 5 + 1.5*chores).

Nigerian example: Intercept = 20, slope = 0.5 (yield = 20 + 0.5*rainfall).

Mini summary: Coefficients tell you the line's equation.


Lesson 6: R-squared – How Good is the Model?

Definition: R-squared (RΒ²) measures the proportion of variance in y that is explained by x. It ranges from 0 to 1.

Why it is important: It tells you how well the line fits the data.

Simple explanation: An RΒ² of 0.8 means 80% of the variation in y is explained by x.

Real-life example: RΒ² = 0.75 means temperature explains 75% of the variation in ice cream sales.

School example: RΒ² = 0.65 means study time explains 65% of score variation.

Home example: RΒ² = 0.50 means chores explain 50% of allowance variation.

Nigerian example: RΒ² = 0.70 means rainfall explains 70% of crop yield variation.

Rule: Higher RΒ² means better fit.

Mini summary: RΒ² measures how well the model predicts.


Lesson 7: Making Predictions – predict()

Definition: predict() uses the model to predict y for new x values.

Why it is important: It lets you forecast future outcomes.

Simple explanation: It's like using the line to read off predicted values.

Code:

   new_data <- data.frame(study_time = c(5, 6, 7))
   predict(model, newdata = new_data)

Mini summary: predict() makes predictions.


Lesson 8: Residuals – The Errors

Definition: Residuals are the differences between actual and predicted y values.

Why they are important: They show how far off the model is. Good models have small residuals.

Simple explanation: It's like the mistakes the model makes.

Code:

   residuals(model)

Mini summary: Residuals measure the model's errors.


Lesson 9: Checking Assumptions

Linear regression assumes:

  • Linearity: The relationship is roughly straight.
  • Independence: Observations are independent.
  • Homoscedasticity: Constant variance of residuals.
  • Normality: Residuals are roughly normal.

Why it is important: Violations can make the model unreliable.

Code:

   plot(model)  # diagnostic plots

Mini summary: Check assumptions to ensure valid model.


Lesson 10: Simple vs Multiple Regression

Definition:

  • Simple: One predictor (x).
  • Multiple: Two or more predictors (x1, x2, ...).

Why it is important: Multiple regression can capture more complex relationships.

Code:

   model_multiple <- lm(score ~ study_time + sleep_hours, data = students)

Mini summary: Multiple regression uses multiple predictors.


Lesson 11: Real Nigerian Example – Predicting Crop Yield

Let's predict crop yield based on rainfall and fertiliser usage.

   model <- lm(yield ~ rainfall + fertiliser, data = farms)
   summary(model)
   # Predict yield for a farm with 100mm rainfall and 50kg fertiliser
   predict(model, newdata = data.frame(rainfall = 100, fertiliser = 50))

Mini summary: Regression is useful for Nigerian agriculture.


Lesson 12: Visualising Regression – Adding the Line

Definition: Add the regression line to a scatter plot.

Code:

   ggplot(data, aes(x = study_time, y = score)) +
       geom_point() +
       geom_smooth(method = "lm", se = FALSE) +
       labs(title = "Scatter Plot with Regression Line")
   Scatter plot with line:
   Score
   100 |        .    /
    80 |      .   . /
    60 |    .     / .
    40 |  .     /
    20 |.     /
     0 |___._.__/.___.___.___
        0   20  40  60  80 100
        Study Time

Mini summary: Visualising the line helps interpret the model.


Lesson 13: Overfitting – Too Much of a Good Thing

Definition: Overfitting is when a model fits the training data too closely but performs poorly on new data.

Why it is important: A good model should generalise to new data.

Simple explanation: It's like memorising a test instead of learning the material.

Mini summary: Avoid overfitting by keeping the model simple.


Lesson 14: Categorical Predictors

You can use categorical variables (like gender) as predictors. R automatically converts them to dummy variables.

   model <- lm(score ~ gender, data = data)

Mini summary: Categorical predictors can be used in regression.


Lesson 15: Summary of Regression

Regression models relationships and makes predictions. We learned about simple and multiple regression, coefficients, R-squared, and assumptions.


Key Vocabulary

Regression
Modeling the relationship between variables.
Linear Regression
Using a straight line to model relationships.
Intercept
Value of y when x = 0.
Slope
Change in y for a 1-unit change in x.
R-squared
Proportion of variance explained.
Residuals
Differences between actual and predicted values.
Predictor
Independent variable (x).
Response
Dependent variable (y).
Overfitting
Model that fits training data too well.

Important Concepts

  • Linear regression fits a line.
  • R-squared measures fit.
  • Coefficients define the line.
  • Residuals show errors.
  • Assumptions must be checked.

Step-by-Step Explanations

How to build a regression model:

  1. Prepare and clean your data.
  2. Choose your response (y) and predictor (x).
  3. Use lm(y ~ x, data = data).
  4. Use summary() to view results.
  5. Check R-squared and p-values.
  6. Make predictions with predict().
  7. Check assumptions with diagnostic plots.

Real-life Examples

  • A business predicts sales based on advertising spend.
  • A doctor predicts patient recovery time based on age.
  • A school predicts student performance based on attendance.

Nigerian Examples

  • Predict crop yield based on rainfall and fertiliser.
  • Predict house prices based on size and location.
  • Predict student scores based on study time and resources.

Fun Examples Children Relate To

  • Predict your score based on how many hours you play video games.
  • Predict your pocket money based on chores completed.
  • Predict the number of friends at your party based on the weather.

Everyday Examples

  • You predict how long it will take you to get to school based on time of day.
  • You predict how much you will eat based on how hungry you are.
  • You predict your sleep quality based on your bedtime.

Teacher Notes

  • Use visual demonstrations to show regression lines.
  • Emphasise that correlation does not imply causation.
  • Discuss the importance of model evaluation.
  • Use real datasets for practice.

Parent Tips

  • Help your child collect data at home and make predictions.
  • Discuss how regression is used in everyday life.
  • Encourage them to think about cause and effect.

Interesting Facts

  • Linear regression was first used in the 19th century.
  • It is one of the most widely used statistical methods.
  • Many machine learning algorithms are based on regression.

Did You Know?

  • You can perform regression with non-linear relationships using transformations.
  • R has packages for advanced regression like glm for logistic regression.
  • Regression is used in weather forecasting, economics, and medicine.

Remember This

  • Regression models relationships.
  • Linear regression uses a straight line.
  • R-squared tells you how good the model is.
  • Check assumptions before trusting the model.
  • Use predict() for forecasting.

Common Mistakes

  • Using regression when relationships are not linear.
  • Ignoring assumptions.
  • Confusing correlation with causation.
  • Overfitting the model.
  • Not checking residuals.

Best Practices

  • Visualise data before modelling.
  • Check for outliers.
  • Use multiple predictors when appropriate.
  • Validate the model with new data.
  • Interpret coefficients carefully.

ASCII Illustrations

Regression line

   y
   ^
   |   /    (line)
   |  /
   | /
   |/
   +--------------------> x

Residuals

   Actual points:   *     *
   Predicted line:  /
   Residuals:       | (vertical distances)

Comparison Tables

Simple vs Multiple Regression
FeatureSimpleMultiple
Number of predictors12 or more
Equationy = a + b*xy = a + b1*x1 + b2*x2 + ...
UseSimple relationshipsComplex relationships

End-of-Module Summary

In this module, we learned about regression – a tool for predicting one variable from another. We used linear regression to fit a line to data, interpreted coefficients and R-squared, and made predictions using predict(). We also discussed assumptions, residuals, and the difference between simple and multiple regression. Regression is a powerful way to understand and forecast data.


Frequently Asked Questions

1. What is regression?
A method to model relationships and make predictions.
2. What is linear regression?
Using a straight line to model relationships.
3. What is the intercept?
The value of y when x = 0.
4. What is the slope?
Change in y for a 1-unit change in x.
5. What is R-squared?
Proportion of variance explained by the model.
6. What are residuals?
Differences between actual and predicted values.
7. How do you make predictions?
Use predict().
8. What is the difference between simple and multiple regression?
Simple has one predictor; multiple has two or more.
9. What are assumptions of regression?
Linearity, independence, homoscedasticity, normality.
10. What is overfitting?
When a model fits training data too closely but performs poorly on new data.

Review Questions

  1. What is regression?
  2. What is linear regression?
  3. What is the intercept in a regression equation?
  4. What is the slope?
  5. What does R-squared measure?
  6. What are residuals?
  7. How do you make predictions using a regression model in R?
  8. What is the difference between simple and multiple regression?
  9. What are the assumptions of linear regression?
  10. What is overfitting?
  11. What function in R is used to build a regression model?
  12. What function gives a summary of the model?
  13. How can you check assumptions?
  14. What is a categorical predictor?
  15. Why is it important to check residuals?

Fill-in-the-Blank Exercises

  1. ______ models relationships between variables. (Regression)
  2. ______ regression uses a straight line. (Linear)
  3. The ______ is the value of y when x = 0. (intercept)
  4. The ______ is the change in y for a 1-unit change in x. (slope)
  5. ______ measures the proportion of variance explained. (R-squared)
  6. ______ are the differences between actual and predicted values. (Residuals)
  7. Use ______ to make predictions from a model. (predict)
  8. ______ regression has more than one predictor. (Multiple)
  9. ______ happens when a model fits training data too well. (Overfitting)
  10. ______ are conditions that must be met for a valid model. (Assumptions)

True or False Exercises

  1. Regression is used to make predictions. (True)
  2. Linear regression fits a curved line. (False)
  3. The intercept is the slope. (False)
  4. R-squared tells you how well the model fits. (True)
  5. Residuals are the predicted values. (False)
  6. Multiple regression has two or more predictors. (True)
  7. Overfitting is good for prediction. (False)
  8. Assumptions must be checked for a valid model. (True)
  9. predict() is used to build a model. (False)
  10. lm() is the function for linear regression. (True)

Multiple Choice Questions

  1. What is regression used for?
    A. Cleaning data B. Predicting C. Visualising D. Importing
    Answer: B
  2. Which function builds a regression model in R?
    A. lm() B. glm() C. predict() D. summary()
    Answer: A
  3. What does the intercept represent?
    A. y when x=0 B. slope C. R-squared D. residual
    Answer: A
  4. What does the slope represent?
    A. y when x=0 B. change in y per unit x C. R-squared D. residual
    Answer: B
  5. What does R-squared measure?
    A. Model fit B. Slope C. Intercept D. Residuals
    Answer: A
  6. What are residuals?
    A. Predicted values B. Actual values C. Differences D. Slopes
    Answer: C
  7. Which function makes predictions?
    A. lm() B. summary() C. predict() D. plot()
    Answer: C
  8. What is multiple regression?
    A. One predictor B. Two or more predictors C. No predictors D. Categorical predictor
    Answer: B
  9. What is overfitting?
    A. Good model B. Model that fits training data too well C. Simple model D. Model with high R-squared
    Answer: B
  10. Which is NOT an assumption of linear regression?
    A. Linearity B. Independence C. Normality D. Multicollinearity
    Answer: D
  11. What does summary(model) show?
    A. Coefficients B. R-squared C. p-values D. All of the above
    Answer: D
  12. Which plot helps check assumptions?
    A. Scatter plot B. Residual plot C. Bar chart D. Pie chart
    Answer: B
  13. Can categorical variables be used as predictors?
    A. Yes B. No C. Only if numeric D. Only in simple regression
    Answer: A
  14. What is the equation of a simple linear regression?
    A. y = a + b*x B. y = a*x + b C. y = a*x D. y = a + b*x^2
    Answer: A
  15. Why check residuals?
    A. To see predictions B. To check assumptions C. To increase R-squared D. To change slope
    Answer: B

Matching Exercises

TermDefinition
1. InterceptA. Change in y per unit x
2. SlopeB. Value of y when x=0
3. R-squaredC. Measure of fit
4. ResidualD. Difference between actual and predicted

Answers: 1-B, 2-A, 3-C, 4-D


Short Answer Questions

  1. What is linear regression?
  2. Explain the meaning of the slope and intercept.
  3. What does R-squared tell you?
  4. What are residuals and why are they important?
  5. What is the difference between simple and multiple regression?

Scenario-based Exercises

  1. Scenario: You want to predict students' final exam scores based on their midterm scores. Which method do you use?
  2. Scenario: You want to predict house prices based on size, number of bedrooms, and location. Which method do you use?
  3. Scenario: You have data on temperature and ice cream sales. You want to predict sales for a 30Β°C day. What do you do?

Group Activity

In groups, collect data on study time and test scores (or use built-in data). Build a regression model, interpret the results, and make predictions. Present your findings.


Individual Activity

Use the mtcars dataset to build a regression model predicting mpg (miles per gallon) from hp (horsepower). Interpret the coefficients and R-squared. Make a prediction for a car with 150 hp.


Classroom Discussion Questions

  1. How can regression help in business decision-making?
  2. What are the risks of using regression incorrectly?
  3. How can regression be used in Nigerian agriculture?
  4. What would you do if a regression model has low R-squared?

Mini Project

Title: "Predicting Nigerian Student Performance"
Find or create a dataset of Nigerian student scores and study habits. Build a multiple regression model to predict scores. Identify the most important predictors. Write a report with your findings.


Practical Assignment

Using the airquality dataset, build a regression model to predict ozone levels based on temperature and wind speed. Interpret the model and check assumptions.


Challenge Exercise

Find a real dataset online. Build a multiple regression model with at least 3 predictors. Evaluate the model, check assumptions, and make predictions. Write a comprehensive report.


Quiz Answers

Fill-in-the-Blank: 1. Regression, 2. Linear, 3. intercept, 4. slope, 5. R-squared, 6. Residuals, 7. predict, 8. Multiple, 9. Overfitting, 10. Assumptions.

True/False: 1T, 2F, 3F, 4T, 5F, 6T, 7F, 8T, 9F, 10T.

Multiple Choice: 1B, 2A, 3A, 4B, 5A, 6C, 7C, 8B, 9B, 10D, 11D, 12B, 13A, 14A, 15B.


Key Takeaways

  • Regression predicts one variable from another.
  • Linear regression uses a straight line.
  • R-squared measures model fit.
  • Residuals are the errors.
  • Multiple regression uses many predictors.
  • Always check assumptions.

Preparation for the Next Module

In Module 16, we will learn about data communication – how to present your findings clearly. We will combine all our skills to create reports and dashboards. Practise building regression models and interpreting their outputs.


17

Module Sixteen

Module 16 Β· R for Data Analysis Β· Data Communication

Module 16 Β· Data Communication – Telling Your Data Story

Hello, young data explorer! In Module 15, we learned how to predict outcomes using regression. But what good is an analysis if you cannot share it with others? This module is about data communication – how to tell a clear and compelling story with your data.

Data communication is like being a storyteller. You take your data, your analysis, and your insights, and you present them in a way that is easy to understand. You use reports, slides, dashboards, and visualisations to share your findings.

By the end of this module, you will be able to create a data report, use R Markdown, and present your work confidently. You will be a data communicator!


Learning Objectives

  • Understand what data communication is.
  • Learn the principles of effective communication.
  • Use R Markdown to create reports.
  • Create clear and informative visualisations.
  • Write a narrative around data.
  • Use tables to summarise findings.
  • Present your work to an audience.
  • Apply these skills to Nigerian data.

Warm-up Story: The School Data Presentation

Chidi had spent weeks analysing data on school attendance and test scores. He had found interesting patterns, but his teacher asked him to present his findings to the class. He needed to communicate his results clearly.

He used R Markdown to create a report. He included a title, an introduction, his data, his analysis, and his conclusions. He added graphs and tables to make it visual. He also prepared a short presentation. When he presented, the class understood everything. They even asked good questions!

Chidi learned that good communication makes your hard work useful.


Main Lessons

Lesson 1: What is Data Communication?

Definition: Data communication is the process of sharing your data findings with others in a clear and effective way.

Why it is important: If you cannot communicate your results, your analysis has no impact.

Simple explanation: It's like telling a story – you need a beginning, middle, and end.

Real-life example: A business analyst presents sales data to the CEO.

School example: A student presents a science project to the class.

Home example: You tell your family about your savings.

Nigerian example: A researcher presents findings on agriculture to farmers.

Illustration:

   Data --> Analysis --> Insights --> Communication --> Impact

Mini summary: Data communication is sharing your findings to create impact.


Lesson 2: Know Your Audience

Definition: Understand who you are communicating with and what they need to know.

Why it is important: Different audiences need different levels of detail.

Simple explanation: You talk differently to a friend than to a teacher.

Real-life example: You explain a game to a younger child differently than to a friend.

School example: You present your project to the teacher differently than to your classmates.

Home example: You explain your daily routine to a visitor.

Nigerian example: You present data to a community leader differently than to a government official.

Mini summary: Tailor your message to your audience.


Lesson 3: The Structure of a Data Report

A good report has:

  1. Title – What is the report about?
  2. Introduction – Why did you do the analysis?
  3. Data – Where did the data come from?
  4. Analysis – What did you find?
  5. Visualisations – Show the results.
  6. Conclusions – What did you learn?
  7. Recommendations – What should be done?

Mini summary: Reports should be clear and structured.


Lesson 4: Introduction to R Markdown

Definition: R Markdown is a tool that combines R code and narrative text to create dynamic documents (HTML, PDF, Word).

Why it is important: It makes reproducible reporting easy.

Simple explanation: It's like a notebook where you write text and code together.

Code:

   ---
   title: "My Data Report"
   author: "Chidi"
   date: "2026-07-04"
   output: html_document
   ---

   ```{r}
   summary(cars)
   ```

Mini summary: R Markdown combines code and text in one document.


Lesson 5: Creating an R Markdown Document

Steps:

  1. Open RStudio.
  2. Click File > New File > R Markdown.
  3. Choose an output format (HTML, PDF, Word).
  4. Write your text and code in chunks.
  5. Click Knit to create the report.

Mini summary: R Markdown is easy to use in RStudio.


Lesson 6: Writing in R Markdown

Simple formatting:

  • # for headings (e.g., # Introduction)
  • * or _ for italics
  • ** or __ for bold
  • - for bullet points
  • Code chunks with ```{r}
   # This is a heading
   This is *italic* and **bold**.
   - Bullet point

Mini summary: R Markdown uses simple formatting.


Lesson 7: Adding Code Chunks

Definition: Code chunks are sections where you write R code. They are enclosed in ```{r} ... ```.

Why it is important: They run your analysis and show results in the report.

   ```{r}
   library(ggplot2)
   ggplot(mtcars, aes(x = hp, y = mpg)) +
       geom_point()
   ```

Mini summary: Code chunks run R code in your report.


Lesson 8: Inline R Code

Definition: You can embed R code within text using `r code`.

Why it is important: It makes dynamic text, e.g., "The mean is `r mean(data$score)`".

   The average score is `r mean(scores)`.

Mini summary: Inline code makes reports dynamic.


Lesson 9: Visualising for Communication

Definition: Good visualisations are clear and informative. They should tell a story.

Why it is important: A picture is worth a thousand words.

Tips:

  • Use simple, clear titles.
  • Label axes properly.
  • Use colours wisely.
  • Avoid clutter.
  • Include captions.
   ggplot(data, aes(x = study_time, y = score)) +
       geom_point() +
       labs(title = "Study Time vs Score",
            x = "Study Time (hours)",
            y = "Test Score") +
       theme_minimal()

Mini summary: Clear visualisations enhance communication.


Lesson 10: Using Tables in Reports

Definition: Tables organise and present numeric data clearly.

Why it is important: They provide detailed information.

Code:

   library(knitr)
   kable(head(data), caption = "First few rows of data")

Mini summary: Tables present data in a structured way.


Lesson 11: Writing a Narrative

Definition: A narrative is the story you tell about your data. It connects the data to the real world.

Why it is important: It makes your report engaging and meaningful.

Simple explanation: Instead of just showing numbers, explain what they mean.

Example: "The data shows that students who studied more got higher scores. This suggests that study time is important for success."

Mini summary: A narrative makes data meaningful.


Lesson 12: Presenting with Slides

Definition: R Markdown can also create presentations (e.g., reveal.js, ioslides).

Why it is important: Slides are great for live presentations.

   ---
   title: "My Presentation"
   output: ioslides_presentation
   ---

Mini summary: R Markdown creates slide presentations.


Lesson 13: Dashboards with flexdashboard

Definition: Dashboards are interactive reports with multiple panels.

Why it is important: They allow users to explore data themselves.

Mini summary: Dashboards are interactive reports.


Lesson 14: Nigerian Example – Report on Education

Create a report on Nigerian education data: import, clean, analyse, and visualise. Include a narrative about the state of education.

Mini summary: Apply all skills to Nigerian data.


Lesson 15: Summary of Data Communication

We learned to communicate data effectively through reports, visualisations, and narratives. Good communication makes data valuable.


Key Vocabulary

Data Communication
Sharing data findings clearly.
R Markdown
A tool for combining code and text.
Report
A document presenting data findings.
Narrative
The story behind the data.
Visualisation
A graph or chart.
Dashboard
An interactive report.
Audience
The people you are communicating with.
Knit
Render the R Markdown document.

Important Concepts

  • Know your audience – tailor your message.
  • Structure – reports have a clear flow.
  • Visuals – make data easy to understand.
  • Narrative – connect data to the real world.
  • Reproducibility – R Markdown ensures your work is reproducible.

Step-by-Step Explanations

How to create an R Markdown report:

  1. Open RStudio.
  2. File > New File > R Markdown.
  3. Give it a title and author.
  4. Choose output format (HTML).
  5. Write your narrative in the text.
  6. Insert code chunks for analysis.
  7. Add visualisations and tables.
  8. Click Knit to generate the report.

Real-life Examples

  • A company creates a quarterly sales report.
  • A researcher publishes a paper with data analysis.
  • A teacher shares a report on student performance.

Nigerian Examples

  • A report on Nigerian agricultural yields.
  • A presentation on education statistics in Lagos.
  • A dashboard on COVID-19 cases in Nigeria.

Fun Examples Children Relate To

  • A report on your favourite video game scores.
  • A presentation on your pet's habits.
  • A dashboard on your chores completion.

Everyday Examples

  • You create a report on your weekly screen time.
  • You present your savings plan to your family.
  • You share a chart of your reading habits.

Teacher Notes

  • Emphasise the importance of clear communication.
  • Show examples of good and bad reports.
  • Encourage students to present their work.
  • Use R Markdown for assignments.

Parent Tips

  • Help your child create a report on something they are interested in.
  • Discuss what makes a good story.
  • Encourage them to share their findings with the family.

Interesting Facts

  • R Markdown was created by the same team as ggplot2.
  • It is used by data scientists worldwide.
  • R Markdown documents can include interactive elements.

Did You Know?

  • You can create a website with R Markdown.
  • R Markdown supports many languages (R, Python, SQL).
  • It can produce reports in over 20 formats.

Remember This

  • Good communication is key to impact.
  • Know your audience.
  • Use clear visuals and a structured narrative.
  • R Markdown makes reporting easy.
  • Always review your report for clarity.

Common Mistakes

  • Too much text or too many numbers.
  • Cluttered graphs.
  • Not explaining the meaning of the analysis.
  • Ignoring the audience.
  • Not checking for errors in the report.

Best Practices

  • Keep it simple and clear.
  • Use headlines and subheadings.
  • Highlight key insights.
  • Use captions for all visuals.
  • Review and refine your report.

ASCII Illustrations

Report structure

   +-----------------------+
   |       Title           |
   +-----------------------+
   |    Introduction       |
   +-----------------------+
   |      Data             |
   +-----------------------+
   |    Analysis           |
   +-----------------------+
   | Visualisations        |
   +-----------------------+
   |   Conclusions         |
   +-----------------------+

Comparison Tables

Communication Tools
ToolUseOutput
R MarkdownReportsHTML, PDF, Word
flexdashboardDashboardsHTML
ShinyInteractive appsWeb app
SlidesPresentationsHTML, PDF

End-of-Module Summary

In this module, we learned about data communication – how to share your data findings effectively. We used R Markdown to create reports that combine code, text, and visualisations. We learned about structuring reports, writing a narrative, and presenting to an audience. Good communication makes data analysis valuable and impactful.


Frequently Asked Questions

1. What is data communication?
Sharing data findings clearly.
2. What is R Markdown?
A tool for dynamic reports.
3. Why is R Markdown useful?
It combines code and text.
4. What is a narrative?
The story behind the data.
5. How do I create an R Markdown report?
New File > R Markdown.
6. What are code chunks?
Blocks of R code in R Markdown.
7. What is knitting?
Generating the report.
8. What makes a good visualisation?
Clear, simple, and labelled.
9. What is a dashboard?
An interactive report.
10. How do I choose the right communication tool?
Based on your audience and purpose.

Review Questions

  1. What is data communication?
  2. Why is it important?
  3. What is R Markdown?
  4. What is the structure of a good report?
  5. What is a narrative?
  6. How do you create an R Markdown document?
  7. What are code chunks?
  8. What does "knit" mean?
  9. What are the principles of good visualisation?
  10. What is a dashboard?
  11. What is the difference between a report and a dashboard?
  12. How do you tailor your communication to your audience?
  13. What are common mistakes in data communication?
  14. What are best practices?
  15. Give an example of a Nigerian data communication project.

Fill-in-the-Blank Exercises

  1. ______ is sharing data findings clearly. (Data communication)
  2. ______ combines code and text in a report. (R Markdown)
  3. The ______ is the story behind the data. (narrative)
  4. ______ are blocks of R code in R Markdown. (Code chunks)
  5. ______ generates the report from R Markdown. (Knit)
  6. A ______ is an interactive report. (dashboard)
  7. ______ should be clear, simple, and labelled. (Visualisations)
  8. Good communication is tailored to the ______. (audience)
  9. A report should have a ______, introduction, analysis, and conclusion. (title)
  10. ______ is a tool for creating presentations in R Markdown. (Slides)

True or False Exercises

  1. Data communication is not important. (False)
  2. R Markdown combines code and text. (True)
  3. A narrative is just the data. (False)
  4. Code chunks run R code. (True)
  5. Knitting creates a report. (True)
  6. Dashboards are static reports. (False)
  7. Good visualisations should be cluttered. (False)
  8. You should know your audience. (True)
  9. Reports have no structure. (False)
  10. R Markdown can create slides. (True)

Multiple Choice Questions

  1. What is R Markdown?
    A. A package B. A report format C. A tool for dynamic reports D. A visualisation tool
    Answer: C
  2. What does "knit" do?
    A. Runs code B. Generates report C. Cleans data D. Imports data
    Answer: B
  3. What is a narrative?
    A. A graph B. The story of data C. A table D. A code chunk
    Answer: B
  4. What are code chunks?
    A. Text blocks B. R code blocks C. Graphs D. Tables
    Answer: B
  5. What is a dashboard?
    A. A static report B. An interactive report C. A graph D. A table
    Answer: B
  6. What is the first step in creating an R Markdown report?
    A. Write code B. New File > R Markdown C. Knit D. Add visualisations
    Answer: B
  7. What is important in a visualisation?
    A. Complexity B. Clarity C. Many colours D. Small text
    Answer: B
  8. Why should you know your audience?
    A. To make it hard B. To tailor your message C. To ignore them D. To confuse them
    Answer: B
  9. What is the difference between a report and a dashboard?
    A. Report is interactive B. Dashboard is interactive C. No difference D. Both are static
    Answer: B
  10. Which of these is a common mistake?
    A. Clear visuals B. Cluttered graphs C. Simple narrative D. Structured report
    Answer: B
  11. What does inline code do?
    A. Runs code in text B. Creates graphs C. Imports data D. Cleans data
    Answer: A
  12. What is the output of R Markdown?
    A. HTML, PDF, Word B. Only HTML C. Only PDF D. Only Word
    Answer: A
  13. What is a good practice in data communication?
    A. Use many colours B. Highlight key insights C. Ignore conclusions D. No visuals
    Answer: B
  14. What does flexdashboard create?
    A. Reports B. Dashboards C. Slides D. Shiny apps
    Answer: B
  15. What is the benefit of reproducible reports?
    A. They are fast B. They can be updated easily C. They are always correct D. They are simple
    Answer: B

Matching Exercises

ConceptDescription
1. R MarkdownA. Interactive report
2. DashboardB. Dynamic reports
3. NarrativeC. Story behind data
4. VisualisationD. Graph or chart
5. AudienceE. People you communicate with

Answers: 1-B, 2-A, 3-C, 4-D, 5-E


Short Answer Questions

  1. What is data communication?
  2. Why is R Markdown useful?
  3. What is the structure of a good report?
  4. How do you tailor communication to an audience?
  5. What are the key elements of good visualisation?

Scenario-based Exercises

  1. Scenario: You have analysed data on student performance. You need to present it to the school principal. What would you include in your report?
  2. Scenario: You want to share your findings with a group of fellow students. How would you make the presentation engaging?
  3. Scenario: You have created a dashboard on Nigerian population data. How would you explain it to a government official?

Group Activity

In groups, analyse a dataset (e.g., iris). Create a report in R Markdown including an introduction, analysis, visualisations, and conclusions. Present your report to the class.


Individual Activity

Create an R Markdown report on a topic of your choice (e.g., your hobbies, school data). Include at least one visualisation and a narrative. Knit it to HTML and share it.


Classroom Discussion Questions

  1. What makes a data story compelling?
  2. How can we make data accessible to everyone?
  3. What is the role of visuals in communication?
  4. How can we ensure our communication is ethical and honest?

Mini Project

Title: "Nigerian Data Report"
Find a dataset about Nigeria (e.g., education, health, agriculture). Perform a complete analysis: import, clean, explore, visualise, and model. Create a comprehensive R Markdown report that tells a story about the data.


Practical Assignment

Using the economics dataset, create a report that shows trends in unemployment, population, and GDP. Include visualisations, summaries, and a narrative. Knit to HTML.


Challenge Exercise

Find a complex dataset online. Create a dashboard using flexdashboard or Shiny that allows users to explore the data. Include filters and interactive elements.


Quiz Answers

Fill-in-the-Blank: 1. Data communication, 2. R Markdown, 3. narrative, 4. Code chunks, 5. Knit, 6. dashboard, 7. Visualisations, 8. audience, 9. title, 10. Slides.

True/False: 1F, 2T, 3F, 4T, 5T, 6F, 7F, 8T, 9F, 10T.

Multiple Choice: 1C, 2B, 3B, 4B, 5B, 6B, 7B, 8B, 9B, 10B, 11A, 12A, 13B, 14B, 15B.


Key Takeaways

  • Data communication is essential for impact.
  • R Markdown combines code and text.
  • Use a clear structure in reports.
  • Visualisations should be simple and clear.
  • A narrative makes data meaningful.
  • Tailor your message to your audience.

Preparation for the Next Module

In the next module, we will bring everything together in a capstone project. You will apply all the skills you have learned to a real-world data analysis project. Start thinking about a dataset and a question you want to answer.


πŸ† Get Certified

πŸ”’

Earn this certificate

Every lesson is already free to read. Sign up, pass the exam, and unlock Practice Tools plus a verified certificate with your name on it β€” ₦4,000/month.

πŸŽ“ Sign Up & Unlock for ₦4,000/month
πŸ› οΈ Practice Tools
Hands-on simulators & labs - subscription required.
β†’
🎯 Internship Tasks
Real-world tasks to build your portfolio - try them free for 7 days, no card required.
β†’