Understand the R ecosystem β language, RStudio, packages, and the data analysis workflow.
Import and clean data β readr, tidyr, and the tidyverse.
Master data visualization β ggplot2, themes, and advanced graphics.
Advanced data wrangling β dplyr, tidyr, and the tidyverse.
Perform statistical analysis β descriptive statistics, hypothesis tests, and confidence intervals.
Build predictive models β regression, classification, and model evaluation.
Share your analysis β R Markdown, Shiny, and interactive dashboards.
Prepare for certification β R certifications, portfolio building, and career paths.
Welcome, young data explorer! Let us begin your journey into the world of R.
Welcome to the Introduction to R module! Have you ever wondered how businesses turn piles of numbers into beautiful charts and insights? The answer is R.
R is a powerful tool that helps people understand their data. It is like a super-smart assistant that takes your data and turns it into colorful charts, graphs, and dashboards that tell a story.
Imagine you have a big box of Lego bricks. You can build anything you want with them. R is like that β it takes your data (the Lego bricks) and helps you build amazing visual stories.
In this module, you will learn what R is, why it is so important, and how to get started. You will discover the R ecosystem, the interface, and how to create your first report. By the end, you will understand the basics of R and be ready to explore more.
π‘ Think about it: Have you ever seen a chart or a graph that made you understand something better? That is what R does β it turns numbers into pictures that make sense.
By the time you finish this module, you will be able to:
Chidi is a 12-year-old boy living in Enugu. He loves to collect data β he counts how many cars pass by his house, how many mangoes fall from the tree, and how many goals his football team scores. He writes everything down in his notebook.
One day, his uncle visited him. His uncle works in a big company in Lagos. Chidi showed his uncle all the numbers in his notebook. His uncle smiled and said, "Chidi, these numbers are amazing! But they are hard to understand. What if you could turn them into beautiful pictures?"
His uncle opened his laptop and showed Chidi a tool called R. He took Chidi's numbers and turned them into colorful charts and graphs. Chidi could see at a glance how many cars passed on different days, which month had the most mangoes, and when his team scored the most goals.
Chidi was fascinated. He said, "Uncle, this is like magic! Numbers become pictures!" His uncle replied, "It is not magic β it is R. And you can learn it too."
π§ Think about it: Have you ever collected data? What did you do with it? Could you turn it into a picture?
Definition: R is a programming language used for data analysis, statistics, and creating charts.
Why it is important: R helps people make better decisions by understanding their data. It is like a magnifying glass for numbers!
Simple explanation: Think of R as a magical paintbrush. You give it data, and it paints you a picture that tells a story.
π« School example: A teacher can use R to see which students are improving and which need extra help.
π Home example: Your parents could use R to see how much they spend on groceries each month.
π³π¬ Nigerian example: A Nigerian business owner can use R to see which products are selling best.
Illustration:
+-------------------+
| YOUR DATA | β Numbers, sales, records
+-------------------+
|
V
+-------------------+
| R | β The magic tool
+-------------------+
|
V
+-------------------+
| BEAUTIFUL CHARTS | β Pictures that tell a story
+-------------------+
π Mini summary: R is a tool that turns data into pictures to help you understand it better.
Definition: R is important because it helps people make sense of data. Data is everywhere, but understanding it can be hard. R makes it easy.
Why it is important: Without tools like R, businesses would have to look at thousands of numbers in spreadsheets. That takes a long time and is hard to understand.
Simple explanation: Imagine reading a book with no pictures. It is still a good book, but pictures make it more interesting and easier to understand. R is like the pictures for your data.
π« School example: A school principal can use R to see attendance patterns and plan better.
π Home example: Your family can use R to track savings and spending.
π³π¬ Nigerian example: A Nigerian bank can use R to see which branches are performing best.
π Mini summary: R is important because it helps people understand data quickly and easily.
Definition: R is the programming language. RStudio is the tool we use to write R code.
Why it is important: You need both to start using R.
Simple explanation: Think of R like the engine of a car and RStudio like the steering wheel and dashboard.
π« School example: A teacher installs R and RStudio to start analyzing student grades.
π Home example: Your parents install R and RStudio to track their expenses.
π³π¬ Nigerian example: A business analyst installs R and RStudio to analyze sales data.
π Mini summary: R is the language; RStudio is the tool you use to write R code.
Definition: RStudio is the tool you use to write and run R code.
Why it is important: Knowing your way around RStudio makes it easier to work with R.
Simple explanation: Think of RStudio like the cockpit of a plane. Each window has a job.
π« School example: A teacher uses the Editor to write code and the Console to see results.
π Home example: Your parents use the Environment to see their data.
π³π¬ Nigerian example: A business analyst uses the Plots pane to see charts.
+---------------------------------------------------+
| EDITOR | CONSOLE |
| (Write code) | (Run code) |
+---------------------------------------------------+
| ENVIRONMENT | PLOTS |
| (View data) | (View charts) |
+---------------------------------------------------+
π Mini summary: RStudio has different panes for writing code, viewing data, and seeing results.
Definition: Syntax is the set of rules for writing R code.
Why it is important: You need to learn the rules to write R code correctly.
Simple explanation: Think of syntax like the grammar of a language.
π« School example: A teacher uses <- to store student scores.
π Home example: Your parents use <- to store expense amounts.
π³π¬ Nigerian example: A business uses <- to store sales data.
π Mini summary: R syntax is the set of rules for writing R code.
Definition: Importing data means loading data into R.
Why it is important: You cannot analyze data without loading it first.
Simple explanation: Think of importing data like opening a book before you read it.
π« School example: A teacher imports student grades from a CSV file.
π Home example: Your parents import expense data from an Excel file.
π³π¬ Nigerian example: A business imports sales data from a CSV file.
π Mini summary: Importing data is the first step in analysis.
Definition: A vector is a collection of values of the same type.
Why it is important: Vectors are the building blocks of data in R.
Simple explanation: Think of a vector like a list of numbers.
π« School example: A teacher creates a vector of student grades.
π Home example: Your parents create a vector of expense amounts.
π³π¬ Nigerian example: A business creates a vector of sales figures.
π Mini summary: Vectors are collections of values.
Definition: Basic operations include addition, subtraction, multiplication, and division.
Why it is important: You need to perform basic operations to analyze data.
Simple explanation: Think of operations like doing math with your data.
π« School example: A teacher calculates the sum and mean of student grades.
π Home example: Your parents calculate the sum of expenses.
π³π¬ Nigerian example: A business calculates the mean of sales figures.
π Mini summary: Basic operations help you analyze data.
Definition: R is used by many Nigerian businesses and organizations to understand their data.
Why it is important: Seeing local examples helps you understand how R is used in your country.
Simple explanation: Think of it like seeing your favourite food at a local restaurant. It makes you feel connected.
π³π¬ Nigerian example: A Lagos-based supermarket uses R to see which products are selling best.
π Mini summary: R is helping Nigerian businesses grow and make better decisions.
Definition: Getting help means finding information about functions and packages.
Why it is important: You will need help as you learn R.
Simple explanation: Think of getting help like asking a teacher for help.
π« School example: A teacher uses ?mean to understand the mean function.
π Home example: Your parents use online resources to learn R.
π³π¬ Nigerian example: A business analyst uses Stack Overflow to solve problems.
π Mini summary: Getting help is an important part of learning R.
| Word | Simple Meaning |
|---|---|
| R | A tool that turns data into pictures. |
| RStudio | The tool used to write R code. |
| Data | Information, like numbers and words. |
| Visualization | A chart or graph that shows data in a picture. |
| Package | A collection of R functions. |
| Function | A command that performs a task. |
| Vector | A collection of values. |
| Data Frame | A table of data. |
| Variable | A container for a value. |
| Import | Load data into R. |
| Console | Where you type commands. |
| Editor | Where you write scripts. |
| Environment | Shows your data and variables. |
| Plot | A chart or graph. |
| Syntax | The rules for writing code. |
+-------------------+
| YOUR DATA | β Numbers, sales, records
+-------------------+
|
V
+-------------------+
| R | β The magic tool
+-------------------+
|
V
+-------------------+
| BEAUTIFUL CHARTS | β Pictures that tell a story
+-------------------+
+---------------------------------------------------+
| EDITOR | CONSOLE |
| (Write code) | (Run code) |
+---------------------------------------------------+
| ENVIRONMENT | PLOTS |
| (View data) | (View charts) |
+---------------------------------------------------+
+-------------------+
| numbers <- c(1,2,3,4,5) β Create a vector
+-------------------+
|
V
+-------------------+
| sum(numbers) β Add them up
+-------------------+
|
V
+-------------------+
| mean(numbers) β Find the average
+-------------------+
| R | Excel |
|---|---|
| Free and open-source | Paid |
| Handles large data | Limited to small data |
| Powerful visualizations | Basic charts |
| Reproducible | Not reproducible |
| Programming language | Spreadsheet tool |
Lesson 1: R is a tool that turns data into pictures.
Lesson 2: R is important because it helps people understand data.
Lesson 3: R is the language; RStudio is the tool.
Lesson 4: RStudio has different panes for writing code and viewing data.
Lesson 5: R syntax is the set of rules for writing R code.
Lesson 6: Importing data is the first step in analysis.
Lesson 7: Vectors are collections of values.
Lesson 8: Basic operations help you analyze data.
Lesson 9: R is used by Nigerian businesses.
Lesson 10: Getting help is an important part of learning R.
In this module, you learned about R β a tool that turns data into beautiful pictures. You discovered the R ecosystem, RStudio, and how to create your first report. You also learned about the importance of R and how it is used in Nigeria.
π― You can now:
Match the word on the left with the correct meaning on the right:
| Word | Meaning |
|---|---|
| R | A tool for writing R code |
| RStudio | A collection of values |
| Vector | A tool that turns data into pictures |
| Package | A collection of functions |
Answers: R β A tool that turns data into pictures; RStudio β A tool for writing R code; Vector β A collection of values; Package β A collection of functions.
Scenario 1: Chidi has collected data on how many cars pass his house each day. He wants to turn this data into a chart so he can see patterns. What should Chidi do? How can R help him?
Scenario 2: A Lagos supermarket wants to see which products are selling best. They have a lot of sales data in Excel. How can R help them?
In groups of 4β5, discuss a type of data you could collect (e.g., sales, attendance, expenses). Think about how you could use R to visualize this data. Present your ideas to the class.
Think about data you collect in your daily life (e.g., how much time you spend on homework, how many books you read). Write a short paragraph (about 100 words) about how you could use R to visualize this data.
Create a Simple R Report
Find a simple dataset (e.g., a list of sales or expenses). Use R to create a report with at least one chart. Save and share your report with the class.
Install R and RStudio on your computer. Import a simple CSV file and create a bar chart. Write a short report (about 150 words) about what you did and what you learned.
The Challenge: Imagine you are a business analyst in Lagos. You have sales data for four regions: Lagos, Abuja, Kano, and Port Harcourt. Create an R script that imports the data, creates a bar chart, and calculates the average sales per region.
Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A
Fill-in-the-Blank: 1. data, 2. R, 3. vector, 4. Importing, 5. Nigerian
True or False: 1. False, 2. True, 3. False, 4. True, 5. True
Matching: R β A tool that turns data into pictures; RStudio β A tool for writing R code; Vector β A collection of values; Package β A collection of functions.
In the next module, we will explore Data Import and Cleaning. You will learn how to clean and prepare your data for analysis.
π Congratulations! You have completed Module One of the R for Data Analysis course.
π You are now ready to move to Module Two: Data Import and Cleaning.
Welcome back, young data cleaner! Today we learn how to import and clean our data.
In Module One, you learned what R is and how to use RStudio. Now you are ready to work with real data. But real data is often messy. It has mistakes, missing values, and inconsistent formats.
Imagine you have a box of toys. Some toys are broken, some are missing pieces, and some are dirty. Before you can play with them, you need to clean them. Data is the same. You need to clean it before you can analyze it.
In this module, you will learn how to import data into R and how to clean it. You will learn about missing values, duplicates, and inconsistent formats. By the end, you will be able to turn messy data into clean, analysis-ready data.
π‘ Think about it: Have you ever had to clean a messy room? Data cleaning is like that β you organize and tidy up your data.
By the time you finish this module, you will be able to:
Nneka is a 12-year-old girl who loves to collect data. She asked her classmates about their favourite foods and wrote down their answers. But her notebook was messy. Some names were spelled wrong. Some answers were missing. Some were written in different ways.
Nneka tried to make a chart, but it was hard. She showed her notebook to her uncle, who works with data. Her uncle said, "Nneka, your data is messy. You need to clean it before you can make a chart."
He opened R and showed her how to import the data and clean it. Together, they fixed the misspelled names, filled in missing answers, and made everything consistent. Then they created a beautiful chart showing the favourite foods of the class.
Nneka learned that cleaning data is just as important as making charts. Without clean data, your charts can be wrong.
π§ Think about it: Have you ever had data that was messy or hard to understand? What did you do?
Definition: Data cleaning is the process of fixing errors and making data consistent.
Why it is important: Clean data leads to accurate analysis. Dirty data leads to wrong conclusions.
Simple explanation: Think of data cleaning like washing your hands before eating. You need to remove the dirt so you don't get sick.
π« School example: A teacher cleans student data by fixing misspelled names.
π Home example: Your parents clean expense data by removing duplicate entries.
π³π¬ Nigerian example: A business cleans sales data by fixing inconsistent region names.
+-------------------+
| MESSY DATA | β Data with errors, blanks, and mistakes
+-------------------+
|
V
+-------------------+
| DATA CLEANING | β The cleaning process
+-------------------+
|
V
+-------------------+
| CLEAN DATA | β Ready for analysis
+-------------------+
π Mini summary: Data cleaning is the process of fixing errors and making data consistent.
Definition: Importing data means loading data from a file into R.
Why it is important: You need to import data before you can analyze it.
Simple explanation: Think of importing data like opening a book before you read it.
π« School example: A teacher imports student grades from a CSV file.
π Home example: Your parents import expense data from an Excel file.
π³π¬ Nigerian example: A business imports sales data from a CSV file.
π Mini summary: Importing data is the first step in analysis.
Definition: Exploring your data means looking at it to understand what you have.
Why it is important: You need to know what your data looks like before you can clean it.
Simple explanation: Think of exploring your data like looking at a map before a journey.
π« School example: A teacher uses head() to see the first few student grades.
π Home example: Your parents use summary() to see expense statistics.
π³π¬ Nigerian example: A business uses head() to see the first few sales records.
π Mini summary: Exploring your data helps you understand what you have.
Definition: Missing values are blank cells in your data.
Why it is important: Missing values can cause errors in your analysis.
Simple explanation: Think of missing values like missing pieces in a puzzle.
π« School example: A teacher replaces missing grades with 0.
π Home example: Your parents remove rows with missing expense amounts.
π³π¬ Nigerian example: A business replaces missing sales with 0.
π Mini summary: Handling missing values prevents errors in your analysis.
Definition: Duplicates are rows that appear more than once.
Why it is important: Duplicates can make your analysis inaccurate.
Simple explanation: Think of duplicates like copying a page in a book.
π« School example: A teacher removes duplicate student entries.
π Home example: Your parents remove duplicate expense entries.
π³π¬ Nigerian example: A business removes duplicate customer records.
π Mini summary: Removing duplicates keeps your data accurate.
Definition: Inconsistent formats are when data is written in different ways.
Why it is important: Inconsistent formats make it hard to analyze data.
Simple explanation: Think of inconsistent formats like different spellings of the same word.
π« School example: A teacher fixes inconsistent student names.
π Home example: Your parents fix inconsistent category names.
π³π¬ Nigerian example: A business fixes inconsistent region names.
π Mini summary: Fixing inconsistent formats makes your data consistent.
Definition: Renaming columns means changing the names of columns.
Why it is important: Clear column names make your data easier to understand.
Simple explanation: Think of renaming columns like labeling boxes in a storage room.
π« School example: A teacher renames "Grade" to "Score".
π Home example: Your parents rename "Amount" to "Expense".
π³π¬ Nigerian example: A business renames "Sales" to "Revenue".
π Mini summary: Renaming columns makes your data easier to understand.
Definition: Filtering means keeping only the rows that meet certain conditions.
Why it is important: Filtering helps you focus on the data that matters.
Simple explanation: Think of filtering like using a sieve to separate what you want.
π« School example: A teacher filters students by grade.
π Home example: Your parents filter expenses by category.
π³π¬ Nigerian example: A business filters sales by region.
π Mini summary: Filtering helps you focus on the data that matters.
Definition: Selecting means keeping only the columns you need.
Why it is important: Selecting helps you focus on the columns that matter.
Simple explanation: Think of selecting like choosing the right tools for a job.
π« School example: A teacher selects only the grade column.
π Home example: Your parents select only the expense column.
π³π¬ Nigerian example: A business selects only the sales column.
π Mini summary: Selecting helps you focus on the columns that matter.
Definition: Nigerian businesses use data cleaning to prepare their data for analysis.
Why it is important: Data cleaning helps Nigerian businesses make better decisions.
Simple explanation: Think of it like cleaning a shop before customers come in.
π³π¬ Nigerian example: A Lagos supermarket cleans sales data by fixing inconsistent product names.
π Mini summary: Nigerian businesses use data cleaning to make better decisions.
| Word | Simple Meaning |
|---|---|
| Data Cleaning | Fixing errors and making data consistent. |
| Import | Loading data into R. |
| Missing Value | A blank cell in your data. |
| Duplicate | A row that appears more than once. |
| Inconsistent Format | Data written in different ways. |
| Filter | Keeping only certain rows. |
| Select | Keeping only certain columns. |
| NA | A missing value in R. |
| head() | Shows the first few rows. |
| summary() | Shows a summary of the data. |
| str() | Shows the structure of the data. |
| View() | Opens the data in a spreadsheet view. |
| mutate() | Creates a new column. |
| rename() | Changes column names. |
| distinct() | Removes duplicate rows. |
+-------------------+
| MESSY DATA | β Data with errors, blanks, and mistakes
+-------------------+
|
V
+-------------------+
| IMPORT DATA | β Load data into R
+-------------------+
|
V
+-------------------+
| EXPLORE DATA | β Look at the data
+-------------------+
|
V
+-------------------+
| HANDLE MISSING | β Fix missing values
+-------------------+
|
V
+-------------------+
| REMOVE DUPLICATES | β Remove duplicates
+-------------------+
|
V
+-------------------+
| FIX INCONSISTENT | β Fix inconsistent formats
+-------------------+
|
V
+-------------------+
| CLEAN DATA | β Ready for analysis
+-------------------+
+-------------------+
| ALL DATA | β All rows
+-------------------+
|
V
+-------------------+
| FILTER | β Keep only rows that meet conditions
+-------------------+
|
V
+-------------------+
| FILTERED DATA | β Only the rows you want
+-------------------+
+-------------------+
| ALL COLUMNS | β All columns
+-------------------+
|
V
+-------------------+
| SELECT | β Keep only certain columns
+-------------------+
|
V
+-------------------+
| SELECTED DATA | β Only the columns you want
+-------------------+
| Dirty Data | Clean Data |
|---|---|
| Has missing values | No missing values |
| Has duplicates | No duplicates |
| Inconsistent formats | Consistent formats |
| Unclear column names | Clear column names |
| Hard to analyze | Easy to analyze |
Lesson 1: Data cleaning fixes errors and makes data consistent.
Lesson 2: Importing data is the first step in analysis.
Lesson 3: Exploring your data helps you understand what you have.
Lesson 4: Handling missing values prevents errors.
Lesson 5: Removing duplicates keeps your data accurate.
Lesson 6: Fixing inconsistent formats makes your data consistent.
Lesson 7: Renaming columns makes your data easier to understand.
Lesson 8: Filtering helps you focus on the data that matters.
Lesson 9: Selecting helps you focus on the columns that matter.
Lesson 10: Nigerian businesses use data cleaning to make better decisions.
In this module, you learned about data import and cleaning. You discovered how to import data from different sources and explore it. You learned how to handle missing values, remove duplicates, and fix inconsistent formats. You also learned how to rename columns, filter data, and select columns.
π― You can now:
Match the word on the left with the correct meaning on the right:
| Word | Meaning |
|---|---|
| read.csv() | Removes duplicates |
| distinct() | Shows the first few rows |
| head() | Imports a CSV file |
| filter() | Keeps only certain rows |
Answers: read.csv() β Imports a CSV file; distinct() β Removes duplicates; head() β Shows the first few rows; filter() β Keeps only certain rows.
Scenario 1: Nneka has a list of students with their grades. Some students are listed twice, and some grades are missing. What should Nneka do? How can R help her?
Scenario 2: A Lagos supermarket has sales data in different formats. Some product names are written in uppercase, some in lowercase. Some have extra spaces. How can R help clean this data?
In groups of 4β5, discuss a type of messy data you have encountered. Think about how you could use R to clean it. Present your ideas to the class.
Think about a dataset you could clean. Write a short paragraph (about 100 words) about how you would use R to clean it.
Clean a Dataset
Find a messy dataset (e.g., a list of sales with duplicates and missing values). Use R to clean it. Document the steps you took and the final result.
Download a sample dataset with errors. Use R to clean it by handling missing values, removing duplicates, and fixing inconsistent formats. Write a short report (about 150 words) about what you did and what you learned.
The Challenge: Imagine you are a business analyst in Lagos. You have sales data from four different stores, each in a different format. Use R to clean and combine the data. Create a single dataset with consistent formatting.
Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-D, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A
Fill-in-the-Blank: 1. cleaning, 2. Importing, 3. missing, 4. distinct(), 5. Filtering
True or False: 1. False, 2. True, 3. True, 4. False, 5. True
Matching: read.csv() β Imports a CSV file; distinct() β Removes duplicates; head() β Shows the first few rows; filter() β Keeps only certain rows.
In the next module, we will explore Data Visualization with ggplot2. You will learn how to create beautiful charts and graphs to visualize your data.
π Congratulations! You have completed Module Two of the R for Data Analysis course.
π You are now ready to move to Module Three: Data Visualization with ggplot2.
Welcome back, young artist! Today we learn how to turn data into beautiful pictures.
In Module Two, you learned how to clean your data. Now you have clean, ready-to-use data. But data is hard to understand when it is just numbers. You need pictures to tell the story.
Imagine you have a story to tell. You could just read it out loud, but it would be much more interesting with pictures. Data visualizations are the pictures for your data.
ggplot2 is a package in R that helps you create beautiful charts and graphs. It is like a paintbrush for your data.
In this module, you will learn about the grammar of graphics, how to create bar charts, scatter plots, line charts, and more. You will also learn how to customize your charts to make them look professional.
π‘ Think about it: Have you ever seen a chart that told you a story? That is what data visualization does.
By the time you finish this module, you will be able to:
Ada is a 12-year-old girl who loves art. She had a lot of data about her class's favourite foods. She wanted to show it to her teacher, but numbers were boring.
Her uncle, who works with data, said, "Ada, you need to visualize your data. Turn it into a picture. That way, everyone can understand it."
Ada learned how to use ggplot2 in R. She created a beautiful bar chart showing the favourite foods of her class. Her teacher was amazed. The chart told a story that numbers alone could not.
π§ Think about it: Have you ever turned numbers into a picture? That is what data visualization is all about.
Definition: Data visualization is the process of turning data into pictures like charts and graphs.
Why it is important: Pictures make it easier to understand data. A picture is worth a thousand numbers.
Simple explanation: Think of data visualization like drawing a picture of your data.
π« School example: A teacher uses a bar chart to show student grades.
π Home example: Your parents use a pie chart to show expenses by category.
π³π¬ Nigerian example: A business uses a line chart to show sales over time.
+-------------------+
| DATA | β Numbers
+-------------------+
|
V
+-------------------+
| VISUALIZATION | β Pictures
+-------------------+
|
V
+-------------------+
| UNDERSTANDING | β Insights
+-------------------+
π Mini summary: Data visualization turns numbers into pictures.
Definition: ggplot2 is a package in R that helps you create beautiful charts.
Why it is important: ggplot2 is the most popular package for data visualization in R.
Simple explanation: Think of ggplot2 like a paintbrush for your data.
π« School example: A teacher uses ggplot2 to create a bar chart of grades.
π Home example: Your parents use ggplot2 to create a pie chart of expenses.
π³π¬ Nigerian example: A business uses ggplot2 to create a line chart of sales.
π Mini summary: ggplot2 is a package for creating charts in R.
Definition: Installing means downloading the package. Loading means making it available to use.
Why it is important: You need to install and load ggplot2 before you can use it.
Simple explanation: Think of installing like buying a new tool. Loading is like taking it out of the toolbox.
π« School example: A teacher installs ggplot2 on the school computer.
π Home example: Your parents install ggplot2 on their computer.
π³π¬ Nigerian example: A business installs ggplot2 on the company computer.
π Mini summary: You need to install and load ggplot2 before using it.
Definition: The grammar of graphics is a system for building charts layer by layer.
Why it is important: It makes it easy to build complex charts.
Simple explanation: Think of it like building a house β you start with the foundation and add layers.
π« School example: A teacher uses facets to show grades by class.
π Home example: Your parents use themes to make charts look professional.
π³π¬ Nigerian example: A business uses facets to show sales by region.
π Mini summary: The grammar of graphics builds charts layer by layer.
Definition: A bar chart uses bars to compare categories.
Why it is important: Bar charts are the most common type of chart.
Simple explanation: Think of a bar chart like blocks of different heights.
π« School example: A teacher creates a bar chart of student grades.
π Home example: Your parents create a bar chart of expenses by category.
π³π¬ Nigerian example: A business creates a bar chart of sales by region.
+-------------------+
| SALES BY REGION |
+-------------------+
| Lagos ββββββββ
| Abuja ββββββ
| Kano ββββ
| Port H ββββββ
+-------------------+
π Mini summary: Bar charts compare categories.
Definition: A scatter plot uses points to show relationships between two variables.
Why it is important: Scatter plots help you see patterns and correlations.
Simple explanation: Think of a scatter plot like dots on a map.
π« School example: A teacher creates a scatter plot of study time vs grades.
π Home example: Your parents create a scatter plot of income vs expenses.
π³π¬ Nigerian example: A business creates a scatter plot of advertising spend vs sales.
+-------------------+
| PROFIT VS SALES |
+-------------------+
| * * |
| * * * |
| * * * |
| * * * |
| * * * |
+-------------------+
π Mini summary: Scatter plots show relationships between variables.
Definition: A line chart uses lines to show trends over time.
Why it is important: Line charts are the best way to show how things change over time.
Simple explanation: Think of a line chart like connecting the dots.
π« School example: A teacher creates a line chart of grades over time.
π Home example: Your parents create a line chart of expenses over months.
π³π¬ Nigerian example: A business creates a line chart of sales over quarters.
+-------------------+
| SALES OVER TIME |
+-------------------+
| 200| *--*--* |
| 150| * * |
| 100|* * |
| 50| * |
| 0|___________ |
+-------------------+
π Mini summary: Line charts show trends over time.
Definition: Customizing means changing the look of your chart.
Why it is important: Customization makes your charts look professional.
Simple explanation: Think of customization like decorating your chart.
π« School example: A teacher adds a title to a chart.
π Home example: Your parents change the colours of a chart.
π³π¬ Nigerian example: A business uses a professional theme for a report.
π Mini summary: Customizing makes your charts look professional.
Definition: Faceting creates multiple charts side by side.
Why it is important: Faceting helps you compare different groups.
Simple explanation: Think of faceting like multiple windows into your data.
π« School example: A teacher creates charts for each class.
π Home example: Your parents create charts for each month.
π³π¬ Nigerian example: A business creates charts for each region.
π Mini summary: Faceting creates multiple charts.
Definition: Nigerian businesses use data visualization to make decisions.
Why it is important: Visualizations help Nigerian businesses understand their data.
Simple explanation: Think of it like seeing a map before a journey.
π³π¬ Nigerian example: A Lagos supermarket uses a bar chart to show sales by product.
π Mini summary: Nigerian businesses use data visualization to make decisions.
| Word | Simple Meaning |
|---|---|
| Visualization | A picture of your data. |
| ggplot2 | A package for creating charts. |
| Bar Chart | A chart with bars. |
| Scatter Plot | A chart with points. |
| Line Chart | A chart with lines. |
| Grammar of Graphics | A system for building charts. |
| Aesthetics | What you map to axes. |
| Geometries | The type of chart. |
| Facets | Multiple charts. |
| Themes | Styling and appearance. |
| geom_bar() | Creates a bar chart. |
| geom_point() | Creates a scatter plot. |
| geom_line() | Creates a line chart. |
| labs() | Adds titles and labels. |
| facet_wrap() | Creates multiple charts. |
+-------------------+
| SALES BY REGION |
+-------------------+
| Lagos ββββββββ
| Abuja ββββββ
| Kano ββββ
| Port H ββββββ
+-------------------+
+-------------------+
| PROFIT VS SALES |
+-------------------+
| * * |
| * * * |
| * * * |
| * * * |
| * * * |
+-------------------+
+-------------------+
| SALES OVER TIME |
+-------------------+
| 200| *--*--* |
| 150| * * |
| 100|* * |
| 50| * |
| 0|___________ |
+-------------------+
| Chart Type | Best For | Example |
|---|---|---|
| Bar Chart | Comparing categories | Sales by region |
| Scatter Plot | Showing relationships | Sales vs profit |
| Line Chart | Showing trends over time | Sales over time |
| Pie Chart | Showing proportions | Market share |
Lesson 1: Data visualization turns numbers into pictures.
Lesson 2: ggplot2 is a package for creating charts.
Lesson 3: You need to install and load ggplot2 before using it.
Lesson 4: The grammar of graphics builds charts layer by layer.
Lesson 5: Bar charts compare categories.
Lesson 6: Scatter plots show relationships.
Lesson 7: Line charts show trends over time.
Lesson 8: Customizing makes your charts look professional.
Lesson 9: Faceting creates multiple charts.
Lesson 10: Nigerian businesses use data visualization.
In this module, you learned about data visualization with ggplot2. You discovered the grammar of graphics and how to create bar charts, scatter plots, and line charts. You also learned how to customize your charts and use faceting to create multiple charts. You saw examples from Nigeria and learned best practices for data visualization.
π― You can now:
Match the word on the left with the correct meaning on the right:
| Word | Meaning |
|---|---|
| Bar Chart | Shows trends over time |
| Scatter Plot | Compares categories |
| Line Chart | Shows relationships |
| Faceting | Creates multiple charts |
Answers: Bar Chart β Compares categories; Scatter Plot β Shows relationships; Line Chart β Shows trends over time; Faceting β Creates multiple charts.
Scenario 1: Ada has data on her class's favourite foods. She wants to show it to her teacher. What chart should she use? Why?
Scenario 2: A Lagos supermarket wants to see how sales have changed over the last year. What chart should they use? Why?
In groups of 4β5, discuss how you would visualize sales data for a Nigerian business. What charts would you create? Present your ideas to the class.
Think about a dataset you would like to visualize. Write a short paragraph (about 100 words) about the chart you would create.
Create a Visualization
Find a dataset and create a visualization using ggplot2. Include a title, labels, and a theme. Save and share your visualization with the class.
Use ggplot2 to create a bar chart and a scatter plot. Customize both charts with titles, colours, and themes. Write a short report (about 150 words) about what you did and what you learned.
The Challenge: Imagine you are a data analyst in Lagos. You have sales data for four regions. Create a bar chart, a scatter plot, and a line chart. Customize all charts and use faceting.
Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A
Fill-in-the-Blank: 1. visualization, 2. ggplot2, 3. bar, 4. scatter, 5. line
True or False: 1. False, 2. True, 3. False, 4. True, 5. True
Matching: Bar Chart β Compares categories; Scatter Plot β Shows relationships; Line Chart β Shows trends over time; Faceting β Creates multiple charts.
In the next module, we will explore Data Manipulation with dplyr. You will learn how to filter, select, mutate, and summarise data.
π Congratulations! You have completed Module Three of the R for Data Analysis course.
π You are now ready to move to Module Four: Data Manipulation with dplyr.
Welcome back, young data wrangler! Today we learn how to manipulate and transform our data.
In Module Three, you learned how to create beautiful charts with ggplot2. But sometimes your data is not in the right shape for analysis. You need to manipulate it β filter, select, mutate, and summarise.
Imagine you have a big box of Lego bricks. You need to find the right pieces, sort them, and build something new. Data manipulation is like that β you take your data and shape it into what you need.
dplyr is a package in R that makes data manipulation easy. It is like a toolbox for your data.
In this module, you will learn about the dplyr verbs: filter(), select(), mutate(), summarise(), arrange(), and group_by(). You will also learn how to join datasets together.
π‘ Think about it: Have you ever had to sort through a big pile of things to find what you need? That is what data manipulation is all about.
By the time you finish this module, you will be able to:
Chidi is a 12-year-old boy who loves Lego. He has a big box of Lego bricks. He wanted to build a spaceship, but he needed to find the right pieces. He had to sort through all the bricks to find the ones he needed.
His uncle, who works with data, said, "Chidi, sorting through Lego is like sorting through data. You need to filter out the pieces you don't need, select the ones you want, and mutate them into something new."
Chidi learned how to use dplyr in R. He could filter, select, and mutate his data just like he sorted his Lego bricks.
π§ Think about it: Have you ever had to sort through a big pile of things? That is what data manipulation is all about.
Definition: dplyr is a package in R that helps you manipulate data.
Why it is important: dplyr makes data manipulation easy and fast.
Simple explanation: Think of dplyr like a toolbox for your data. It has tools to filter, select, mutate, and summarise.
π« School example: A teacher uses dplyr to filter student grades.
π Home example: Your parents use dplyr to summarise expenses.
π³π¬ Nigerian example: A business uses dplyr to analyse sales data.
+-------------------+
| dplyr TOOLBOX |
+-------------------+
| filter() | β Select rows
| select() | β Select columns
| mutate() | β Create new columns
| summarise() | β Summarise data
| arrange() | β Sort data
| group_by() | β Group data
+-------------------+
π Mini summary: dplyr is a toolbox for manipulating data.
Definition: Installing means downloading the package. Loading means making it available to use.
Why it is important: You need to install and load dplyr before you can use it.
Simple explanation: Think of installing like buying a new tool. Loading is like taking it out of the toolbox.
π« School example: A teacher installs dplyr on the school computer.
π Home example: Your parents install dplyr on their computer.
π³π¬ Nigerian example: A business installs dplyr on the company computer.
π Mini summary: You need to install and load dplyr before using it.
Definition: The pipe operator (%>%) passes data from one function to the next.
Why it is important: The pipe makes your code easier to read and write.
Simple explanation: Think of the pipe like a conveyor belt that moves data from one step to the next.
π« School example: A teacher uses the pipe to chain operations.
π Home example: Your parents use the pipe to chain operations.
π³π¬ Nigerian example: A business uses the pipe to chain operations.
π Mini summary: The pipe passes data from one function to the next.
Definition: filter() selects rows that meet certain conditions.
Why it is important: filter() helps you focus on the data that matters.
Simple explanation: Think of filter() like using a sieve to separate what you want.
π« School example: A teacher filters students by grade.
π Home example: Your parents filter expenses by category.
π³π¬ Nigerian example: A business filters sales by region.
π Mini summary: filter() selects rows based on conditions.
Definition: select() chooses specific columns.
Why it is important: select() helps you focus on the columns that matter.
Simple explanation: Think of select() like choosing the right tools for a job.
π« School example: A teacher selects only the grade column.
π Home example: Your parents select only the expense column.
π³π¬ Nigerian example: A business selects only the sales column.
π Mini summary: select() chooses specific columns.
Definition: mutate() creates new columns from existing columns.
Why it is important: mutate() helps you create new variables.
Simple explanation: Think of mutate() like adding new pieces to your Lego creation.
π« School example: A teacher creates a new column for letter grades.
π Home example: Your parents create a new column for expense categories.
π³π¬ Nigerian example: A business creates a new column for profit.
π Mini summary: mutate() creates new columns.
Definition: summarise() creates summary statistics for your data.
Why it is important: summarise() helps you understand your data.
Simple explanation: Think of summarise() like getting a summary of a book.
π« School example: A teacher summarises grades.
π Home example: Your parents summarise expenses.
π³π¬ Nigerian example: A business summarises sales.
π Mini summary: summarise() creates summary statistics.
Definition: arrange() sorts your data.
Why it is important: arrange() helps you order your data.
Simple explanation: Think of arrange() like organising your books on a shelf.
π« School example: A teacher sorts students by grade.
π Home example: Your parents sort expenses by amount.
π³π¬ Nigerian example: A business sorts sales by region.
π Mini summary: arrange() sorts your data.
Definition: group_by() groups your data by one or more variables.
Why it is important: group_by() allows you to perform operations on groups.
Simple explanation: Think of group_by() like sorting your Lego bricks by colour.
π« School example: A teacher groups students by class.
π Home example: Your parents group expenses by category.
π³π¬ Nigerian example: A business groups sales by region.
π Mini summary: group_by() groups your data.
Definition: Joining combines two datasets based on a common column.
Why it is important: Joining allows you to combine data from different sources.
Simple explanation: Think of joining like connecting two Lego pieces.
π« School example: A teacher joins student data with grade data.
π Home example: Your parents join expense data with category data.
π³π¬ Nigerian example: A business joins sales data with customer data.
π Mini summary: Joining combines two datasets.
| Word | Simple Meaning |
|---|---|
| dplyr | A package for manipulating data. |
| filter() | Selects rows based on conditions. |
| select() | Chooses specific columns. |
| mutate() | Creates new columns. |
| summarise() | Creates summary statistics. |
| arrange() | Sorts data. |
| group_by() | Groups data. |
| Pipe (%>%) | Passes data from one function to the next. |
| Join | Combines two datasets. |
| inner_join() | Keeps only matching rows. |
| left_join() | Keeps all rows from the left table. |
| right_join() | Keeps all rows from the right table. |
| full_join() | Keeps all rows from both tables. |
| summarise() | Creates summary statistics. |
| n() | Counts the number of rows. |
+-------------------+
| filter() | β Select rows
+-------------------+
| select() | β Select columns
+-------------------+
| mutate() | β Create new columns
+-------------------+
| summarise() | β Summarise data
+-------------------+
| arrange() | β Sort data
+-------------------+
| group_by() | β Group data
+-------------------+
+-------------------+
| data | β Start with data
+-------------------+
|
V
+-------------------+
| filter() | β Filter rows
+-------------------+
|
V
+-------------------+
| select() | β Select columns
+-------------------+
|
V
+-------------------+
| summarise() | β Summarise data
+-------------------+
+-------------------+ +-------------------+
| Table A | | Table B |
| ID | Name | | ID | Sales |
+-------------------+ +-------------------+
| |
+-------------+---------------+
|
V
+-------------------+
| Joined Table |
| ID | Name | Sales|
+-------------------+
| dplyr | Base R |
|---|---|
| filter() | data[data$column == condition, ] |
| select() | data[, c("col1", "col2")] |
| mutate() | data$new <- expression |
| summarise() | sum(data$column) |
| arrange() | data[order(data$column), ] |
| group_by() | aggregate(data, by = list(group), FUN = sum) |
Lesson 1: dplyr is a toolbox for manipulating data.
Lesson 2: You need to install and load dplyr before using it.
Lesson 3: The pipe (%>%) passes data from one function to the next.
Lesson 4: filter() selects rows based on conditions.
Lesson 5: select() chooses specific columns.
Lesson 6: mutate() creates new columns.
Lesson 7: summarise() creates summary statistics.
Lesson 8: arrange() sorts your data.
Lesson 9: group_by() groups your data.
Lesson 10: Joining combines two datasets.
In this module, you learned about data manipulation with dplyr. You discovered the dplyr verbs: filter(), select(), mutate(), summarise(), arrange(), and group_by(). You also learned how to use the pipe operator and how to join datasets together.
π― You can now:
Match the word on the left with the correct meaning on the right:
| Word | Meaning |
|---|---|
| filter() | Sorts data |
| select() | Selects rows |
| mutate() | Selects columns |
| arrange() | Creates new columns |
Answers: filter() β Selects rows; select() β Selects columns; mutate() β Creates new columns; arrange() β Sorts data.
Scenario 1: Chidi has a dataset of student grades. He wants to filter students who scored above 80 and create a new column for letter grades. How can dplyr help him?
Scenario 2: A Lagos supermarket has sales data for different products. They want to summarise total sales by product category. How can dplyr help them?
In groups of 4β5, discuss how you would manipulate sales data for a Nigerian business. What dplyr verbs would you use? Present your ideas to the class.
Think about a dataset you would like to manipulate. Write a short paragraph (about 100 words) about what dplyr verbs you would use.
Manipulate a Dataset
Find a dataset and use dplyr to manipulate it. Filter, select, mutate, summarise, and group the data. Save and share your results with the class.
Use dplyr to filter, select, mutate, summarise, and group a dataset. Write a short report (about 150 words) about what you did and what you learned.
The Challenge: Imagine you are a data analyst in Lagos. You have sales data for four regions. Use dplyr to filter each region, summarise total sales, and arrange by sales. Create a new column for profit.
Multiple Choice: 1-A, 2-A, 3-B, 4-C, 5-D, 6-D, 7-C, 8-A, 9-A, 10-D, 11-A, 12-A, 13-B, 14-C, 15-D
Fill-in-the-Blank: 1. manipulating, 2. filter(), 3. select(), 4. mutate(), 5. summarise()
True or False: 1. True, 2. False, 3. True, 4. True, 5. False
Matching: filter() β Selects rows; select() β Selects columns; mutate() β Creates new columns; arrange() β Sorts data.
In the next module, we will explore Statistical Analysis and Hypothesis Testing. You will learn how to perform statistical tests and analyse your data.
π Congratulations! You have completed Module Four of the R for Data Analysis course.
π You are now ready to move to Module Five: Statistical Analysis and Hypothesis Testing.
Welcome back, young data scientist! Today we learn how to understand and test our data.
In Module Four, you learned how to manipulate data. Now you have clean, organised data. But how do you understand it? How do you know if there is a real pattern or just random chance?
Imagine you are a detective. You have clues (your data). You need to figure out what they mean. Statistics is like your detective toolkit. It helps you understand your data and make decisions.
Hypothesis testing is a way to test if something is true. It is like a lie detector for your data.
In this module, you will learn about descriptive statistics, t-tests, ANOVA, correlation, and linear regression. By the end, you will be able to understand and test your data.
π‘ Think about it: Have you ever wondered if something was true or just a coincidence? Statistics can help you find out.
By the time you finish this module, you will be able to:
Ada is a 12-year-old girl who loves mysteries. She noticed that her class performed better on tests on Fridays than on Mondays. She wondered, "Is this just a coincidence, or is there a real difference?"
Her uncle, who works with data, said, "Ada, you need to use statistics. Statistics can help you find out if the difference is real or just random chance."
Ada learned about hypothesis testing. She tested her data and found that the difference was real. She was like a detective solving a mystery.
π§ Think about it: Have you ever wondered if something was true or just a coincidence? Statistics can help you find out.
Definition: Statistics is the science of collecting, analysing, and interpreting data.
Why it is important: Statistics helps you make sense of data and make decisions.
Simple explanation: Think of statistics like a magnifying glass for your data. It helps you see what is really there.
π« School example: A teacher uses statistics to understand student performance.
π Home example: Your parents use statistics to understand their spending.
π³π¬ Nigerian example: A business uses statistics to understand sales.
+-------------------+
| STATISTICS |
+-------------------+
| Descriptive | β Summarise data
| Inferential | β Make predictions
+-------------------+
π Mini summary: Statistics is the science of understanding data.
Definition: Descriptive statistics summarise your data.
Why it is important: They help you understand your data at a glance.
Simple explanation: Think of descriptive statistics like a summary of a book.
π« School example: A teacher calculates the mean grade.
π Home example: Your parents calculate the mean expense.
π³π¬ Nigerian example: A business calculates the mean sales.
π Mini summary: Descriptive statistics summarise your data.
Definition: A population is the entire group you are studying. A sample is a subset of the population.
Why it is important: You often can't study the whole population, so you use a sample.
Simple explanation: Think of a population like all the students in a school. A sample is like one class.
π« School example: All students in the school (population) vs. one class (sample).
π Home example: All family members (population) vs. one person (sample).
π³π¬ Nigerian example: All customers (population) vs. a survey group (sample).
π Mini summary: Population is the whole group; sample is a part of it.
Definition: Hypothesis testing is a way to test if something is true.
Why it is important: It helps you make decisions based on data.
Simple explanation: Think of hypothesis testing like a lie detector for your data.
π« School example: Testing if students who study more get better grades.
π Home example: Testing if eating breakfast improves test scores.
π³π¬ Nigerian example: Testing if a new marketing campaign increases sales.
π Mini summary: Hypothesis testing helps you make decisions.
Definition: A t-test compares the means of two groups.
Why it is important: It helps you see if two groups are different.
Simple explanation: Think of a t-test like comparing two things.
π« School example: Comparing grades of two classes.
π Home example: Comparing expenses before and after a budget.
π³π¬ Nigerian example: Comparing sales of two products.
π Mini summary: A t-test compares two groups.
Definition: ANOVA compares the means of three or more groups.
Why it is important: It helps you see if multiple groups are different.
Simple explanation: Think of ANOVA like comparing many things at once.
π« School example: Comparing grades across three classes.
π Home example: Comparing expenses across three categories.
π³π¬ Nigerian example: Comparing sales across four regions.
π Mini summary: ANOVA compares three or more groups.
Definition: Correlation measures the relationship between two variables.
Why it is important: It helps you see if two things are related.
Simple explanation: Think of correlation like a friendship between two variables.
π« School example: Correlation between study time and grades.
π Home example: Correlation between income and savings.
π³π¬ Nigerian example: Correlation between advertising spend and sales.
π Mini summary: Correlation measures relationships.
Definition: Linear regression models the relationship between a dependent variable and one or more independent variables.
Why it is important: It helps you predict outcomes.
Simple explanation: Think of linear regression like drawing a line through your data.
π« School example: Predicting grades based on study time.
π Home example: Predicting expenses based on income.
π³π¬ Nigerian example: Predicting sales based on advertising spend.
π Mini summary: Linear regression predicts outcomes.
Definition: Interpreting results means understanding what your analysis tells you.
Why it is important: You need to understand your results to make decisions.
Simple explanation: Think of interpreting results like reading a map.
π« School example: A teacher interprets test scores.
π Home example: Your parents interpret expense data.
π³π¬ Nigerian example: A business interprets sales data.
π Mini summary: Interpreting results helps you make decisions.
Definition: Nigerian businesses use statistics to make decisions.
Why it is important: Statistics help Nigerian businesses understand their data.
Simple explanation: Think of it like using a compass to find your way.
π³π¬ Nigerian example: A Lagos supermarket uses statistics to understand sales patterns.
π Mini summary: Nigerian businesses use statistics to make decisions.
| Word | Simple Meaning |
|---|---|
| Statistics | The science of understanding data. |
| Mean | The average. |
| Median | The middle value. |
| Standard Deviation | How spread out the data is. |
| Population | The entire group. |
| Sample | A part of the group. |
| Hypothesis | A testable statement. |
| p-value | The probability of seeing the data if the null hypothesis is true. |
| t-test | Compares two groups. |
| ANOVA | Compares three or more groups. |
| Correlation | Measures relationships. |
| Linear Regression | Predicts outcomes. |
| R-squared | How well the model fits. |
| Confidence Interval | The range where the true value lies. |
| Significance Level | Usually 0.05. |
+-------------------+
| DATA | β 10, 20, 30, 40, 50
+-------------------+
|
V
+-------------------+
| MEAN | β 30
+-------------------+
| MEDIAN | β 30
+-------------------+
| STANDARD DEV | β 15.81
+-------------------+
+-------------------+
| GROUP A | β 10, 20, 30
+-------------------+
| GROUP B | β 40, 50, 60
+-------------------+
|
V
+-------------------+
| t-TEST | β p-value = 0.02
+-------------------+
|
V
+-------------------+
| SIGNIFICANT | β Reject H0
+-------------------+
+-------------------+
| DATA POINTS | β * * *
+-------------------+
|
V
+-------------------+
| REGRESSION LINE | β y = 2x + 3
+-------------------+
|
V
+-------------------+
| PREDICT | β y = 2(5) + 3 = 13
+-------------------+
| Test | Purpose | Number of Groups |
|---|---|---|
| t-test | Compare means | 2 |
| ANOVA | Compare means | 3+ |
| Correlation | Measure relationship | 2 variables |
| Linear Regression | Predict outcomes | 2+ variables |
Lesson 1: Statistics is the science of understanding data.
Lesson 2: Descriptive statistics summarise data.
Lesson 3: Population is the whole group; sample is a part.
Lesson 4: Hypothesis testing helps you make decisions.
Lesson 5: A t-test compares two groups.
Lesson 6: ANOVA compares three or more groups.
Lesson 7: Correlation measures relationships.
Lesson 8: Linear regression predicts outcomes.
Lesson 9: Interpreting results helps you make decisions.
Lesson 10: Nigerian businesses use statistics to make decisions.
In this module, you learned about statistical analysis and hypothesis testing. You discovered descriptive statistics, the difference between population and sample, and hypothesis testing. You also learned about t-tests, ANOVA, correlation, and linear regression.
π― You can now:
Match the word on the left with the correct meaning on the right:
| Word | Meaning |
|---|---|
| t-test | Measures relationships |
| ANOVA | Compares two groups |
| Correlation | Compares three or more groups |
| Regression | Predicts outcomes |
Answers: t-test β Compares two groups; ANOVA β Compares three or more groups; Correlation β Measures relationships; Regression β Predicts outcomes.
Scenario 1: Ada wants to know if students who study more get better grades. What statistical test should she use?
Scenario 2: A Lagos supermarket wants to know if there is a relationship between advertising spend and sales. What statistical test should they use?
In groups of 4β5, discuss how you would use statistics to analyse a dataset. What tests would you use? Present your ideas to the class.
Think about a dataset you would like to analyse. Write a short paragraph (about 100 words) about what statistical tests you would use.
Analyse a Dataset
Find a dataset and perform statistical analysis. Calculate descriptive statistics, perform a t-test, and create a linear regression model. Save and share your results with the class.
Use R to perform statistical analysis on a dataset. Calculate descriptive statistics, perform a t-test, and create a linear regression model. Write a short report (about 150 words) about what you did and what you learned.
The Challenge: Imagine you are a data analyst in Lagos. You have sales data for four regions. Perform statistical analysis to see if there is a significant difference in sales between regions. Use ANOVA and interpret the results.
Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A
Fill-in-the-Blank: 1. Statistics, 2. Descriptive, 3. Population, sample, 4. t, 5. Correlation
True or False: 1. False, 2. True, 3. False, 4. True, 5. False
Matching: t-test β Compares two groups; ANOVA β Compares three or more groups; Correlation β Measures relationships; Regression β Predicts outcomes.
In the next module, we will explore Advanced Modeling and Machine Learning. You will learn how to build predictive models and use machine learning algorithms.
π Congratulations! You have completed Module Five of the R for Data Analysis course.
π You are now ready to move to Module Six: Advanced Modeling and Machine Learning.
Welcome back, young data scientist! Today we learn how to build models and make predictions.
In Module Five, you learned how to test hypotheses and understand relationships. Now you are ready for the next step: predicting the future!
Imagine you are a weather forecaster. You look at past weather data to predict tomorrow's weather. That is what machine learning does β it uses past data to predict future outcomes.
Machine learning is a type of artificial intelligence that helps computers learn from data. It is like teaching a computer to recognise patterns.
In this module, you will learn about linear regression, logistic regression, decision trees, random forests, and model evaluation. By the end, you will be able to build your own predictive models.
π‘ Think about it: Have you ever tried to predict something? Machine learning helps you do that with data.
By the time you finish this module, you will be able to:
Chidi is a 12-year-old boy who loves to predict the weather. He noticed that when it was cloudy in the morning, it often rained in the afternoon. He wanted to predict if it would rain.
His uncle, who works with data, said, "Chidi, you are doing machine learning. You are using past data (cloudy mornings) to predict future outcomes (rain)."
Chidi learned how to use R to build predictive models. He could now predict the weather with data. He was like a real data scientist.
π§ Think about it: Have you ever tried to predict something? Machine learning helps you do that with data.
Definition: Machine learning is a type of artificial intelligence that allows computers to learn from data.
Why it is important: Machine learning helps you make predictions and decisions.
Simple explanation: Think of machine learning like teaching a computer to recognise patterns.
π« School example: A teacher uses past test scores to predict future performance.
π Home example: Your parents use past spending to predict future expenses.
π³π¬ Nigerian example: A business uses past sales to predict future sales.
+-------------------+
| MACHINE LEARNING |
+-------------------+
| Supervised | β Labelled data
| Unsupervised | β Unlabelled data
| Reinforcement | β Rewards and punishments
+-------------------+
π Mini summary: Machine learning is teaching computers to learn from data.
Definition: Supervised learning uses labelled data. Unsupervised learning uses unlabelled data.
Why it is important: You need to choose the right type for your problem.
Simple explanation: Think of supervised learning like learning with a teacher. Unsupervised learning is like learning by yourself.
π« School example: Supervised: predicting grades. Unsupervised: grouping students by interests.
π Home example: Supervised: predicting expenses. Unsupervised: grouping expenses by category.
π³π¬ Nigerian example: Supervised: predicting sales. Unsupervised: grouping customers by behaviour.
π Mini summary: Supervised learning uses labelled data; unsupervised learning uses unlabelled data.
Definition: Linear regression predicts a continuous outcome.
Why it is important: It is the simplest and most common predictive model.
Simple explanation: Think of linear regression like drawing a line through your data.
π« School example: Predicting grades from study time.
π Home example: Predicting expenses from income.
π³π¬ Nigerian example: Predicting sales from advertising spend.
π Mini summary: Linear regression predicts continuous outcomes.
Definition: Logistic regression predicts a binary outcome (yes/no).
Why it is important: It is used for classification problems.
Simple explanation: Think of logistic regression like answering yes/no questions.
π« School example: Predicting if a student will pass or fail.
π Home example: Predicting if you will overspend this month.
π³π¬ Nigerian example: Predicting if a customer will churn.
π Mini summary: Logistic regression predicts binary outcomes.
Definition: A decision tree is a tree-like model that makes decisions based on rules.
Why it is important: Decision trees are easy to understand and interpret.
Simple explanation: Think of a decision tree like a flowchart.
π« School example: A decision tree to determine if a student passes.
π Home example: A decision tree to decide if you should buy something.
π³π¬ Nigerian example: A decision tree to determine if a customer will churn.
+-------------------+
| DECISION TREE |
+-------------------+
| Tenure > 12? |
| / \ |
| Yes No |
| / \ |
| Churn? ... |
+-------------------+
π Mini summary: Decision trees make decisions based on rules.
Definition: Random forest is a collection of many decision trees.
Why it is important: Random forests are more accurate than single decision trees.
Simple explanation: Think of random forest like many experts voting on a decision.
π« School example: Predicting student performance with many teachers.
π Home example: Predicting expenses with many family members.
π³π¬ Nigerian example: Predicting sales with many predictors.
π Mini summary: Random forest combines many decision trees.
Definition: Model evaluation is the process of assessing how well your model performs.
Why it is important: You need to know if your model is any good.
Simple explanation: Think of model evaluation like grading your model.
π« School example: Testing a student's knowledge with an exam.
π Home example: Testing a recipe before serving it.
π³π¬ Nigerian example: Testing a sales model before using it.
π Mini summary: Model evaluation assesses how well your model performs.
Definition: Overfitting is when a model learns the training data too well. Underfitting is when a model is too simple.
Why it is important: You want a model that generalises well to new data.
Simple explanation: Think of overfitting like memorising the answers. Underfitting is like not studying enough.
π« School example: A student who memorises the textbook but can't answer new questions.
π Home example: A recipe that works only with specific ingredients.
π³π¬ Nigerian example: A sales model that works only for one region.
π Mini summary: Overfitting is memorising the training data; underfitting is too simple.
Definition: Cross-validation is a technique to evaluate model performance by splitting data multiple times.
Why it is important: It gives a more reliable estimate of model performance.
Simple explanation: Think of cross-validation like taking multiple tests instead of just one.
π« School example: Testing a student with multiple quizzes.
π Home example: Testing a recipe with multiple taste tests.
π³π¬ Nigerian example: Testing a sales model with multiple time periods.
π Mini summary: Cross-validation gives a reliable estimate of model performance.
Definition: Nigerian businesses use machine learning to make predictions and decisions.
Why it is important: Machine learning helps Nigerian businesses grow.
Simple explanation: Think of it like a crystal ball for business.
π³π¬ Nigerian example: A Lagos supermarket uses machine learning to predict sales.
π Mini summary: Nigerian businesses use machine learning to grow.
| Word | Simple Meaning |
|---|---|
| Machine Learning | Teaching computers to learn from data. |
| Supervised Learning | Learning from labelled data. |
| Unsupervised Learning | Learning from unlabelled data. |
| Linear Regression | Predicts continuous outcomes. |
| Logistic Regression | Predicts binary outcomes. |
| Decision Tree | Makes decisions based on rules. |
| Random Forest | Combines many decision trees. |
| Overfitting | Memorising the training data. |
| Underfitting | Too simple to learn patterns. |
| Cross-Validation | Evaluates model performance. |
| Accuracy | Percentage of correct predictions. |
| Confusion Matrix | Shows true/false positives and negatives. |
| R-squared | How well the model fits the data. |
| RMSE | Root mean squared error. |
| Ensemble Method | Combines multiple models. |
+-------------------+
| MACHINE LEARNING |
+-------------------+
| Supervised | β Labelled data
| Unsupervised | β Unlabelled data
| Reinforcement | β Rewards and punishments
+-------------------+
+-------------------+
| DATA POINTS | β * * *
+-------------------+
|
V
+-------------------+
| REGRESSION LINE | β y = 2x + 3
+-------------------+
|
V
+-------------------+
| PREDICT | β y = 2(5) + 3 = 13
+-------------------+
+-------------------+
| Tenure > 12? |
| / \ |
| Yes No |
| / \ |
| Churn? ... |
+-------------------+
+-------------------+
| Tree 1 |
+-------------------+
| Tree 2 |
+-------------------+
| Tree 3 |
+-------------------+
|
V
+-------------------+
| Vote | β Combine predictions
+-------------------+
Lesson 1: Machine learning teaches computers to learn from data.
Lesson 2: Supervised learning uses labelled data; unsupervised uses unlabelled.
Lesson 3: Linear regression predicts continuous outcomes.
Lesson 4: Logistic regression predicts binary outcomes.
Lesson 5: Decision trees make decisions based on rules.
Lesson 6: Random forest combines many decision trees.
Lesson 7: Model evaluation assesses performance.
Lesson 8: Overfitting is memorising; underfitting is too simple.
Lesson 9: Cross-validation gives reliable performance estimates.
Lesson 10: Nigerian businesses use machine learning to grow.
In this module, you learned about advanced modeling and machine learning. You discovered supervised and unsupervised learning, linear and logistic regression, decision trees, random forest, and model evaluation. You also learned about overfitting, underfitting, and cross-validation.
π― You can now:
Match the word on the left with the correct meaning on the right:
| Word | Meaning |
|---|---|
| Regression | Predicts categories |
| Classification | Predicts numbers |
| Decision Tree | Combines many trees |
| Random Forest | Makes decisions based on rules |
Answers: Regression β Predicts numbers; Classification β Predicts categories; Decision Tree β Makes decisions based on rules; Random Forest β Combines many trees.
Scenario 1: Chidi wants to predict if it will rain tomorrow based on temperature and humidity. What type of model should he use?
Scenario 2: A Lagos supermarket wants to predict sales for the next month. What type of model should they use?
In groups of 4β5, discuss how you would use machine learning to solve a problem. What model would you use? Present your ideas to the class.
Think about a problem you could solve with machine learning. Write a short paragraph (about 100 words) about what model you would use.
Build a Machine Learning Model
Find a dataset and build a machine learning model. Split the data, train the model, and evaluate its performance. Save and share your results with the class.
Use R to build a machine learning model on a dataset. Split the data, train the model, and evaluate its performance. Write a short report (about 150 words) about what you did and what you learned.
The Challenge: Imagine you are a data scientist in Lagos. You have customer data and want to predict churn. Build a machine learning model to predict churn. Use logistic regression, decision tree, and random forest. Compare their performance.
Multiple Choice: 1-A, 2-A, 3-A, 4-A, 5-A, 6-A, 7-A, 8-A, 9-A, 10-D, 11-A, 12-A, 13-A, 14-A, 15-A
Fill-in-the-Blank: 1. Supervised, 2. Linear, 3. Logistic, 4. decision, 5. Random
True or False: 1. False, 2. True, 3. False, 4. True, 5. False
Matching: Regression β Predicts numbers; Classification β Predicts categories; Decision Tree β Makes decisions based on rules; Random Forest β Combines many trees.
In the next module, we will explore Reporting with R Markdown and Shiny. You will learn how to create dynamic reports and interactive dashboards.
π Congratulations! You have completed Module Six of the R for Data Analysis course.
π You are now ready to move to Module Seven: Reporting with R Markdown and Shiny.
Hello, young data explorer! In this module, we are going to learn how to tell stories with data. Numbers and tables are useful, but sometimes they are hard to understand. That is why we use pictures β also called graphs or charts β to show our data in a fun and clear way.
Imagine you have a big bag of colourful sweets. If I just tell you the numbers, you might forget. But if I draw a picture with coloured bars, you will see which colour is the most popular right away! That is what this module is about: turning numbers into pictures so everyone can understand them easily.
We will use the R programming language to create these pictures. R has special tools β called packages β
that help us draw beautiful graphs. We will learn the most important ones: ggplot2 and plotly.
By the end of this module, you will be able to look at a table of numbers, choose the best graph, and draw it using R. You will be a data storyteller!
ggplot2.ggplot2.plotly.At Sunshine Primary School, the pupils voted for their favourite school lunch. The choices were: Jollof Rice, Fried Plantain, Pounded Yam, and Beans. The head teacher, Mrs. Ade, received the votes as numbers:
She wrote these numbers on the board. The children could not quickly see which food was the winner. It was confusing! Then her assistant, Mr. Bello, said, βLet me draw a picture.β He made a bar chart with colourful bars. Instantly, everyone saw that Jollof Rice was the most loved food. The children cheered!
That is the power of data visualisation. It turns boring numbers into a story that everyone can understand in seconds. In this module, you will learn how to create these magical pictures using R.
Definition: Data visualisation means making pictures (graphs, charts, maps) from data (numbers and facts).
Why it is important: Our brains understand pictures faster than numbers. A picture can show patterns, trends, and outliers (things that don't fit) very quickly.
Simple explanation: Think of your data as a story. The graph is the illustration that makes the story exciting.
Real-life example: Weather forecasters use maps with colours to show temperature β red for hot, blue for cold.
School example: Teachers draw bar charts to show how many pupils got A, B, C in a test.
Home example: You can draw a picture of how you spend your day: school, play, homework, sleep.
Nigerian example: The National Bureau of Statistics uses graphs to show Nigeria's population growth.
Illustration:
DATA (numbers) --> VISUALISATION (pictures) --> UNDERSTANDING
(45, 30, 20, 5) (bar chart with bars) (Jollof wins!)
Mini summary: Visualisation turns numbers into pictures so we can see what the data is telling us.
Definition: ggplot2 is a special R package that helps us create beautiful graphs. It is based on the "grammar of graphics" β a set of rules for building graphs layer by layer.
Why it is important: ggplot2 is the most popular drawing tool in R. It is flexible and powerful.
Simple explanation: Imagine you are building a house with LEGO blocks. ggplot2 gives you blocks (layers) to build your graph: data, axes, bars, colours, titles.
Real-life example: An artist uses layers of paint to make a painting. ggplot2 uses layers to make graphs.
School example: In art class, you draw a background, then add a tree, then add birds. ggplot2 adds one layer at a time.
Home example: When you make a sandwich, you add bread, then cheese, then ham, then bread. Layers!
Nigerian example: A Nigerian data analyst uses ggplot2 to show the yield of cassava in different states.
Illustration:
DATA + AESTHETICS (x, y) + GEOM (bars, lines) + THEME = GRAPH (layer 1) (layer 2) (layer 3) (layer 4)
Mini summary: ggplot2 builds graphs in layers like you build a LEGO castle.
Definition: To use ggplot2, we first need to install it (only once) and then load it (every time we start R).
Why it is important: You cannot use a tool until you bring it to your workspace.
Simple explanation: It is like buying a new board game (install) and then taking it out of the box to play (load).
Real-life example: You install a game app on your phone, then you open it to play.
School example: Your teacher gives you a textbook (install) and you open it to read (load).
Home example: Mum buys flour (install) and then opens the bag to bake (load).
Nigerian example: A student in Lagos installs R packages once, then loads them for each project.
Code:
# Install (do this once)
install.packages("ggplot2")
# Load (do this every time)
library(ggplot2)
Mini summary: Install once, load every session.
Definition: A bar chart uses bars to show counts or values for different categories.
Why it is important: It is the easiest way to compare categories.
Simple explanation: The taller the bar, the bigger the number.
Real-life example: A supermarket uses a bar chart to show which fruit sells most.
School example: A bar chart of favourite sports: football, basketball, tennis.
Home example: A bar chart of how many hours you read each day.
Nigerian example: A bar chart showing the number of pupils in each class in a school in Abuja.
Code & Illustration:
# Sample data: favourite foods
foods <- data.frame(
Food = c("Jollof", "Plantain", "Yam", "Beans"),
Votes = c(45, 30, 20, 5)
)
# ggplot bar chart
ggplot(foods, aes(x = Food, y = Votes)) +
geom_bar(stat = "identity") +
labs(title = "Favourite School Lunch", x = "Food", y = "Votes")
Bar chart (concept):
Votes
50 | β
40 | β
30 | β β
20 | β β β
10 | β β β β
0 |____β ___β ___β ___β ____
J P Y B
Mini summary: geom_bar() draws the bars. Use stat="identity" when you have values.
Definition: A line chart connects data points with lines to show how something changes over time.
Why it is important: It helps us see trends: going up (increase), going down (decrease), or staying flat.
Simple explanation: It is like connecting the dots in a dot-to-dot puzzle, but the dots are data points.
Real-life example: A line chart of temperature every day of the week.
School example: A line chart of your test scores over the term.
Home example: A line chart of your piggy bank savings each month.
Nigerian example: A line chart showing the price of tomatoes in Lagos market over the year.
Code:
# Sample data: monthly rainfall (mm)
rain <- data.frame(
Month = c("Jan", "Feb", "Mar", "Apr"),
Rainfall = c(15, 20, 35, 40)
)
ggplot(rain, aes(x = Month, y = Rainfall, group = 1)) +
geom_line() +
geom_point() + # adds dots
labs(title = "Monthly Rainfall", x = "Month", y = "Rainfall (mm)")
Line chart concept:
Rainfall
40 | β
35 | β
30 |
25 |
20 | β
15 | β
0 |___β___β___β___β____
J F M A
Mini summary: Use geom_line() for trends, add geom_point() to show data points.
Definition: A pie chart is a circle divided into slices. Each slice represents a part of the total (100%).
Why it is important: It shows proportions (what fraction of the whole each category is).
Simple explanation: Imagine a pizza cut into slices. Bigger slice = bigger share.
Real-life example: A pie chart of how you spend your 24 hours: sleep, school, play, chores.
School example: A pie chart showing the percentage of pupils who like each subject.
Home example: A pie chart of your weekly allowance spending.
Nigerian example: A pie chart of Nigeria's exports: oil, agriculture, etc.
Code (using ggplot2 with coord_polar):
# We use geom_bar() + coord_polar()
foods <- data.frame(
Food = c("Jollof", "Plantain", "Yam", "Beans"),
Votes = c(45, 30, 20, 5)
)
ggplot(foods, aes(x = "", y = Votes, fill = Food)) +
geom_bar(stat = "identity", width = 1) +
coord_polar("y", start = 0) +
labs(title = "Pie Chart of Favourite Foods")
Pie chart concept:
________
/ Jollof \
| 45% |
| (big) |
\________/
/ Plantain \
| 30% |
\________/
/ Yam 20% \
/ Beans 5% \
Mini summary: Pie charts show parts of a whole. Use coord_polar() to turn a bar chart into a pie chart.
Definition: Titles and labels are words that tell the reader what the graph is about.
Why it is important: Without labels, your graph is just a picture. Labels make it meaningful.
Simple explanation: Labels are like name tags for your graph β they tell what each part means.
Real-life example: A map without place names is useless. Labels tell you the names of cities.
School example: A graph of test scores needs a title "Test Scores" and labels for subjects.
Home example: A bar chart of chores needs labels: "Wash dishes", "Sweep", etc.
Nigerian example: A graph of state populations must label each state.
Code (using labs()):
ggplot(foods, aes(x = Food, y = Votes)) +
geom_bar(stat = "identity") +
labs(
title = "Favourite School Lunch",
subtitle = "Sunshine Primary School Election",
x = "Food Items",
y = "Number of Votes",
caption = "Data from the school election"
)
[Title] Favourite School Lunch
[Subtitle] Sunshine Primary School Election
Votes |
45 | [bar]
30 | [bar] [bar]
20 | [bar] [bar] [bar]
...
[x-axis label] Food Items
[y-axis label] Number of Votes
[caption] Data from the school election
Mini summary: Always use labs() to add title, subtitle, x, y, and caption.
Definition: Colours make your graph attractive and help distinguish different groups.
Why it is important: Colours catch the eye and make comparisons easier.
Simple explanation: Colour is like using different crayons to colour different parts of your drawing.
Real-life example: Traffic lights use red, yellow, green to convey meaning.
School example: In a bar chart, you colour each bar differently to show different classes.
Home example: You use a red marker for important events on your calendar.
Nigerian example: Using green-white-green for Nigerian data highlights.
Code (using fill and scale_fill_manual):
ggplot(foods, aes(x = Food, y = Votes, fill = Food)) +
geom_bar(stat = "identity") +
scale_fill_manual(values = c("red", "yellow", "brown", "green"))
[Red bar for Jollof] [Yellow bar for Plantain] [Brown bar for Yam] [Green bar for Beans]
Mini summary: Use fill inside aes() to colour by category, and scale_fill_manual() to choose colours.
Definition: A theme controls the non-data elements of a graph: background, grid lines, font size, etc.
Why it is important: A clean, professional look helps people focus on the data.
Simple explanation: Think of themes as different outfits for your graph β you can dress it up or keep it simple.
Real-life example: A formal report uses a clean theme; a fun poster uses a colourful theme.
School example: Your teacher might prefer a plain background for clarity.
Home example: You can choose a "cartoon" theme for a fun family graph.
Nigerian example: Use theme_minimal() for a clean, professional look in business reports.
Code:
ggplot(foods, aes(x = Food, y = Votes, fill = Food)) +
geom_bar(stat = "identity") +
theme_minimal() # or theme_classic(), theme_bw()
[Graph with a clean white background, no gridlines, simple fonts]
Mini summary: Themes change the appearance. theme_minimal() is a good default.
Definition: Saving a graph means exporting it as an image file (PNG, JPEG, PDF) so you can use it in reports or presentations.
Why it is important: You often need to share your graphs with others who don't use R.
Simple explanation: It is like taking a screenshot of your graph, but much better quality.
Real-life example: You save a photo to show your friends.
School example: You save your graph to paste into your project document.
Home example: You save a graph of your savings to show your family.
Nigerian example: A researcher saves a graph to include in a report for the government.
Code:
# Create a graph and save it
p <- ggplot(foods, aes(x = Food, y = Votes)) + geom_bar(stat = "identity")
ggsave("favourite_food.png", plot = p, width = 6, height = 4)
Mini summary: Use ggsave() to save your graph as an image.
Definition: plotly is an R package that creates interactive graphs β you can hover, zoom, and click.
Why it is important: Interactive graphs let users explore data themselves.
Simple explanation: It is like a living graph that responds to your mouse.
Real-life example: On a weather website, you can hover over a map to see temperatures.
School example: An interactive graph in a science project where you can click on bars to see details.
Home example: A graph of family expenses that shows details when you hover.
Nigerian example: An interactive dashboard showing COVID-19 cases in Nigeria.
Code:
install.packages("plotly")
library(plotly)
p <- ggplot(foods, aes(x = Food, y = Votes, fill = Food)) +
geom_bar(stat = "identity")
ggplotly(p) # makes it interactive!
[Interactive graph: hover over a bar to see "Votes: 45"]
Mini summary: ggplotly() turns your static graph into an interactive one.
Definition: Not every graph works for every type of data. You need to choose wisely.
Why it is important: The wrong graph can confuse people; the right graph makes the story clear.
Simple explanation: It is like choosing the right shoe for the occasion β sandals for the beach, boots for rain.
Real-life example: Use a line chart for trends over time, bar chart for comparing categories.
School example: Bar chart for favourite subjects; line chart for temperature change.
Home example: Pie chart for budget; bar chart for chores done.
Nigerian example: Bar chart for state populations; line chart for GDP growth.
Decision table:
| Purpose | Best Graph |
|---|---|
| Compare categories | Bar chart |
| Show trends over time | Line chart |
| Show parts of a whole | Pie chart |
| Show distribution | Histogram |
| Show relationship | Scatter plot |
Mini summary: Match the graph to your question: compare? use bar. trend? use line. part of whole? use pie.
Definition: Mistakes are things we do wrong that make our graph hard to understand.
Why it is important: Avoiding mistakes makes our graphs clear and trustworthy.
Simple explanation: It is like baking: if you add too much salt, the cake tastes bad. With graphs, small errors can confuse.
Real-life example: Forgetting to label the y-axis means no one knows what the numbers mean.
School example: Using too many colours makes the graph look messy.
Home example: Making a pie chart with too many slices β hard to read.
Nigerian example: Not including the source of data can make the graph unreliable.
List of common mistakes:
ggsave().Mini summary: Check labels, colours, graph type β keep it simple and honest.
Let's use data about Nigeria's states and population (example). We will create a bar chart.
# Sample data (population in millions)
states <- data.frame(
State = c("Lagos", "Kano", "Oyo", "Rivers"),
Population = c(20, 15, 8, 7)
)
ggplot(states, aes(x = State, y = Population, fill = State)) +
geom_bar(stat = "identity") +
labs(title = "Population of Selected Nigerian States",
x = "State", y = "Population (millions)") +
theme_minimal()
[Bar chart with Lagos tallest, then Kano, Oyo, Rivers]
Mini summary: You can apply the same R code to any data, including Nigerian data.
We have learned about bar charts, line charts, pie charts, colours, labels, themes, saving, and interactive graphs.
Remember the golden rule: Keep it clear, keep it simple, and always label your axes!
library(ggplot2)ggplot(data, aes(x, y))+ geom_bar(stat = "identity")+ labs(title = "...", x = "...", y = "...")+ theme_minimal()ggsave("filename.png")gganimate.ggsave().library(ggplot2)).geom_bar() without stat="identity" when you have pre-summarised data.
Votes
50 | β
40 | β
30 | β β
20 | β β β
10 | β β β β
0 |____β ___β ___β ___β ____
J P Y B
40 | β
35 | β
30 |
25 |
20 | β
15 | β
0 |___β___β___β___β____
J F M A
________
/ Jollof \
| 45% |
\________/
/ Plantain \
| 30% |
\________/
/ Yam 20% \
/ Beans 5% \
| Graph Type | Best Used For | Example |
|---|---|---|
| Bar Chart | Comparing categories | Favourite foods |
| Line Chart | Showing trends over time | Temperature changes |
| Pie Chart | Showing parts of a whole | Budget allocation |
In this module, we learned that data visualisation is the art of turning numbers into pictures. We discovered the ggplot2 package, which builds graphs layer by layer. We created bar charts to compare categories, line charts to show trends, and pie charts to show parts of a whole. We added titles, labels, and colours to make our graphs clear and beautiful. We also learned to save our graphs and make them interactive with plotly. Remember: a good graph tells a story without confusion. Keep practising with data from your school, home, and country!
install.packages("ggplot2").labs(title = "My Title").scale_fill_manual().plotly package and ggplotly().geom_bar() do?geom_line() do?plotly package do?ggsave(). (True)geom_bar(stat = "identity") do?| Term | Definition |
|---|---|
| 1. Bar chart | A. Shows parts of a whole |
| 2. Line chart | B. Compares categories with bars |
| 3. Pie chart | C. Shows trends over time |
| 4. ggplot2 | D. Interactive graphs package |
| 5. plotly | E. R package for layering graphs |
Answers: 1-B, 2-C, 3-A, 4-E, 5-D
ggsave() do?In groups of four, collect data on the favourite fruits of your classmates (e.g., mango, orange, banana, apple). Create a bar chart using R. Present your graph to the class and explain what it shows.
Collect data on how you spend your time in a day (sleep, school, play, eating, homework). Create a pie chart using R. Write one sentence about what the pie chart tells you.
Title: "Our School Data Story"
Collect data on the number of pupils in each class in your school (or use sample data). Create a bar chart and a line chart (if you have data over time). Write a short report (2-3 paragraphs) explaining what the graphs show and why it is important.
Using R, create the following graphs from the mtcars dataset (built-in):
cyl).mpg (miles per gallon) vs wt (weight) β use geom_line().Find a real dataset about Nigeria (e.g., population, agriculture, weather). Create at least three different types of graphs (bar, line, pie) and arrange them in a report using R Markdown (optional). Explain the story behind each graph.
Fill-in-the-Blank: 1. visualisation, 2. ggplot2, 3. bar, 4. line, 5. pie, 6. labs, 7. plotly, 8. ggsave, 9. theme, 10. fill.
True/False: 1F, 2T, 3F, 4T, 5F, 6T, 7T, 8F, 9F, 10T.
Multiple Choice: 1B, 2C, 3B, 4B, 5C, 6B, 7C, 8C, 9B, 10B, 11C, 12A, 13A, 14B, 15C.
ggsave().plotly.In Module 8, we will learn about data wrangling β cleaning and preparing data for analysis.
We will use the dplyr package to filter, sort, and summarise data.
Before the next module, try to practise loading dplyr with library(dplyr) and explore the iris dataset.
Great job completing Module 7! You are now a data storyteller.
Hello, young data explorer! In Module 7, we learned how to tell stories with data using pictures and graphs. But before we can make beautiful graphs, we need to make sure our data is clean and ready. Imagine you want to bake a cake β you need to wash the fruits, measure the flour, and mix everything properly. Data is just the same!
In this module, we will learn how to wrangle data. Data wrangling means cleaning, organising, and transforming data so that it is easy to work with. We will use a powerful R package called dplyr. It gives us simple "verbs" β like filter(), select(), mutate(), summarise(), and arrange() β that help us change our data in many useful ways.
By the end of this module, you will be able to take messy data, clean it up, and get it ready for analysis and graphs. You will be a data detective!
dplyr package.filter() to choose rows that meet a condition.select() to choose specific columns.mutate() to create new columns.summarise() to get summary statistics (like mean, sum).arrange() to sort data.%>% to chain commands.Chidi loves helping his mother at the market. One day, his mother gave him a messy list of items to buy:
Chidi was confused. The list had duplicates, the items were not sorted, and he couldn't tell how many of each item he really needed. He decided to clean the list:
Now the list was clear:
Chidi felt proud. He had just wrangled the data! In R, we do the same thing with our datasets β we clean them, combine them, and sort them so we can understand them better.
Definition: Data wrangling is the process of cleaning, transforming, and organising data to make it useful.
Why it is important: Real-world data is often messy β it has mistakes, missing values, or is in the wrong format. Wrangling fixes these issues.
Simple explanation: It's like tidying your room before you can find your toys easily.
Real-life example: A librarian organises books by category and author.
School example: Your teacher sorts test scores from highest to lowest.
Home example: You arrange your clothes by colour in the wardrobe.
Nigerian example: A farmer sorts his harvest by crop type and quality.
Illustration:
Messy Data --> Clean Data --> Ready for Analysis
(duplicates, (no duplicates,
errors, correct format,
unsorted) sorted)
Mini summary: Data wrangling makes data clean and tidy so we can work with it easily.
Definition: dplyr is a popular R package that provides simple functions for data manipulation.
Why it is important: It makes data wrangling fast and easy, with intuitive "verbs" (words that describe actions).
Simple explanation: Think of dplyr as a Swiss Army knife for data β it has many tools in one.
Real-life example: A chef has different knives for cutting, peeling, and slicing β dplyr has different verbs for different tasks.
School example: A pencil case contains a pencil, an eraser, and a ruler β each does a specific job.
Home example: A toolbox has a hammer, a screwdriver, and a wrench.
Nigerian example: A mechanic uses different spanners for different bolts.
Code to install and load:
install.packages("dplyr") # do this once
library(dplyr) # do this every session
Mini summary: dplyr is a toolbox of verbs for cleaning and changing data.
Definition: The pipe operator %>% (pronounced "then") lets you chain multiple operations together in a sequence.
Why it is important: It makes your code easier to read and write, like a recipe with steps.
Simple explanation: It means "and then do this". Instead of writing many nested functions, you write a clear list of actions.
Real-life example: A recipe: "Take eggs, then break them, then whisk them."
School example: "Open your book, then read page 10, then answer the questions."
Home example: "Take the laundry, then put it in the washing machine, then turn it on."
Nigerian example: "Get the yam, then peel it, then cut it, then boil it."
Illustration:
data %>%
filter(...) %>%
select(...) %>%
arrange(...)
# This means: take data, THEN filter, THEN select, THEN arrange.
Mini summary: The pipe %>% helps us chain commands in a logical order.
Definition: filter() picks rows that meet a condition (like "age > 10" or "city == 'Lagos'").
Why it is important: Often we only need a part of the data β e.g., only pupils from a certain class.
Simple explanation: It's like using a sieve to separate big pieces from small pieces.
Real-life example: A shopkeeper filters out expired products.
School example: Your teacher filters the list to only show pupils who scored above 80%.
Home example: You filter your toys to only show the ones you want to give away.
Nigerian example: A researcher filters survey data to only include respondents from Kano.
Code:
library(dplyr)
# Sample data: pupils
pupils <- data.frame(
name = c("Ade", "Bola", "Chidi", "Dami"),
age = c(10, 11, 9, 12),
score = c(85, 70, 90, 65)
)
# Filter to only those with score > 80
pupils %>%
filter(score > 80)
# Result: Ade (85) and Chidi (90)
Before filter: name age score Ade 10 85 Bola 11 70 Chidi 9 90 Dami 12 65 After filter(score > 80): name age score Ade 10 85 Chidi 9 90
Mini summary: filter() keeps only rows that satisfy a condition.
Definition: select() picks specific columns (variables) from your data.
Why it is important: Sometimes we have too many columns and only need a few.
Simple explanation: It's like highlighting important words in a paragraph.
Real-life example: When you read a long article, you focus on the main points.
School example: A teacher might only look at the "name" and "score" columns.
Home example: When you check your phone, you only look at messages, not all apps.
Nigerian example: An economist selects only "State" and "GDP" columns from a big table.
Code:
# Select name and score
pupils %>%
select(name, score)
# Result: only those two columns.
Before select: name age score Ade 10 85 Bola 11 70 After select(name, score): name score Ade 85 Bola 70
Mini summary: select() keeps only the columns you choose.
Definition: mutate() creates a new column (variable) from existing columns, or modifies an existing one.
Why it is important: Often we need to calculate new values, like total or percentage.
Simple explanation: It's like adding a new row to your table with new information.
Real-life example: You calculate your total pocket money by adding your allowance and gifts.
School example: You create a column for "percentage" from test scores.
Home example: You convert minutes into hours and minutes.
Nigerian example: A farmer calculates total harvest in kg by adding different crops.
Code:
# Add a new column: score in percentage (out of 100)
pupils %>%
mutate(percentage = score / 100 * 100) # just score itself, but let's add 10 bonus points
# Actually, let's add bonus: total = score + 5
mutate(total_score = score + 5)
Before mutate: name age score Ade 10 85 After mutate(total_score = score + 5): name age score total_score Ade 10 85 90
Mini summary: mutate() adds new columns by performing calculations.
Definition: summarise() (or summarize()) collapses your data into a single summary row β like mean, sum, count.
Why it is important: It gives you the big picture (e.g., average score, total votes).
Simple explanation: It's like asking "What is the total?" or "What is the average?"
Real-life example: You calculate the total money spent on your shopping trip.
School example: You find the average test score of your class.
Home example: You sum up the minutes you spent on homework each day.
Nigerian example: You summarise the total population of all states.
Code:
# Get average age and average score
pupils %>%
summarise(avg_age = mean(age), avg_score = mean(score))
# Result: one row with the averages.
Before summarise: name age score Ade 10 85 Bola 11 70 Chidi 9 90 Dami 12 65 After summarise(avg_age = mean(age), avg_score = mean(score)): avg_age avg_score 10.5 77.5
Mini summary: summarise() gives you one number or a few numbers that describe your data.
Definition: group_by() splits your data into groups, and then you apply summarise() to each group.
Why it is important: It lets you compare groups (e.g., average score by class).
Simple explanation: It's like sorting your toys by colour and then counting how many of each colour you have.
Real-life example: A shop counts total sales per product category.
School example: Find the average score for each class.
Home example: Count how many hours you spend on each activity (play, study, sleep).
Nigerian example: Calculate the average income per state.
Code:
# Group by age, then get average score per age group
pupils %>%
group_by(age) %>%
summarise(avg_score = mean(score))
Data grouped by age: age avg_score 9 90 10 85 11 70 12 65
Mini summary: group_by() + summarise() gives summaries for each group.
Definition: arrange() sorts your rows by one or more columns, in ascending (default) or descending order.
Why it is important: It helps you see the highest, lowest, or alphabetical order.
Simple explanation: It's like putting your books on a shelf from A to Z.
Real-life example: A teacher sorts pupils' names alphabetically.
School example: You sort test scores from highest to lowest.
Home example: You arrange your clothes from biggest to smallest.
Nigerian example: A business owner sorts sales by highest revenue.
Code:
# Sort by score ascending (lowest to highest)
pupils %>%
arrange(score)
# Sort by score descending (highest first)
pupils %>%
arrange(desc(score))
Original: name score Ade 85 Bola 70 Chidi 90 Dami 65 arrange(desc(score)): name score Chidi 90 Ade 85 Bola 70 Dami 65
Mini summary: arrange() sorts rows by column values.
We can combine filter(), select(), mutate(), arrange(), and summarise() in a single pipeline.
Example: We want to:
pupils %>%
filter(score > 70) %>%
select(name, score) %>%
mutate(bonus = score + 5) %>%
arrange(desc(score))
Result: name score bonus Chidi 90 95 Ade 85 90
Mini summary: Pipes let you chain multiple verbs in one clear sequence.
Definition: Missing data are empty values, shown as NA in R.
Why it is important: Missing values can break calculations if not handled.
Simple explanation: It's like a page missing from your book β you need to decide what to do.
Real-life example: A survey form where someone didn't answer a question.
School example: A pupil was absent on the day of the test.
Home example: You forgot to record the time you spent reading.
Nigerian example: A farmer didn't record the harvest for one month.
How to handle: Use filter() to remove NAs, or use na.rm = TRUE inside functions like mean().
# Remove rows with NA in score pupils %>% filter(!is.na(score)) # Get mean ignoring NAs summarise(avg = mean(score, na.rm = TRUE))
Mini summary: Always check for NA values and decide to remove them or ignore them in calculations.
Definition: rename() changes the names of columns.
Why it is important: Column names might be too long, misspelled, or unclear.
Simple explanation: It's like giving your pet a new nickname.
Real-life example: You change a file name to make it more descriptive.
School example: Your teacher renames "Test_1" to "Math_Test".
Home example: You rename your folders from "Stuff" to "Schoolwork".
Nigerian example: An analyst renames a column from "Pop" to "Population".
pupils %>%
rename(Score = score, Name = name)
Mini summary: rename() changes column names to be clearer.
Definition: distinct() removes duplicate rows from your data.
Why it is important: Duplicates can overcount and give wrong results.
Simple explanation: It's like removing double entries in your list.
Real-life example: You delete duplicate contacts on your phone.
School example: You remove duplicate names from the class register.
Home example: You make sure you don't have two of the same toy listed.
Nigerian example: A voter register removes duplicate names.
pupils %>% distinct() # removes exact duplicate rows pupils %>% distinct(name, .keep_all = TRUE) # keep only unique names
Mini summary: distinct() removes duplicate rows.
Let's use a small dataset about Nigerian states and their populations (example). We will wrangle it.
# Sample data (population in millions)
states <- data.frame(
State = c("Lagos", "Kano", "Oyo", "Lagos", "Rivers", "Kano"),
Population = c(20, 15, 8, 20, 7, 15),
Region = c("SW", "NW", "SW", "SW", "SS", "NW")
)
# We have duplicates! Let's remove them.
states_unique <- states %>% distinct(State, .keep_all = TRUE)
# Filter to only SW states
sw_states <- states_unique %>% filter(Region == "SW")
# Select State and Population
sw_pop <- sw_states %>% select(State, Population)
# Arrange by population descending
sw_pop %>% arrange(desc(Population))
Result: State Population Lagos 20 Oyo 8
Mini summary: You can use all dplyr verbs on Nigerian data to clean and explore it.
We have learned six main verbs:
filter() β pick rows.select() β pick columns.mutate() β create new columns.summarise() β get summaries (mean, sum, etc.).arrange() β sort rows.group_by() β split data into groups.And the pipe %>% connects them.
Remember: These verbs make data wrangling fun and easy!
%>% operator makes code readable.dplyr functions do not change the original data unless you assign the result.library(dplyr)data %>% distinct()data %>% filter(condition)data %>% select(col1, col2)data %>% mutate(new = old * 2)data %>% summarise(mean = mean(col))data %>% arrange(desc(col))filter() to find products with low stock.arrange() to sort students by grades.summarise() to get average patient age.filter() to get data for only Lagos State.group_by() and summarise() to calculate average income by region.mutate() to add a "pass/fail" column based on scores.dplyr package was created by Hadley Wickham, a famous data scientist.dplyr is used by companies like Facebook and Google.dplyr is one of the most popular.dplyr can work with large datasets very quickly because it uses C++ behind the scenes.dplyr with databases without loading all data into memory.dtplyr that uses dplyr syntax on data.table for even more speed.dplyr with library(dplyr).%>% to chain commands.filter() picks rows, select() picks columns.mutate() creates new columns, summarise() gives summaries.arrange() sorts, group_by() splits into groups.dplyr before using its functions.filter() with the wrong condition (e.g., == instead of =).na.rm = TRUE when calculating means with missing values.dplyr does not modify in place.NA values early.glimpse() or str() to understand your data before wrangling.
data
|
V
filter() --> (keeps rows)
|
V
select() --> (keeps columns)
|
V
mutate() --> (adds columns)
|
V
arrange() --> (sorts rows)
Before filter: +------+-----+-------+ | name | age | score | +------+-----+-------+ | Ade | 10 | 85 | | Bola | 11 | 70 | | Chidi| 9 | 90 | +------+-----+-------+ filter(score > 80): +------+-----+-------+ | name | age | score | +------+-----+-------+ | Ade | 10 | 85 | | Chidi| 9 | 90 | +------+-----+-------+
Before select: +------+-----+-------+ | name | age | score | +------+-----+-------+ | Ade | 10 | 85 | +------+-----+-------+ select(name, score): +------+-------+ | name | score | +------+-------+ | Ade | 85 | +------+-------+
| Verb | What it does | Example |
|---|---|---|
| filter() | Picks rows based on condition | filter(score > 80) |
| select() | Picks columns by name | select(name, score) |
| mutate() | Creates new columns | mutate(total = score + 5) |
| summarise() | Calculates summary stats | summarise(avg = mean(score)) |
| arrange() | Sorts rows | arrange(desc(score)) |
| group_by() | Splits data into groups | group_by(age) |
In this module, we learned how to wrangle data β that is, clean and prepare it for analysis.
We used the dplyr package, which provides simple verbs: filter(), select(), mutate(),
summarise(), arrange(), and group_by(). We also learned the pipe %>% to chain commands.
We handled missing values (NA), removed duplicates, and renamed columns.
Data wrangling is an essential step before any analysis or visualisation.
Remember: clean data leads to clear insights!
install.packages("dplyr")distinct().arrange().| Verb | Description |
|---|---|
| 1. filter | A. Adds new columns |
| 2. select | B. Picks rows that meet condition |
| 3. mutate | C. Chooses columns |
| 4. summarise | D. Sorts rows |
| 5. arrange | E. Calculates summary statistics |
Answers: 1-B, 2-C, 3-A, 4-E, 5-D
In groups of four, create a small dataset (10 rows) about your favourite foods, with columns: food, cost, rating. Then use dplyr to:
Present your code and output to the class.
Use the iris dataset (built-in) and dplyr to:
Title: "Clean the School Data"
You are given a messy dataset of school attendance (columns: name, class, days_present, days_absent) with duplicates, missing values, and errors. Use dplyr to:
Using the mtcars dataset, perform the following using dplyr:
Find a real, messy dataset online (e.g., from Kaggle or Nigerian open data). Use dplyr to perform a complete cleaning and summarisation. Write a short report (2 paragraphs) describing your steps and what you discovered.
Fill-in-the-Blank: 1. wrangling, 2. dplyr, 3. %>%, 4. filter, 5. select, 6. mutate, 7. summarise, 8. arrange, 9. group_by, 10. NA.
True/False: 1F, 2T, 3F, 4F, 5F, 6T, 7T, 8F, 9T, 10T.
Multiple Choice: 1B, 2B, 3B, 4B, 5B, 6A, 7A, 8B, 9A, 10C, 11A, 12C, 13B, 14A, 15C.
filter(), select(), mutate(), summarise(), arrange(), group_by().%>% chains commands.In Module 9, we will learn about joining datasets β combining information from different tables, like merging pupil data with class data.
We will use dplyr joins: left_join(), inner_join(), and more.
Before that, practise the verbs you learned in this module on different datasets.
Well done! You are becoming a data wrangler!
Hello, young data explorer! In Module 8, we learned how to clean and prepare our data using dplyr. We used verbs like filter(), select(), and mutate(). But what if our data is spread across two or more tables? For example, one table has pupils' names and ages, and another table has their test scores. How do we bring them together?
In this module, we will learn how to join tables. Joining means combining rows or columns from different tables based on a common key (like an ID or name). It's like putting puzzle pieces together to get a complete picture.
We will use the dplyr package again, because it has special functions for joining: inner_join(), left_join(), right_join(), and full_join(). By the end of this module, you will be able to merge different datasets and unlock new insights.
inner_join() to keep only rows that match in both tables.left_join() to keep all rows from the left table and matching rows from the right.right_join() to keep all rows from the right table and matching rows from the left.full_join() to keep all rows from both tables.At Sunshine Primary School, the principal, Mrs. Ade, kept two lists. The first list had pupils' names and their class:
The second list had pupils' names and their favourite subject:
Mrs. Ade wanted a single list with each pupil's class and favourite subject. She needed to join the two lists using the pupils' names as the common key. She used a left_join() to keep all pupils from the first list and add their favourite subjects. For Efe, who wasn't in the first list, she had no class β so it showed NA.
Now, the combined list looked like this:
Mrs. Ade was happy because she could see everything in one place. That is the power of joining tables!
Definition: A join is a way to combine two tables by matching rows based on a common column (called a key).
Why it is important: Data is often stored in separate tables to avoid duplication. Joins let us bring them back together for analysis.
Simple explanation: It's like joining two pieces of a puzzle that have matching edges.
Real-life example: A shop has a list of products (product ID, name) and a list of sales (product ID, quantity). Joining them gives each product's sales.
School example: A teacher has a table of pupils' names and a table of their test scores. Joining gives a complete view.
Home example: You have a list of family members and a list of their birthdays. Joining them gives everyone's birthday.
Nigerian example: A government agency has a table of states and a table of governors. Joining gives the governor for each state.
Illustration:
Table A (Pupils) Table B (Favourite Subject) +------+-------+ +------+----------+ | name | class | | name | subject | +------+-------+ +------+----------+ | Ade | 5 | | Ade | Maths | | Bola | 6 | | Bola | English | +------+-------+ +------+----------+ Join on name: +------+-------+----------+ | name | class | subject | +------+-------+----------+ | Ade | 5 | Maths | | Bola | 6 | English | +------+-------+----------+
Mini summary: Joins combine tables using a common column (key).
Definition: A key is a column that appears in both tables and is used to match rows.
Why it is important: Without a key, we cannot know which rows belong together.
Simple explanation: The key is like the "glue" that sticks the two tables together.
Real-life example: In school, your admission number is a key that links your name to your grades.
School example: The "student ID" column is a key.
Home example: Your phone number can be a key to find your address in a phone book.
Nigerian example: The "state code" is a key to join state names with their capital cities.
Illustration:
Table A Table B +-------+------+ +-------+--------+ | id | name | | id | score | +-------+------+ +-------+--------+ | 1 | Ade | | 1 | 85 | | 2 | Bola | | 2 | 70 | +-------+------+ +-------+--------+ The key is "id". It appears in both tables.
Mini summary: A key is the column that matches rows between tables.
Definition: inner_join() keeps only rows that have a match in both tables.
Why it is important: It gives you only the complete cases where information is available from both sides.
Simple explanation: It's like inviting only the friends who RSVP'd "yes" to your party.
Real-life example: A company wants to see only employees who have a department assigned.
School example: You want to see only pupils who have both a name and a score.
Home example: You want to see only family members who have both a birthday and a favourite meal.
Nigerian example: You want to see only states that have both a governor and a population figure.
Code:
library(dplyr)
# Table A: pupils
pupils <- data.frame(
name = c("Ade", "Bola", "Chidi"),
class = c(5, 6, 5)
)
# Table B: subjects
subjects <- data.frame(
name = c("Ade", "Bola", "Dami"),
subject = c("Maths", "English", "Art")
)
# Inner join
pupils %>%
inner_join(subjects, by = "name")
# Result: only Ade and Bola (common names)
Result: name class subject Ade 5 Maths Bola 6 English
Mini summary: inner_join() keeps only rows that match in both tables.
Definition: left_join() keeps all rows from the left table and adds matching rows from the right. If no match, it fills with NA.
Why it is important: Often we want to keep all our main data and just "decorate" it with extra information.
Simple explanation: It's like starting with your full guest list and only adding RSVP notes for those who replied.
Real-life example: You have a list of all products and you want to add their current stock from another table.
School example: You have a list of all pupils and you want to add their scores (even if some scores are missing).
Home example: You have a list of all your chores and you want to add the time each took (some may be empty).
Nigerian example: You have a list of all states and you want to add their populations (some may be missing).
Code:
# Left join
pupils %>%
left_join(subjects, by = "name")
# Result: all pupils, with subjects added.
Result: name class subject Ade 5 Maths Bola 6 English Chidi 5 NA (no match in subjects)
Mini summary: left_join() keeps all rows from the left table.
Definition: right_join() keeps all rows from the right table and adds matching rows from the left. If no match, it fills with NA.
Why it is important: It's useful when the right table is the main one you care about.
Simple explanation: It's like starting with the RSVP list and adding guest details from the main list.
Real-life example: You have a list of orders (right) and a product list (left). You want to see all orders with product details.
School example: You have a list of scores (right) and want to add pupil names (left).
Home example: You have a list of expenses (right) and want to add category descriptions (left).
Nigerian example: You have a list of governors (right) and want to add state names (left).
Code:
# Right join
pupils %>%
right_join(subjects, by = "name")
# Result: all subjects, with pupils added.
Result: name class subject Ade 5 Maths Bola 6 English Dami NA Art (no match in pupils)
Mini summary: right_join() keeps all rows from the right table.
Definition: full_join() keeps all rows from both tables, matching where possible and filling with NA where no match.
Why it is important: It gives you the complete union of both datasets.
Simple explanation: It's like inviting everyone from two guest lists and merging them.
Real-life example: You have two customer lists and you want to combine them into one master list.
School example: You have two class lists and want to combine them.
Home example: You have a list of your toys and a list of your sibling's toys β combine them.
Nigerian example: You have a list of states and a list of capital cities β combine them.
Code:
# Full join
pupils %>%
full_join(subjects, by = "name")
# Result: all names from both tables.
Result: name class subject Ade 5 Maths Bola 6 English Chidi 5 NA Dami NA Art
Mini summary: full_join() keeps all rows from both tables.
Sometimes the key columns have different names in each table. You can specify by = c("col1" = "col2").
Example:
# Table A has column "student_id", Table B has column "id"
# We want to join on student_id = id
tableA %>%
left_join(tableB, by = c("student_id" = "id"))
Mini summary: Use by = c("name_in_A" = "name_in_B") when key names differ.
If there are duplicate keys in one table, the join will produce all combinations (cartesian product). This can lead to more rows than expected.
Why it is important: Duplicates can cause your results to be overcounted. Always check for duplicates before joining.
Solution: Use distinct() to remove duplicates from the key column.
tableA %>% distinct(key, .keep_all = TRUE)
Mini summary: Check for duplicate keys; remove them with distinct() if needed.
You can join on two or more columns by providing a vector of column names: by = c("col1", "col2").
Example: Join on "first_name" and "last_name" together to avoid confusion.
tableA %>%
left_join(tableB, by = c("first_name", "last_name"))
Mini summary: Use multiple columns as a composite key when one column is not enough.
Let's join a table of states with a table of governors.
states <- data.frame(
state = c("Lagos", "Kano", "Oyo", "Rivers"),
region = c("SW", "NW", "SW", "SS")
)
governors <- data.frame(
state = c("Lagos", "Kano", "Rivers", "Abia"),
governor = c("Sanwo-Olu", "Yusuf", "Fubara", "Ottu")
)
# Left join to keep all states
states %>%
left_join(governors, by = "state")
# Result: Lagos, Kano, Oyo (NA governor), Rivers
Result: state region governor Lagos SW Sanwo-Olu Kano NW Yusuf Oyo SW NA Rivers SS Fubara
Mini summary: Joins work perfectly with Nigerian data too.
Joins can be visualised as Venn diagrams:
inner_join: only the overlapping part. left_join: all of left circle + overlap. right_join: all of right circle + overlap. full_join: entire union of both circles.
Venn diagram concept: +----------+ +----------+ | left | | right | | table | | table | +----------+ +----------+ Overlap = matching rows
Mini summary: Joins are like Venn diagrams for data.
How to choose:
inner_join when you only want rows that exist in both tables.left_join when you want all rows from the main table and add extra info.right_join when you want all rows from the secondary table.full_join when you want to combine all rows.Mini summary: Choose the join that keeps the rows you need for your analysis.
You can chain multiple joins using pipes.
tableA %>%
left_join(tableB, by = "key1") %>%
left_join(tableC, by = "key2")
Mini summary: You can join many tables one after another.
After joining, always check:
NA values β are they okay?Use str() or glimpse() to see the structure.
result %>% glimpse()
Mini summary: Always inspect your joined data to make sure it's correct.
We have learned four main joins:
inner_join() β only matches.left_join() β all left, matches right.right_join() β all right, matches left.full_join() β all rows from both.And we learned about keys, handling duplicates, and checking results.
NA.dplyr: library(dplyr)left_join(left_table, right_table, by = "key").NA values and row count.semi_join() and anti_join() that filter rather than combine.left_join() is the most commonly used join.NA values after a join.distinct() to remove duplicate keys if necessary.glimpse() or head().by argument when key names are different.NA values mean no match.by explicitly, even if the key names are the same.distinct() before joining.left_join() unless you have a specific reason for another join.
inner_join: [overlap only]
+------+------+
| | |
| A | B |
| | |
+------+------+
left_join: [all of A + overlap]
+------+------+
| | |
| A | B |
| | |
+------+------+
right_join: [all of B + overlap]
+------+------+
| | |
| A | B |
| | |
+------+------+
full_join: [all of A and all of B]
+------+------+
| | |
| A | B |
| | |
+------+------+
Left Table Right Table +----+------+ +----+--------+ | id | name | | id | score | +----+------+ +----+--------+ | 1 | Ade | | 1 | 85 | | 2 | Bola | | 3 | 90 | +----+------+ +----+--------+ left_join on id: +----+------+--------+ | id | name | score | +----+------+--------+ | 1 | Ade | 85 | | 2 | Bola | NA | +----+------+--------+
| Join | Rows Kept | When to Use |
|---|---|---|
| inner_join | Only matches | Complete cases only |
| left_join | All left + matches | Keep main table |
| right_join | All right + matches | Keep secondary table |
| full_join | All rows from both | Union of data |
In this module, we learned how to join tables using dplyr. We discovered that a key is a column that links two tables. We explored four types of joins: inner_join() (only matching rows), left_join() (all rows from the left table), right_join() (all rows from the right table), and full_join() (all rows from both tables). We also learned how to handle duplicate keys, join on different column names, and check our results. Joins are a powerful way to combine information from multiple sources, which is essential in data analysis.
by = c("name_in_A" = "name_in_B").distinct().NA mean after a join?| Join | Description |
|---|---|
| 1. inner_join | A. All rows from left + matches |
| 2. left_join | B. All rows from both |
| 3. right_join | C. Only matching rows |
| 4. full_join | D. All rows from right + matches |
Answers: 1-C, 2-A, 3-D, 4-B
In groups, create two small datasets (e.g., pupils and their classes, and pupils and their favourite subjects). Use all four types of joins and discuss the differences in the results. Present your findings to the class.
Using the mtcars dataset and a small custom dataset (e.g., car names and their country of origin), perform a left join and a full join. Write down the differences in the resulting tables.
Title: "Combining School Data"
You have two tables: students (id, name, class) and grades (id, subject, score). Perform joins to create a complete dataset that shows each student's name, class, subject, and score. Then, summarise the average score by class using group_by() and summarise().
Using the nycflights13 package (flights data), join the flights table with the airlines table to get the airline names for each flight. Then, join with the airports table to get the destination airport names. Write the code and show the first few rows.
Find two real datasets online (e.g., population and GDP for Nigerian states). Perform a join to combine them, then create a visualisation (bar chart) showing the relationship. Write a short report explaining your steps.
Fill-in-the-Blank: 1. join, 2. key, 3. inner_join, 4. left_join, 5. right_join, 6. full_join, 7. duplicate, 8. NA, 9. distinct, 10. by.
True/False: 1T, 2F, 3F, 4T, 5T, 6F, 7T, 8F, 9T, 10T.
Multiple Choice: 1C, 2A, 3B, 4C, 5A, 6B, 7B, 8B, 9A, 10B, 11B, 12B, 13B, 14A, 15A.
NA values.In Module 10, we will learn about tidy data and the tidyr package. We will reshape data, convert between wide and long formats, and use functions like pivot_longer() and pivot_wider(). This will help us prepare data for analysis and visualisation. Practise joining tables and reflect on how combining data makes it more powerful.
Hello, young data explorer! In Module 9, we learned how to join tables to bring data together. Now, we will learn how to reshape data so it is in the best format for analysis. Sometimes data is wide (many columns) and sometimes it is long (many rows). We need to choose the right shape for our work.
This is called tidy data. Tidy data has a simple rule: each variable is a column, each observation is a row, and each value is a cell. But real data is often messy. We will use the tidyr package to clean and reshape data.
By the end of this module, you will be able to convert data between wide and long formats, split columns, and handle missing values. You will be a data shaper!
tidyr package.pivot_longer() to convert wide data to long data.pivot_wider() to convert long data to wide data.separate() to split one column into multiple columns.unite() to combine multiple columns into one.drop_na() to remove rows with missing values.fill() to fill missing values with the previous value.At Sunshine Primary School, the teachers recorded test scores in a wide format:
name Math English Science Ade 85 90 78 Bola 70 65 80 Chidi 95 88 92
This is easy to read, but it is hard to analyse because subjects are in columns. If we want to calculate the average score per subject, we have to do extra work. The data is not tidy because the subjects should be in a column, not spread across columns.
So, the teacher used pivot_longer() to reshape it into a long format:
name subject score Ade Math 85 Ade English 90 Ade Science 78 Bola Math 70 ...
Now, each row is one observation (a pupil and a subject). It is easy to calculate averages, make graphs, and analyse. The teacher was happy because the data was now tidy!
Definition: Tidy data is data that follows three rules: (1) each variable is a column, (2) each observation is a row, and (3) each value is a cell.
Why it is important: Tidy data is easy to analyse, visualise, and model. Most R tools expect data in a tidy format.
Simple explanation: It's like organising your bookshelf: every book has a title (variable), and each book is on its own (observation).
Real-life example: A grocery list where each item is a row and columns are item, quantity, and price.
School example: A class register where each pupil is a row and columns are name, age, and grade.
Home example: A list of chores where each chore is a row and columns are chore, person, and time.
Nigerian example: A government dataset where each state is a row and columns are state, population, and region.
Illustration:
Tidy data: +------+-------+--------+ | name | class | score | +------+-------+--------+ | Ade | 5 | 85 | | Bola | 6 | 70 | +------+-------+--------+ Each variable (name, class, score) is a column. Each observation (pupil) is a row.
Mini summary: Tidy data has one variable per column, one observation per row.
Definition: Wide data has many columns, where each category is a separate column. Long data has fewer columns, but more rows, with one column for category names and one for values.
Why it is important: Different analyses require different shapes. Graphs often need long data, while reports often need wide data.
Simple explanation: Wide is like a spreadsheet with many columns; long is like a database with many rows.
Real-life example: Wide: monthly sales for each month as a column. Long: a column for month and a column for sales.
School example: Wide: scores for each subject as columns. Long: subject and score columns.
Home example: Wide: each family member has a column for each chore. Long: columns for person, chore, and done?
Nigerian example: Wide: population for each state as separate columns. Long: state and population columns.
Illustration:
Wide (not tidy): +------+------+----------+ | name | Math | English | +------+------+----------+ | Ade | 85 | 90 | +------+------+----------+ Long (tidy): +------+---------+-------+ | name | subject | score | +------+---------+-------+ | Ade | Math | 85 | | Ade | English | 90 | +------+---------+-------+
Mini summary: Wide has many columns; long has many rows. Tidy data is often long.
Definition: tidyr is an R package that helps you reshape and organise data to make it tidy.
Why it is important: It provides functions like pivot_longer(), pivot_wider(), separate(), and unite().
Simple explanation: Think of tidyr as a clothes folder β it folds your data into the right shape.
Real-life example: A tailor reshapes cloth to make a shirt β tidyr reshapes data for analysis.
School example: You rearrange your school bag to fit everything neatly.
Home example: You fold clothes to fit in the wardrobe.
Nigerian example: A farmer arranges yams in rows and columns for storage.
install.packages("tidyr")
library(tidyr)
Mini summary: tidyr helps you change the shape of your data.
Definition: pivot_longer() collapses multiple columns into two columns: one for the column names (keys) and one for the values.
Why it is important: It makes data tidy by turning column names into a variable.
Simple explanation: It's like taking a table with subjects as columns and stacking them into a single column.
Real-life example: You have sales data for each month in separate columns. You pivot to a month column and sales column.
School example: You have test scores for Math, English, Science as columns. You pivot to subject and score columns.
Home example: You have a list of expenses for each day as columns. You pivot to day and expense columns.
Nigerian example: You have population data for each state as columns. You pivot to state and population.
Code:
# Wide data
scores_wide <- data.frame(
name = c("Ade", "Bola"),
Math = c(85, 70),
English = c(90, 65)
)
# Pivot to long
scores_long <- scores_wide %>%
pivot_longer(cols = c(Math, English),
names_to = "subject",
values_to = "score")
# Result:
# name subject score
# Ade Math 85
# Ade English 90
# Bola Math 70
# Bola English 65
Wide: +------+------+---------+ | name | Math | English | +------+------+---------+ | Ade | 85 | 90 | +------+------+---------+ Long (after pivot_longer): +------+---------+-------+ | name | subject | score | +------+---------+-------+ | Ade | Math | 85 | | Ade | English | 90 | +------+---------+-------+
Mini summary: pivot_longer() makes data longer by stacking columns.
Definition: pivot_wider() spreads a column into multiple columns, making data wider.
Why it is important: Sometimes we need data in a wider format for reporting or presentation.
Simple explanation: It's like taking a column of categories and making them into separate columns.
Real-life example: You have a column for months and a column for sales. You want months as columns.
School example: You have subject and score columns. You want subjects as columns for a report card.
Home example: You have chore and time columns. You want each chore as a column.
Nigerian example: You have state and population columns. You want each state as a column.
Code:
# Long data
scores_long <- data.frame(
name = c("Ade", "Ade", "Bola", "Bola"),
subject = c("Math", "English", "Math", "English"),
score = c(85, 90, 70, 65)
)
# Pivot to wide
scores_wide <- scores_long %>%
pivot_wider(names_from = subject,
values_from = score)
# Result: name, Math, English columns
Long: +------+---------+-------+ | name | subject | score | +------+---------+-------+ | Ade | Math | 85 | | Ade | English | 90 | +------+---------+-------+ Wide (after pivot_wider): +------+------+---------+ | name | Math | English | +------+------+---------+ | Ade | 85 | 90 | +------+------+---------+
Mini summary: pivot_wider() makes data wider by spreading a column.
Definition: separate() splits a single column into multiple columns based on a separator (like a comma or space).
Why it is important: Sometimes a column contains combined information (e.g., "first_name last_name").
Simple explanation: It's like cutting a sandwich into two halves.
Real-life example: Splitting full name into first name and last name.
School example: Splitting "class_roll" into class and roll number.
Home example: Splitting "meal_time" into meal and time.
Nigerian example: Splitting "state_capital" into state and capital.
Code:
# Column: full_name = "Ade Bola"
data <- data.frame(full_name = c("Ade Bola", "Chidi Ngozi"))
data %>%
separate(full_name, into = c("first", "last"), sep = " ")
# Result: first = "Ade", last = "Bola"
Before: +------------+ | full_name | +------------+ | Ade Bola | +------------+ After separate: +-------+------+ | first | last | +-------+------+ | Ade | Bola | +-------+------+
Mini summary: separate() splits a column into two or more columns.
Definition: unite() combines multiple columns into one column.
Why it is important: Sometimes we need to merge information for a unique identifier.
Simple explanation: It's like gluing two pieces of paper together.
Real-life example: Combining first name and last name into full name.
School example: Combining class and roll number into "class_roll".
Home example: Combining meal and time into "meal_time".
Nigerian example: Combining state and region into "state_region".
Code:
data <- data.frame(first = c("Ade", "Chidi"),
last = c("Bola", "Ngozi"))
data %>%
unite(full_name, first, last, sep = " ")
# Result: full_name = "Ade Bola"
Before: +-------+-------+ | first | last | +-------+-------+ | Ade | Bola | +-------+-------+ After unite: +------------+ | full_name | +------------+ | Ade Bola | +------------+
Mini summary: unite() combines columns into one.
Definition: drop_na() removes rows that contain any missing values (NA).
Why it is important: Missing values can cause errors in analysis. Removing them is a quick fix.
Simple explanation: It's like taking out the rotten apples from a basket.
Real-life example: A survey respondent didn't answer a question β you drop that row.
School example: A pupil was absent, so you drop that score row.
Home example: You forgot to record a chore, so you drop that row.
Nigerian example: A state's data is missing, so you drop that row.
Code:
data <- data.frame(name = c("Ade", "Bola", "Chidi"),
score = c(85, NA, 90))
data %>% drop_na()
# Result: only Ade and Chidi (Bola removed)
Before: +-------+-------+ | name | score | +-------+-------+ | Ade | 85 | | Bola | NA | | Chidi | 90 | +-------+-------+ After drop_na(): +-------+-------+ | name | score | +-------+-------+ | Ade | 85 | | Chidi | 90 | +-------+-------+
Mini summary: drop_na() removes rows with missing values.
Definition: fill() fills missing values (NA) with the previous or next value in the column.
Why it is important: Sometimes missing values are just gaps that can be filled logically.
Simple explanation: It's like copying the last known value to fill a blank.
Real-life example: A time series where missing days can be filled with the previous day's value.
School example: If a pupil's class is missing, you can fill with the class from the previous row.
Home example: If you forget to record your expense for a day, you can use the previous day.
Nigerian example: If a state's region is missing, fill with the region from a nearby state.
Code:
data <- data.frame(
day = c(1, 2, 3, 4),
sales = c(100, NA, NA, 150)
)
data %>% fill(sales, .direction = "down")
# Result: 100, 100, 100, 150
Before: +-----+-------+ | day | sales | +-----+-------+ | 1 | 100 | | 2 | NA | | 3 | NA | | 4 | 150 | +-----+-------+ After fill(down): +-----+-------+ | day | sales | +-----+-------+ | 1 | 100 | | 2 | 100 | | 3 | 100 | | 4 | 150 | +-----+-------+
Mini summary: fill() replaces missing values with nearby values.
Let's reshape Nigerian population data from wide to long.
# Wide data: columns for each year
pop_wide <- data.frame(
state = c("Lagos", "Kano", "Oyo"),
X2020 = c(20, 15, 8),
X2021 = c(21, 16, 9)
)
# Pivot to long
pop_long <- pop_wide %>%
pivot_longer(cols = c(X2020, X2021),
names_to = "year",
values_to = "population")
# Result: state, year, population
Before (wide): +-------+------+------+ | state | 2020 | 2021 | +-------+------+------+ | Lagos | 20 | 21 | +-------+------+------+ After (long): +-------+------+------------+ | state | year | population | +-------+------+------------+ | Lagos | 2020 | 20 | | Lagos | 2021 | 21 | +-------+------+------------+
Mini summary: pivot_longer() works on any data, including Nigerian data.
You can chain tidyr functions with dplyr verbs using pipes.
data %>%
pivot_longer(cols = c(Math, English), names_to = "subject", values_to = "score") %>%
filter(score > 70) %>%
arrange(desc(score))
Mini summary: tidyr and dplyr work together beautifully.
When to use wide: For human-readable tables, reports, and some visualisations.
When to use long: For analysis, modelling, and many ggplot2 graphs.
Mini summary: Use wide for presentation, long for analysis.
After pivoting, you might get duplicate rows. Use distinct() to remove them.
data %>% distinct()
Mini summary: Check for duplicates after reshaping.
pivot_longer() β wide to longpivot_wider() β long to wideseparate() β split one columnunite() β combine columnsdrop_na() β remove missing rowsfill() β fill missing valuesTidy data is the foundation of data analysis in R. Once your data is tidy, everything else becomes easier.
tidyr and dplyr.pivot_longer(cols = c(col1, col2), names_to = "key", values_to = "value").head().drop_na() to remove rows.fill() to propagate values.pivot_longer() with multiple columns at once.pivot_longer_spec() for advanced pivoting.complete() to fill missing combinations.pivot_longer() makes data longer; pivot_wider() makes it wider.separate() and unite() for columns.NA with drop_na() or fill().tidyr.pivot_longer().names_to and values_to correctly.pivot_wider() when pivot_longer() is needed.glimpse() to inspect your data after reshaping.Wide: +------+------+------+ | name | Math | Eng | +------+------+------+ | Ade | 85 | 90 | +------+------+------+ pivot_longer: +------+---------+-------+ | name | subject | score | +------+---------+-------+ | Ade | Math | 85 | | Ade | English | 90 | +------+---------+-------+
Long: +------+---------+-------+ | name | subject | score | +------+---------+-------+ | Ade | Math | 85 | | Ade | English | 90 | +------+---------+-------+ pivot_wider: +------+------+---------+ | name | Math | English | +------+------+---------+ | Ade | 85 | 90 | +------+------+---------+
| Feature | Wide | Long |
|---|---|---|
| Number of columns | Many | Few |
| Number of rows | Few | Many |
| Best for | Reporting | Analysis |
| Example | Subjects as columns | Subject column |
In this module, we learned about tidy data and how to reshape data using the tidyr package. We can convert data from wide to long with pivot_longer() and from long to wide with pivot_wider(). We also learned to split columns with separate(), combine columns with unite(), and handle missing values with drop_na() and fill(). Tidy data is essential for analysis and visualisation in R. Remember, the shape of your data matters!
drop_na().fill().| Function | Action |
|---|---|
| 1. pivot_longer | A. Combines columns |
| 2. pivot_wider | B. Splits a column |
| 3. separate | C. Wide to long |
| 4. unite | D. Long to wide |
| 5. drop_na | E. Removes missing rows |
Answers: 1-C, 2-D, 3-B, 4-A, 5-E
In groups, create a wide dataset (e.g., favourite foods with ratings for each family member). Use pivot_longer() to make it long, then pivot_wider() to make it wide again. Discuss the differences.
Use the iris dataset, but modify it to have wide format (one column per species for Sepal.Length). Then use pivot_longer() to restore it to long format.
Title: "Reshaping Nigerian Education Data"
Create a wide dataset of test scores for pupils in different subjects. Use pivot_longer() to tidy it. Then, use group_by() and summarise() to find the average score per subject. Present your findings.
Using the economics dataset (built-in), which is in long format, use pivot_wider() to make it wide with variables as columns. Then use pivot_longer() to return to long.
Find a messy dataset online (e.g., from Kaggle). Use tidyr and dplyr to clean and reshape it into a tidy format. Write a short report on your steps.
Fill-in-the-Blank: 1. Tidy, 2. Wide, 3. Long, 4. pivot_longer, 5. pivot_wider, 6. separate, 7. unite, 8. drop_na, 9. fill, 10. tidyr.
True/False: 1T, 2F, 3F, 4F, 5T, 6T, 7T, 8T, 9F, 10T.
Multiple Choice: 1C, 2B, 3A, 4B, 5B, 6A, 7B, 8C, 9D, 10B, 11B, 12D, 13C, 14A, 15A.
In Module 11, we will learn about data import β reading data from different file types like CSV, Excel, and databases. We will use packages like readr, readxl, and haven. Practise reshaping data with tidyr and reflect on how tidy data makes analysis easier.
Hello, young data explorer! In Module 10, we learned how to reshape data using tidyr. But where does data come from? Usually, data is stored in files on your computer. These files can be in many formats: CSV (comma-separated values), Excel, or even databases.
In this module, we will learn how to import (read) data into R from different file types, and export (write) data from R to files. This is like opening a book (reading) and writing your own book (saving). We will use packages like readr, readxl, and haven to handle different formats.
By the end of this module, you will be able to load any dataset into R and save your results for sharing. You will be a data librarian!
readr package.read_csv() to read CSV files.write_csv() to write CSV files.readxl package.read_excel() to read Excel files.read_delim(), read_table().Chidi was doing a school project on Nigerian states. He had collected data in a CSV file on his computer. He opened R and typed read_csv("states.csv"), but R gave an error: "File not found." Chidi was confused. He realised he needed to tell R exactly where the file was on his computer β the path. He set his working directory to the folder containing the file, and it worked!
Later, he analysed the data and created a summary. He wanted to share his results with his teacher, so he used write_csv() to save the summary as a new CSV file. He emailed it to his teacher. Chidi learned that importing and exporting data is like opening and saving files β it's essential for working with real data.
Definition: Import means reading data from a file into R. Export means writing data from R to a file.
Why it is important: Data lives in files (CSV, Excel, etc.). To analyse it, we must bring it into R. After analysis, we save results for sharing.
Simple explanation: It's like reading a book (import) and writing your own book (export).
Real-life example: A shopkeeper reads a price list from a spreadsheet (import) and writes a sales report (export).
School example: Your teacher reads a class list from a file (import) and saves grades (export).
Home example: You open a recipe from a file (import) and save your grocery list (export).
Nigerian example: A researcher reads population data from a CSV (import) and saves a summary report (export).
Illustration:
File (on computer) ----Import---> R (data frame) R (data frame) ----Export---> File (saved)
Mini summary: Import = read into R; Export = save from R.
Definition: File formats are ways of storing data. Common ones are:
Mini summary: CSV and Excel are most common for beginners.
Definition: The working directory is the folder where R looks for files.
Why it is important: If your file is not in the working directory, R won't find it.
Simple explanation: It's like the "home base" of your R session.
Real-life example: You keep your school books in a specific bag β you know where to find them.
School example: Your teacher tells you to open a file from the "Class Data" folder.
Home example: You keep your toys in a box β you know where to look.
Nigerian example: A researcher keeps all data files in a "Projects" folder.
Code:
getwd() # shows current working directory
setwd("path/to/folder") # changes it
Mini summary: Set your working directory to the folder containing your files.
Definition: read_csv() is a function from the readr package that reads CSV files.
Why it is important: CSV is the most common data format.
Simple explanation: It's like opening a book β R reads the text and turns it into a table.
Real-life example: A scientist reads a CSV of weather data.
School example: A teacher reads a CSV of pupil scores.
Home example: You read a CSV of your family expenses.
Nigerian example: An economist reads a CSV of Nigeria's GDP data.
Code:
install.packages("readr")
library(readr)
data <- read_csv("file.csv")
# Or with a full path:
data <- read_csv("C:/Users/YourName/Downloads/file.csv")
Illustration:
CSV file (text): name,age,score Ade,10,85 Bola,11,70 read_csv() -> R data frame: +------+-----+-------+ | name | age | score | +------+-----+-------+ | Ade | 10 | 85 | | Bola | 11 | 70 | +------+-----+-------+
Mini summary: read_csv() reads CSV files into R.
Definition: write_csv() saves a data frame as a CSV file.
Why it is important: You can save your results and share them.
Simple explanation: It's like writing a story and saving it as a document.
Real-life example: A shop saves a sales summary as CSV.
School example: A teacher saves grades as CSV.
Home example: You save your chore chart as CSV.
Nigerian example: A researcher saves analysed data as CSV.
Code:
write_csv(data, "output.csv") # Saves to working directory.
Mini summary: write_csv() saves data frames as CSV.
Definition: read_excel() from the readxl package reads Excel (.xlsx, .xls) files.
Why it is important: Excel is widely used in business and schools.
Simple explanation: It's like opening a spreadsheet in R.
Real-life example: A business reads an Excel budget file.
School example: A teacher reads an Excel attendance sheet.
Home example: You read an Excel list of family members.
Nigerian example: A government official reads an Excel file of state budgets.
Code:
install.packages("readxl")
library(readxl)
data <- read_excel("file.xlsx")
# Specify sheet if needed:
data <- read_excel("file.xlsx", sheet = "Sheet1")
Mini summary: read_excel() reads Excel files.
To write Excel files, you can use the writexl package or openxlsx.
install.packages("writexl")
library(writexl)
write_xlsx(data, "output.xlsx")
Mini summary: Use write_xlsx() to save as Excel.
Definition: read_delim() reads files with any separator (e.g., tabs, semicolons).
Real-life example: A tab-separated file.
School example: A file with pipe (|) separators.
Code:
data <- read_delim("file.txt", delim = "\t") # tab-separated
data <- read_delim("file.txt", delim = ";") # semicolon
Mini summary: read_delim() handles any delimiter.
Definition: A file path tells R exactly where the file is located.
Why it is important: If you don't give the full path, R looks in the working directory.
Simple explanation: It's like giving an address to find a house.
Types:
Mini summary: Use full paths to avoid errors.
You can read data directly from a URL.
data <- read_csv("https://example.com/data.csv")
Mini summary: read_csv() works with web URLs too.
After analysis, you might want to save your data in different formats.
write_csv(data, "clean_data.csv") write_xlsx(data, "clean_data.xlsx") write_rds(data, "clean_data.rds") # R's native format
Mini summary: Save your data in the format your audience needs.
Definition: Databases are organised collections of data. R can connect to them using packages like DBI and odbc.
For beginners, we focus on files.
Mini summary: Databases are more advanced β we'll start with files.
read_delim() with correct delim.col_names = TRUE or skip.locale() to specify encoding.Mini summary: Most errors are due to file location or format.
Let's say we have a CSV file "nigeria_population.csv" with columns: State, Population, Region.
pop <- read_csv("nigeria_population.csv")
head(pop)
Mini summary: Import Nigerian data easily with read_csv().
We learned to read and write CSV, Excel, and other formats. We also learned about working directories and file paths.
read.csv().head() or glimpse().readr: install.packages("readr").library(readr).read_csv("filename.csv").data <- read_csv("filename.csv").head(data).write_csv(data, "output.csv").googlesheets4 package.readr is part of the tidyverse.read_csv() for CSV files, read_excel() for Excel.head().write_csv() or write_xlsx().read.csv() instead of read_csv() (both work, but read_csv is faster).read_csv() for tab-separated files).readr functions for better performance.col_types argument for large files.File on disk ----read_csv----> R data frame (data.csv) (data)
R data frame ----write_csv----> File on disk (data) (output.csv)
| File Type | Function | Package |
|---|---|---|
| CSV | read_csv() | readr |
| Excel | read_excel() | readxl |
| Tab-separated | read_delim(file, delim = "\t") | readr |
| RDS | readRDS() | base R |
In this module, we learned how to import (read) data from files into R and export (write) data from R to files. We used read_csv() and write_csv() for CSV files, and read_excel() for Excel files. We also learned about the working directory, file paths, and common errors. These skills are essential because most data comes from files, and we need to share our results.
readr.readxl.getwd().setwd("path").read_csv("http://example.com/data.csv").write_csv(data, "filename.csv").read_delim() do?head() do?| Function | File Type |
|---|---|
| 1. read_csv | A. Excel |
| 2. read_excel | B. CSV |
| 3. read_delim | C. Any delimiter |
| 4. write_csv | D. Save as CSV |
Answers: 1-B, 2-A, 3-C, 4-D
read_csv("grades.csv") but get an error. What might be the problem?In groups, create a small dataset (e.g., favourite foods with ratings). Save it as a CSV file. Then, each group member imports it into R and creates a summary. Export the summary as a new CSV file and share it with the group.
Find a CSV file online (e.g., from a data portal). Import it into R, explore it with glimpse() and head(), and then save a filtered version as a new CSV file.
Title: "Importing and Analysing Nigerian Data"
Find a CSV or Excel file about Nigeria (e.g., population, education, or weather). Import it into R, perform some basic analysis (summary, filtering), and export the results as a new file. Write a short report on what you found.
Using the iris dataset, save it as a CSV file using write_csv(). Then, import it back into R and verify it matches the original. Do the same with an Excel file using write_xlsx().
Find a dataset in a format other than CSV (e.g., JSON, XML) and use appropriate R packages (e.g., jsonlite) to import it. Write a brief explanation of the process.
Fill-in-the-Blank: 1. Import, 2. Export, 3. CSV, 4. readr, 5. readxl, 6. working directory, 7. setwd, 8. write_csv, 9. read_excel, 10. file path.
True/False: 1F, 2T, 3F, 4T, 5T, 6F, 7T, 8T, 9F, 10T.
Multiple Choice: 1B, 2B, 3B, 4B, 5B, 6A, 7B, 8B, 9A, 10C, 11C, 12B, 13B, 14B, 15C.
read_csv().read_excel().head() after import.write_csv() to save your results.In Module 12, we will learn about data cleaning β handling missing values, outliers, and inconsistencies in detail. We will combine skills from Modules 8-11 to create a complete data cleaning workflow. Practise importing and exporting data, and think about how you would clean messy data.
Hello, young data explorer! In Module 11, we learned how to import data from files. But what if the data we import is messy? It might have missing values, wrong data types, duplicates, or inconsistent spellings. We need to clean it before we can analyse it.
Data cleaning is like washing fruits before eating them β you want them to be fresh and ready. In this module, we will learn how to handle missing values, fix data types, remove duplicates, and standardise text. We will use tools from dplyr, tidyr, and base R.
By the end of this module, you will be able to take any messy dataset and make it clean and ready for analysis. You will be a data cleaner!
na.omit(), drop_na(), and replace_na().as.numeric(), as.character(), etc.distinct().mutate() to create corrected columns.Chidi was excited to analyse data from his school's sports day. He imported the data, but it was a mess! Some scores were missing, some names had extra spaces, and one pupil's name was spelled as "Ade" in some rows and "AdE" in others. The ages were stored as text, not numbers. He couldn't make a graph because R didn't understand the data.
Chidi decided to clean the data. He fixed the spellings, converted ages to numbers, filled missing scores with the average, and removed duplicates. Now, the data was perfect. He could analyse it and create beautiful graphs. Chidi learned that cleaning data is the first and most important step in any analysis.
Definition: Data cleaning is the process of fixing errors, handling missing values, and standardising data to make it ready for analysis.
Why it is important: Dirty data leads to wrong conclusions. Cleaning ensures accuracy.
Simple explanation: It's like tidying your room β you put everything in its place.
Real-life example: A shopkeeper corrects typos in product names.
School example: Your teacher checks that all names are spelled correctly.
Home example: You sort your clothes and remove those with holes.
Nigerian example: A researcher cleans census data to remove errors.
Illustration:
Dirty Data --> Clean Data --> Ready for Analysis
(errors, (fixed,
missing, complete,
duplicates) consistent)
Mini summary: Data cleaning fixes problems so your analysis is correct.
Here are some common problems:
Mini summary: We will learn how to fix each of these.
Definition: drop_na() removes rows that have any missing values.
Why it is important: It's a quick way to get complete cases.
Simple explanation: It's like throwing away rotten apples.
Real-life example: A survey where some questions were not answered β you remove those responses.
School example: A pupil was absent β you remove that row.
Home example: You forgot to record a chore β you remove that entry.
Nigerian example: A state's data is missing β you remove that state.
library(tidyr) data %>% drop_na() # removes rows with any NA data %>% drop_na(score) # removes rows where score is NA
Mini summary: drop_na() removes rows with missing values.
Definition: replace_na() replaces missing values with a specified value.
Why it is important: Sometimes we want to fill in missing values instead of removing rows.
Simple explanation: It's like drawing a smile on a blank face.
Real-life example: You replace a missing score with the average.
School example: You fill absent marks with 0.
Home example: You replace a missing chore time with "0 minutes".
Nigerian example: You replace missing population with the national average.
data %>% replace_na(list(score = 0, age = 10)) # Replaces NA in score with 0, in age with 10.
Mini summary: replace_na() fills missing values.
Definition: na.omit() is a base R function that removes rows with NA.
Why it is important: It's a quick alternative to drop_na().
clean_data <- na.omit(data)
Mini summary: na.omit() removes rows with missing values.
Definition: as.numeric() converts text to numbers.
Why it is important: Numbers stored as text cannot be used in calculations.
Simple explanation: It's like changing a word "five" to the number 5.
Real-life example: A column "age" might be stored as "10" (text) β we need it as number.
School example: Scores stored as "85" (text) β we need numbers to calculate averages.
Home example: Time stored as "30" (text) β we need numbers to sum.
Nigerian example: Population stored as "20 million" β we need to extract numbers.
data %>% mutate(score = as.numeric(score)) # Converts score column to numeric.
Mini summary: as.numeric() turns text into numbers.
Definition: as.character() converts to text; as.factor() converts to categorical data.
Why it is important: Sometimes we need text for labels, or factors for grouping.
Real-life example: "Male"/"Female" should be a factor.
School example: "Class" (5,6) should be a factor, not a number.
Home example: "Chore" names should be text.
Nigerian example: "State" should be a factor.
data %>% mutate(class = as.factor(class))
Mini summary: Use as.character() and as.factor() to fix types.
Definition: distinct() removes duplicate rows.
Why it is important: Duplicates can overcount and skew results.
Simple explanation: It's like deleting repeated entries in a list.
Real-life example: A customer appears twice in a mailing list β remove one.
School example: A pupil's name appears twice β remove duplicate.
Home example: You wrote the same chore twice β remove one.
Nigerian example: A state appears twice in a dataset β remove duplicate.
data %>% distinct() # removes all duplicate rows data %>% distinct(name, .keep_all = TRUE) # keep first occurrence
Mini summary: distinct() removes duplicate rows.
Definition: tolower() converts text to lowercase; toupper() to uppercase; trimws() removes extra spaces.
Why it is important: Inconsistent text prevents matching and grouping.
Simple explanation: It's like making sure everyone writes their name the same way.
Real-life example: "Lagos" and "lagos" should be the same β convert to "lagos" or "Lagos".
School example: "Ade" and "aDe" should be "Ade".
Home example: "Chore" and "chore " β remove spaces.
Nigerian example: "Abuja" and "abuja" β standardise.
data %>% mutate(name = tolower(name)) # all lowercase data %>% mutate(name = trimws(name)) # remove leading/trailing spaces
Mini summary: Use tolower(), toupper(), and trimws() to clean text.
Definition: Outliers are values that are very different from the rest.
Why it is important: They can skew averages and graphs.
Simple explanation: A very tall person in a class is an outlier.
How to handle: You can remove them or cap them at a threshold.
# Remove values above 100 or below 0
data %>% filter(score >= 0 & score <= 100)
# Or use quantiles
data %>% filter(score > quantile(score, 0.25) - 1.5*IQR(score) &
score < quantile(score, 0.75) + 1.5*IQR(score))
Mini summary: Outliers are extreme values β handle with care.
We can chain all cleaning steps together:
clean_data <- raw_data %>%
drop_na() %>%
distinct() %>%
mutate(name = tolower(trimws(name))) %>%
mutate(age = as.numeric(age)) %>%
filter(age > 0 & age < 120) # reasonable age range
Mini summary: A pipeline makes cleaning reproducible and clear.
Let's clean a messy dataset of Nigerian states.
# Messy data
messy <- data.frame(
state = c("Lagos", "Lagos ", "lagos", "Kano", "Kano"),
population = c("20", "21", NA, "15", "15"),
region = c("SW", "SW", "SW", "NW", "NW")
)
clean <- messy %>%
distinct() %>%
mutate(state = tolower(trimws(state))) %>%
mutate(population = as.numeric(population)) %>%
drop_na()
# Result: one row per state, with numeric population, no NAs.
Before: state population region Lagos 20 SW Lagos 21 SW lagos NA SW Kano 15 NW Kano 15 NW After cleaning: state population region lagos 20 SW lagos 21 SW (if distinct on all columns, duplicates removed) kano 15 NW
Mini summary: Cleaning makes Nigerian data usable.
Definition: The janitor package has helpers for cleaning column names and more.
install.packages("janitor")
library(janitor)
data %>% clean_names() # makes column names consistent
Mini summary: janitor is a helpful addition.
Always check:
Use glimpse(), summary(), and head().
Mini summary: Validate to ensure cleaning worked.
We learned to handle missing values, fix types, remove duplicates, standardise text, and manage outliers. Cleaning is the foundation of good analysis.
is.na()).str()).distinct()).glimpse().janitor package has a function get_dupes() to find duplicates.skimr to get a quick summary of data quality.
Raw Data
|
V
drop_na() ----> remove missing
|
V
distinct() ----> remove duplicates
|
V
mutate() ----> fix types & text
|
V
Clean Data
| Function | Action | When to Use |
|---|---|---|
| drop_na() | Removes rows | If few missing values |
| replace_na() | Fills with value | To keep all rows |
| na.omit() | Removes rows | Base R alternative |
In this module, we learned how to clean dirty data. We handled missing values with drop_na() and replace_na(), fixed data types with as.numeric(), removed duplicates with distinct(), and standardised text with tolower() and trimws(). We also learned about outliers and validation. Clean data is essential for accurate analysis and visualisation.
drop_na().replace_na().distinct().as.numeric(), as.character(), etc.trimws() do?tolower() do?janitor package?| Problem | Solution |
|---|---|
| 1. Missing values | A. distinct() |
| 2. Duplicates | B. as.numeric() |
| 3. Text as numbers | C. drop_na() |
| 4. Extra spaces | D. trimws() |
| 5. Inconsistent case | E. tolower() |
Answers: 1-C, 2-A, 3-B, 4-D, 5-E
In groups, create a dirty dataset (with missing values, duplicates, inconsistent text). Exchange datasets with another group and clean them using R. Compare your cleaning approaches.
Use the airquality dataset (built-in). It has missing values. Clean it by removing rows with NA and converting month to a factor. Save the cleaned dataset.
Title: "Clean and Analyse Nigerian School Data"
You are given a messy dataset of Nigerian school attendance with missing values, duplicates, and inconsistent state names. Clean the data and then summarise attendance by state. Write a short report.
Using the diamonds dataset (built-in), perform cleaning: check for missing values, duplicates, and outliers (price, carat). Then, create a summary table of cleaned data.
Find a real, messy dataset online (e.g., from Kaggle). Apply a complete cleaning pipeline using dplyr and tidyr. Document each step and explain why you made each choice.
Fill-in-the-Blank: 1. Cleaning, 2. drop_na, 3. replace_na, 4. distinct, 5. as.numeric, 6. trimws, 7. tolower, 8. outlier, 9. Validation, 10. duplicate.
True/False: 1F, 2T, 3T, 4T, 5F, 6T, 7F, 8T, 9F, 10T.
Multiple Choice: 1B, 2B, 3D, 4B, 5A, 6C, 7B, 8D, 9B, 10B, 11C, 12D, 13A, 14B, 15B.
drop_na() or replace_na().as.numeric() etc.distinct().tolower() and trimws().In Module 13, we will learn about exploratory data analysis (EDA) β using summary statistics and visualisations to understand your data. We will combine all our skills: importing, cleaning, wrangling, and visualising. Practise cleaning different datasets to prepare.
Hello, young data explorer! In Module 12, we learned how to clean our data. Now that our data is clean, it's time to explore it. Exploratory Data Analysis (EDA) means looking at the data from many angles to understand its patterns, relationships, and surprises.
EDA is like being a detective. You ask questions, make plots, calculate summaries, and look for clues. The goal is to understand the data before you do any formal analysis. In this module, we will use summary statistics, visualisations, and grouping to explore data.
By the end of this module, you will be able to explore any dataset and tell its story. You will be a data detective!
group_by() and summarise() to compare groups.ggplot2 to create exploratory plots.Chidi noticed that snacks in his school canteen were disappearing quickly. He wanted to find out which snacks were most popular and when they were bought. He collected data on snack sales: snack name, price, quantity sold, and time of day.
He imported the data, cleaned it, and then started to explore. He calculated the average quantity sold per snack, made a bar chart to compare popularity, and used a line chart to see sales over time. He discovered that meat pies were the most popular, and sales peaked during lunchtime. He presented his findings to the canteen manager, who used the information to stock more meat pies.
Chidi had performed Exploratory Data Analysis!
Definition: EDA is the process of exploring data to understand its main characteristics, patterns, and relationships.
Why it is important: EDA helps you know your data before you do any formal analysis. It guides your next steps.
Simple explanation: It's like exploring a new playground β you check out all the equipment, find the best spots, and see what's fun.
Real-life example: A detective gathers clues at a crime scene.
School example: A teacher looks at test scores to see which topics need more review.
Home example: You look at your piggy bank to see how much money you have and where it came from.
Nigerian example: A researcher explores census data to understand population patterns.
Illustration:
Data --> Ask Questions --> Plot/Summarise --> Find Patterns --> Tell Story
Mini summary: EDA helps you understand your data and find interesting insights.
Definition: glimpse() shows a summary of the data structure; head() shows the first few rows.
Why it is important: They give you a quick overview of the data.
Simple explanation: It's like peeking into a box to see what's inside.
Real-life example: You open a book and read the first page.
School example: You look at the class register to see who is in your class.
Home example: You check the pantry to see what food you have.
Nigerian example: You view the first rows of a dataset of Nigerian states.
library(dplyr) glimpse(data) head(data, 10) # first 10 rows
Mini summary: Always start with glimpse() and head().
Definition: summary() gives basic statistics for each column: min, max, mean, median, quartiles.
Why it is important: It tells you about the center and spread of your data.
Simple explanation: It's like a report card for each column.
Real-life example: A teacher summarises test scores: average, highest, lowest.
School example: You summarise your weekly allowance.
Home example: You summarise the time you spend on homework each day.
Nigerian example: You summarise population data to see average state population.
summary(data) # For a specific column: summary(data$score)
Example output for score:
Min. 1st Qu. Median Mean 3rd Qu. Max.
65.00 70.00 85.00 82.75 90.00 95.00
Mini summary: summary() gives a numeric snapshot of your data.
Definition:
Why they are important: They describe the center of the data.
Simple explanation: Mean is like sharing candies equally; median is the middle candy; mode is the most common candy.
Real-life example: A shop uses mean to find average sales.
School example: Teacher calculates mean test score.
Home example: You find the median age in your family.
Nigerian example: You find the mean population of states.
mean(data$score, na.rm = TRUE) median(data$score, na.rm = TRUE) # For mode, we can use table(): table(data$favourite_food) # shows frequency
Mini summary: Mean, median, and mode tell you about the typical value.
Definition:
Why they are important: They measure spread β how much data varies.
Simple explanation: Range is the distance between the smallest and largest; standard deviation is how much scores typically differ from the average.
Real-life example: A teacher sees if scores are close together or spread out.
School example: You compare the spread of test scores in two classes.
Home example: You check how much your daily reading time varies.
Nigerian example: You check the variability of rainfall across states.
range(data$score) var(data$score, na.rm = TRUE) sd(data$score, na.rm = TRUE)
Mini summary: Range, variance, and SD measure how spread out data is.
Definition: A histogram is a bar chart that shows the frequency of values in bins (intervals).
Why it is important: It shows the shape of the distribution (e.g., bell-shaped, skewed).
Simple explanation: It's like sorting your toys into boxes by size.
Real-life example: A teacher looks at a histogram of test scores to see how many students got each grade range.
School example: You see how many friends live in each neighbourhood.
Home example: You make a histogram of the time you spend on different activities.
Nigerian example: A histogram of ages in a community.
library(ggplot2)
ggplot(data, aes(x = score)) +
geom_histogram(binwidth = 5, fill = "blue", color = "black") +
labs(title = "Histogram of Scores")
Histogram concept:
Frequency
8 | ###
6 | ### ###
4 | ### ### ###
2 | ### ### ### ###
0 |___###___###___###___###____
60-64 65-69 70-74 75-79
Mini summary: Histograms show the distribution of a single variable.
Definition: A boxplot shows the median, quartiles, and outliers of a variable.
Why it is important: It quickly shows the spread and identifies outliers.
Simple explanation: It's like a picture of a box with whiskers.
Real-life example: A scientist uses a boxplot to show temperature variation.
School example: You compare test scores across classes with boxplots.
Home example: You compare the amount of pocket money you and your friend get.
Nigerian example: A boxplot of salaries in different states.
ggplot(data, aes(x = "", y = score)) +
geom_boxplot(fill = "orange") +
labs(title = "Boxplot of Scores", y = "Score")
Boxplot concept: +-----+ outlier (o) | | | +--+--+ (box = IQR) | | | | | +--+--+ | | +-----+ (whiskers extend to min/max within 1.5*IQR)
Mini summary: Boxplots show spread and outliers.
Definition: A scatter plot shows the relationship between two numeric variables.
Why it is important: It reveals correlations (positive, negative, or none).
Simple explanation: It's like plotting points on a map to see if they form a pattern.
Real-life example: A doctor plots height vs weight to see if they are related.
School example: You plot study time vs test score to see if more study leads to higher scores.
Home example: You plot age vs screen time.
Nigerian example: A scatter plot of education level vs income.
ggplot(data, aes(x = study_time, y = score)) +
geom_point() +
labs(title = "Study Time vs Score")
Scatter plot concept:
Score
100 | .
80 | . .
60 | . .
40 | .
20 |.
0 |___._.___.___.___.___.___
0 20 40 60 80 100
Study Time
Mini summary: Scatter plots show relationships between variables.
Definition: Use group_by() and summarise() to get statistics for each group.
Why it is important: It lets you compare groups (e.g., boys vs girls).
Simple explanation: It's like sorting your toys and then counting how many in each group.
Real-life example: A store compares sales by product category.
School example: A teacher compares average scores by class.
Home example: You compare the time you spend on school vs play.
Nigerian example: You compare average income by region.
data %>%
group_by(class) %>%
summarise(avg_score = mean(score),
median_score = median(score),
n = n())
Mini summary: Grouped summaries reveal differences between groups.
Definition: Faceting creates separate plots for each group in a single figure.
Why it is important: It makes comparisons easy.
Simple explanation: It's like having a separate page for each group.
ggplot(data, aes(x = score)) +
geom_histogram() +
facet_wrap(~ class)
Mini summary: Faceting helps compare groups visually.
Definition: Correlation measures the strength and direction of a relationship between two variables.
Why it is important: It tells you if variables are related.
Simple explanation: It's like seeing if two friends always walk together.
Real-life example: Height and weight have a positive correlation.
School example: Study time and test scores often have a positive correlation.
Home example: The amount of exercise and health might have a positive correlation.
Nigerian example: Education and income have a positive correlation.
cor(data$study_time, data$score, use = "complete.obs")
Mini summary: Correlation shows if variables change together.
Let's explore a dataset of Nigerian states.
# Example: population and region
data %>%
group_by(region) %>%
summarise(avg_pop = mean(population),
sd_pop = sd(population)) %>%
ggplot(aes(x = region, y = avg_pop)) +
geom_bar(stat = "identity")
Mini summary: Apply EDA to any Nigerian data.
Always ask questions like:
Mini summary: Good questions lead to good explorations.
Keep a record of your findings. Use an R Markdown or a script with comments.
Mini summary: Document your EDA so others can follow.
EDA is a journey of discovery. You use summary stats, plots, and grouping to understand your data. It's the foundation of all analysis.
glimpse() and head().summary().GGally::ggpairs() to create a matrix of plots.Data --> Clean --> Explore --> Summarise --> Visualise --> Insights
Frequency
8 | ###
6 | ### ###
4 | ### ### ###
2 | ### ### ### ###
0 |___###___###___###___###____
60-64 65-69 70-74 75-79
| Statistic | What it shows | Function |
|---|---|---|
| Mean | Average | mean() |
| Median | Middle | median() |
| Range | Min to Max | range() |
| SD | Spread | sd() |
In this module, we learned how to explore data using Exploratory Data Analysis (EDA). We used summary statistics (mean, median, range, SD), histograms, boxplots, scatter plots, and grouped summaries. We asked questions and looked for patterns, outliers, and relationships. EDA is the first and most important step in any data analysis.
glimpse() do?summary() do?summary() do?glimpse() do?| Concept | Description |
|---|---|
| 1. Mean | A. Middle value |
| 2. Median | B. Average |
| 3. Histogram | C. Shows distribution |
| 4. Boxplot | D. Shows outliers |
| 5. Scatter plot | E. Shows relationship |
Answers: 1-B, 2-A, 3-C, 4-D, 5-E
In groups, choose a dataset (e.g., iris). Perform EDA: summary stats, histograms, boxplots, scatter plots, and grouped summaries. Present your findings to the class.
Use the mtcars dataset. Perform EDA to understand the relationship between horsepower (hp) and fuel efficiency (mpg). Write a short paragraph on your findings.
Title: "Explore Nigerian States"
Find a dataset of Nigerian states (population, region, etc.). Perform a complete EDA: summary stats, histograms, boxplots, and grouped summaries. Write a report with your findings and include visualisations.
Using the diamonds dataset, perform EDA: explore price, carat, cut, and their relationships. Create at least 5 visualisations and summarise your findings.
Find a real dataset online (e.g., Kaggle). Perform a complete EDA and write a report. Include at least 3 different types of plots and grouped summaries. Present your findings in a clear and engaging way.
Fill-in-the-Blank: 1. EDA, 2. mean, 3. median, 4. histogram, 5. boxplot, 6. scatter plot, 7. Correlation, 8. group_by, 9. range, 10. head.
True/False: 1F, 2F, 3T, 4T, 5T, 6F, 7T, 8T, 9T, 10F.
Multiple Choice: 1A, 2B, 3B, 4B, 5C, 6A, 7B, 8B, 9A, 10A, 11A, 12B, 13A, 14A, 15B.
In Module 14, we will learn about statistical testing β how to make decisions based on data. We will use tests like t-test, ANOVA, and chi-square to see if differences are real or due to chance. Practise EDA on different datasets to get comfortable with exploring.
Hello, young data explorer! In Module 13, we learned how to explore data and find patterns. But sometimes we need to know if a pattern is real or just due to chance. For example, if boys score higher than girls on a test, is that a real difference or just a coincidence?
This is where statistical testing comes in. Statistical tests help us make decisions based on data. They tell us if the differences we see are significant (real) or not. We will learn about the t-test (comparing two groups), ANOVA (comparing more than two groups), and chi-square test (for categorical data).
By the end of this module, you will be able to test your data and make confident decisions. You will be a data decision-maker!
Chidi and his friend Bola argued about who was the better basketball player. They decided to record the number of points they scored in 10 games. Chidi's average was 12 points, and Bola's was 10 points. But was that difference real or just luck? They used a t-test to find out!
The t-test gave a p-value of 0.20. Since this was greater than 0.05, they concluded that the difference was not significant β it could have happened by chance. They decided they were equally good and stopped arguing. Statistical testing helped them settle the debate.
Definition: Statistical testing is a way to decide if a pattern in data is real or just due to chance.
Why it is important: It helps us make objective decisions, not just guesses.
Simple explanation: It's like a referee who decides if a goal is valid or not.
Real-life example: A doctor tests if a new medicine works better than an old one.
School example: A teacher tests if teaching method A gives better scores than method B.
Home example: You test if you sleep better with or without a nightlight.
Nigerian example: A researcher tests if a new farming method increases crop yield.
Illustration:
Data --> Hypothesis --> Test --> p-value --> Decision
Mini summary: Statistical testing helps us decide if patterns are real.
Definition: The p-value is the probability of getting the observed result (or more extreme) if there is no real difference (i.e., if the null hypothesis is true).
Why it is important: A small p-value (usually < 0.05) means the result is unlikely to be due to chance, so we call it statistically significant.
Simple explanation: It's like a magic number that tells you how surprising your result is. If it's very small (less than 5%), it's a real effect.
Real-life example: A p-value of 0.03 means there's only a 3% chance the difference is due to luck.
School example: If p < 0.05, the new teaching method is likely better.
Home example: If p < 0.05, the nightlight might really affect your sleep.
Nigerian example: If p < 0.05, the new fertiliser really increases crop yield.
Key rule: p < 0.05 = significant (real effect). p > 0.05 = not significant (could be chance).
Mini summary: p-value tells us if a result is real or due to chance.
Definition:
Why it is important: The test decides whether to reject the null hypothesis.
Simple explanation: Null is like saying "nothing is going on"; alternative is "something is happening".
Real-life example: H0: the medicine doesn't work; H1: the medicine works.
School example: H0: boys and girls have same average score; H1: they differ.
Home example: H0: the nightlight has no effect; H1: it affects sleep.
Nigerian example: H0: fertiliser has no effect; H1: it increases yield.
Mini summary: We test if we can reject H0 in favour of H1.
Definition: The t-test compares the means of two groups to see if they are significantly different.
Why it is important: It's the most common test for comparing two groups.
Simple explanation: It's like weighing two boxes to see if one is heavier.
Real-life example: Compare the height of boys and girls.
School example: Compare test scores of class A and class B.
Home example: Compare your screen time with your friend's.
Nigerian example: Compare average income of urban vs rural areas.
Code:
t.test(data$score ~ data$group) # where group has two categories # Example: t.test(scores ~ gender)
t-test output: t = 2.45, df = 18, p-value = 0.025 alternative hypothesis: true difference in means is not equal to 0 95% confidence interval: [2.1, 8.9] sample estimates: mean in group A mean in group B 75.0 69.0
Mini summary: t-test compares means of two groups.
Steps:
Mini summary: p < 0.05 = significant difference.
Definition: t-test assumes:
Why it is important: If assumptions are violated, the test may not be valid.
Simple explanation: It's like rules of a game β you need to follow them.
Mini summary: Check assumptions before using t-test.
Definition: Analysis of Variance (ANOVA) compares means of three or more groups.
Why it is important: When you have more than two groups, you use ANOVA.
Simple explanation: It's like a t-test for many groups.
Real-life example: Compare test scores of three different schools.
School example: Compare scores of students from different teachers.
Home example: Compare time spent on different activities.
Nigerian example: Compare crop yield across four regions.
Code:
# One-way ANOVA aov_model <- aov(score ~ group, data = data) summary(aov_model)
ANOVA output:
Df Sum Sq Mean Sq F value Pr(>F)
group 2 123.4 61.7 4.23 0.023 *
Residuals 27 393.5 14.6
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Mini summary: ANOVA compares means of three or more groups.
Definition: After ANOVA, if the result is significant, post-hoc tests tell us which groups are different.
Why it is important: ANOVA tells us there is a difference, but not where it is.
Simple explanation: It's like finding out which specific pair of teams caused the difference.
Code:
TukeyHSD(aov_model)
Mini summary: Post-hoc tests identify which groups differ.
Definition: The chi-square test is used when data is categorical (counts in categories).
Why it is important: It tests if two categorical variables are related.
Simple explanation: It's like checking if favourite foods are different between boys and girls.
Real-life example: Test if gender and favourite subject are related.
School example: Test if class and preference for school lunch are related.
Home example: Test if age and favourite TV show are related.
Nigerian example: Test if region and voting preference are related.
Code:
table_data <- table(data$gender, data$favourite_food) chisq.test(table_data)
Chi-square output: X-squared = 10.2, df = 2, p-value = 0.006
Mini summary: Chi-square tests relationships between categories.
If p < 0.05, the variables are related (dependent). If p > 0.05, they are not related (independent).
Mini summary: p < 0.05 = relationship exists.
Let's say we have data on education level (primary, secondary, tertiary) and income level (low, medium, high). We want to see if education and income are related.
table_data <- table(education, income) chisq.test(table_data) # If p < 0.05, education and income are related.
Mini summary: Chi-square is great for survey data.
| Question | Data Type | Test |
|---|---|---|
| Compare two groups | Numeric | t-test |
| Compare three+ groups | Numeric | ANOVA |
| Relationship between categories | Categorical | Chi-square |
Mini summary: Choose the test based on your data and question.
Definition: Statistical significance means the result is unlikely to be chance. Practical significance means the result is big enough to matter.
Why it is important: A difference can be statistically significant but very small and unimportant.
Simple explanation: A small difference might be real (statistically significant) but not meaningful (e.g., 0.1 kg weight loss).
Mini summary: Always consider if the result is meaningful in real life.
Mini summary: Be careful not to misuse tests.
We learned about p-values, t-tests, ANOVA, and chi-square tests. Testing helps us make decisions based on data.
t.test() with your data.
Group A: **** Group B: ***
(mean = 75) (mean = 69)
t-test checks if the gap is real.
Gap (difference)
|--------------|
75 69
Group A: **** Group B: *** Group C: ***** ANOVA checks if any group differs.
Table of counts:
Food A Food B
Boys 10 5
Girls 6 9
Chi-square checks if food preference is independent of gender.
| Test | Use | Data Type | R Function |
|---|---|---|---|
| t-test | Compare two groups | Numeric | t.test() |
| ANOVA | Compare three+ groups | Numeric | aov() |
| Chi-square | Test relationship | Categorical | chisq.test() |
In this module, we learned about statistical testing β a way to make decisions based on data. We learned about the p-value and how it helps us decide if a result is significant (p < 0.05). We used the t-test for two groups, ANOVA for three or more groups, and chi-square for categorical data. We also learned about assumptions and choosing the right test. Testing is a powerful tool in data analysis.
| Test | Use |
|---|---|
| 1. t-test | A. Categorical data |
| 2. ANOVA | B. Compare two groups |
| 3. Chi-square | C. Compare three+ groups |
Answers: 1-B, 2-C, 3-A
In groups, create a dataset with two groups (e.g., boys and girls) and a numeric variable (e.g., test score). Perform a t-test and interpret the results. Then, create a categorical dataset and perform a chi-square test.
Use the iris dataset. Perform a t-test to compare petal length between two species (e.g., setosa and versicolor). Perform ANOVA to compare petal length among all three species.
Title: "Testing a Nigerian Hypothesis"
Find a dataset from Nigeria (e.g., education, health, or agriculture). Formulate a hypothesis (e.g., "Rural and urban areas have different access to clean water"). Use an appropriate statistical test to test your hypothesis. Write a report of your findings.
Using the mtcars dataset, compare the mpg (miles per gallon) of automatic and manual transmission cars using a t-test. Then, perform ANOVA to compare mpg across different numbers of cylinders (4, 6, 8).
Find a real dataset online. Perform at least three different statistical tests on it (t-test, ANOVA, chi-square). Write a complete report explaining the data, the tests used, and the conclusions.
Fill-in-the-Blank: 1. p-value, 2. 0.05, 3. null, 4. t-test, 5. ANOVA, 6. Chi-square, 7. Post-hoc, 8. Practical, 9. Assumptions, 10. Data dredging.
True/False: 1T, 2F, 3F, 4F, 5T, 6F, 7T, 8F, 9F, 10T.
Multiple Choice: 1A, 2A, 3B, 4C, 5B, 6A, 7A, 8B, 9B, 10A, 11B, 12C, 13A, 14B, 15C.
In Module 15, we will learn about regression β how to predict one variable using another. We will use linear regression to model relationships and make predictions. Practise testing and reflecting on how tests guide decisions.
Hello, young data explorer! In Module 14, we learned how to test if patterns are real. Now, we will learn how to predict one variable from another. Imagine you know a person's height β can you predict their weight? Or if you know how many hours you study, can you predict your test score? This is what regression does.
Regression helps us understand the relationship between two or more variables and use that to make predictions. The simplest type is linear regression, where we draw a straight line through our data to predict outcomes.
By the end of this module, you will be able to build regression models, interpret them, and make predictions. You will be a data predictor!
lm() to build a regression model in R.Chidi sells ice cream at school. He noticed that on hot days, he sells more ice cream. He wanted to predict how many ice creams he would sell based on the temperature. He collected data for 10 days: temperature and sales.
He used linear regression to draw a line through the data points. The line showed that for every degree increase in temperature, he sells 2 more ice creams. Now, when the weather forecast says it will be 30Β°C, he predicts he will sell about 60 ice creams. He can prepare enough stock.
Chidi learned that regression helps him make better decisions!
Definition: Regression is a statistical method to model the relationship between a dependent (response) variable and one or more independent (predictor) variables.
Why it is important: It helps us understand how variables are related and make predictions.
Simple explanation: It's like finding a line that best fits the data points.
Real-life example: A real estate agent uses regression to predict house prices based on size and location.
School example: A teacher predicts test scores based on study time.
Home example: You predict your allowance based on your chores.
Nigerian example: A farmer predicts crop yield based on rainfall.
Illustration:
Data points (dots) --> Fit a line --> Use line to predict new points
Mini summary: Regression models relationships and predicts outcomes.
Definition: Linear regression fits a straight line to the data: y = a + b*x, where:
Simple explanation: It's like drawing a straight line through the cloud of dots.
Real-life example: y = 2x + 10, where x is temperature, y is ice cream sales.
School example: y = 2*study_time + 50 (predicting score).
Home example: y = 1.5*chores + 5 (predicting allowance).
Nigerian example: y = 0.5*rainfall + 20 (predicting crop yield).
Illustration:
y ^ | / (line: y = a + b*x) | / | / |/__________________ x
Mini summary: Linear regression fits a straight line to data.
Definition: lm() is the R function for linear regression. It stands for "linear model".
Why it is important: It's the standard way to perform regression in R.
Simple explanation: It's like giving R a recipe to draw the best line.
Code:
model <- lm(y ~ x, data = data) # y ~ x means "y is predicted by x"
Example: model <- lm(score ~ study_time, data = students)
Mini summary: Use lm() to create a regression model.
Definition: summary(model) gives detailed output including coefficients, R-squared, and p-values.
Why it is important: It tells you how good the model is and the strength of relationships.
Simple explanation: It's like a report card for your model.
Code:
summary(model)
Output example:
Call:
lm(formula = score ~ study_time, data = students)
Residuals:
Min 1Q Median 3Q Max
-8.345 -3.234 0.123 2.456 7.890
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 50.1234 2.3456 21.376 <2e-16 ***
study_time 2.3456 0.4567 5.135 0.0002 ***
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 4.567 on 28 degrees of freedom
Multiple R-squared: 0.485, Adjusted R-squared: 0.467
F-statistic: 26.38 on 1 and 28 DF, p-value: 0.0002
Mini summary: summary() evaluates your model.
Definition:
Why they are important: They define the line and tell you the relationship.
Simple explanation: Intercept is where the line starts; slope is how steep it is.
Real-life example: In our ice cream example, intercept = 10 (sales when 0Β°C) and slope = 2 (each degree adds 2 sales).
School example: Intercept = 50, slope = 2 (score = 50 + 2*study_time).
Home example: Intercept = 5, slope = 1.5 (allowance = 5 + 1.5*chores).
Nigerian example: Intercept = 20, slope = 0.5 (yield = 20 + 0.5*rainfall).
Mini summary: Coefficients tell you the line's equation.
Definition: R-squared (RΒ²) measures the proportion of variance in y that is explained by x. It ranges from 0 to 1.
Why it is important: It tells you how well the line fits the data.
Simple explanation: An RΒ² of 0.8 means 80% of the variation in y is explained by x.
Real-life example: RΒ² = 0.75 means temperature explains 75% of the variation in ice cream sales.
School example: RΒ² = 0.65 means study time explains 65% of score variation.
Home example: RΒ² = 0.50 means chores explain 50% of allowance variation.
Nigerian example: RΒ² = 0.70 means rainfall explains 70% of crop yield variation.
Rule: Higher RΒ² means better fit.
Mini summary: RΒ² measures how well the model predicts.
Definition: predict() uses the model to predict y for new x values.
Why it is important: It lets you forecast future outcomes.
Simple explanation: It's like using the line to read off predicted values.
Code:
new_data <- data.frame(study_time = c(5, 6, 7)) predict(model, newdata = new_data)
Mini summary: predict() makes predictions.
Definition: Residuals are the differences between actual and predicted y values.
Why they are important: They show how far off the model is. Good models have small residuals.
Simple explanation: It's like the mistakes the model makes.
Code:
residuals(model)
Mini summary: Residuals measure the model's errors.
Linear regression assumes:
Why it is important: Violations can make the model unreliable.
Code:
plot(model) # diagnostic plots
Mini summary: Check assumptions to ensure valid model.
Definition:
Why it is important: Multiple regression can capture more complex relationships.
Code:
model_multiple <- lm(score ~ study_time + sleep_hours, data = students)
Mini summary: Multiple regression uses multiple predictors.
Let's predict crop yield based on rainfall and fertiliser usage.
model <- lm(yield ~ rainfall + fertiliser, data = farms) summary(model) # Predict yield for a farm with 100mm rainfall and 50kg fertiliser predict(model, newdata = data.frame(rainfall = 100, fertiliser = 50))
Mini summary: Regression is useful for Nigerian agriculture.
Definition: Add the regression line to a scatter plot.
Code:
ggplot(data, aes(x = study_time, y = score)) +
geom_point() +
geom_smooth(method = "lm", se = FALSE) +
labs(title = "Scatter Plot with Regression Line")
Scatter plot with line:
Score
100 | . /
80 | . . /
60 | . / .
40 | . /
20 |. /
0 |___._.__/.___.___.___
0 20 40 60 80 100
Study Time
Mini summary: Visualising the line helps interpret the model.
Definition: Overfitting is when a model fits the training data too closely but performs poorly on new data.
Why it is important: A good model should generalise to new data.
Simple explanation: It's like memorising a test instead of learning the material.
Mini summary: Avoid overfitting by keeping the model simple.
You can use categorical variables (like gender) as predictors. R automatically converts them to dummy variables.
model <- lm(score ~ gender, data = data)
Mini summary: Categorical predictors can be used in regression.
Regression models relationships and makes predictions. We learned about simple and multiple regression, coefficients, R-squared, and assumptions.
lm(y ~ x, data = data).summary() to view results.predict().glm for logistic regression.predict() for forecasting.y ^ | / (line) | / | / |/ +--------------------> x
Actual points: * * Predicted line: / Residuals: | (vertical distances)
| Feature | Simple | Multiple |
|---|---|---|
| Number of predictors | 1 | 2 or more |
| Equation | y = a + b*x | y = a + b1*x1 + b2*x2 + ... |
| Use | Simple relationships | Complex relationships |
In this module, we learned about regression β a tool for predicting one variable from another. We used linear regression to fit a line to data, interpreted coefficients and R-squared, and made predictions using predict(). We also discussed assumptions, residuals, and the difference between simple and multiple regression. Regression is a powerful way to understand and forecast data.
predict().| Term | Definition |
|---|---|
| 1. Intercept | A. Change in y per unit x |
| 2. Slope | B. Value of y when x=0 |
| 3. R-squared | C. Measure of fit |
| 4. Residual | D. Difference between actual and predicted |
Answers: 1-B, 2-A, 3-C, 4-D
In groups, collect data on study time and test scores (or use built-in data). Build a regression model, interpret the results, and make predictions. Present your findings.
Use the mtcars dataset to build a regression model predicting mpg (miles per gallon) from hp (horsepower). Interpret the coefficients and R-squared. Make a prediction for a car with 150 hp.
Title: "Predicting Nigerian Student Performance"
Find or create a dataset of Nigerian student scores and study habits. Build a multiple regression model to predict scores. Identify the most important predictors. Write a report with your findings.
Using the airquality dataset, build a regression model to predict ozone levels based on temperature and wind speed. Interpret the model and check assumptions.
Find a real dataset online. Build a multiple regression model with at least 3 predictors. Evaluate the model, check assumptions, and make predictions. Write a comprehensive report.
Fill-in-the-Blank: 1. Regression, 2. Linear, 3. intercept, 4. slope, 5. R-squared, 6. Residuals, 7. predict, 8. Multiple, 9. Overfitting, 10. Assumptions.
True/False: 1T, 2F, 3F, 4T, 5F, 6T, 7F, 8T, 9F, 10T.
Multiple Choice: 1B, 2A, 3A, 4B, 5A, 6C, 7C, 8B, 9B, 10D, 11D, 12B, 13A, 14A, 15B.
In Module 16, we will learn about data communication β how to present your findings clearly. We will combine all our skills to create reports and dashboards. Practise building regression models and interpreting their outputs.
Hello, young data explorer! In Module 15, we learned how to predict outcomes using regression. But what good is an analysis if you cannot share it with others? This module is about data communication β how to tell a clear and compelling story with your data.
Data communication is like being a storyteller. You take your data, your analysis, and your insights, and you present them in a way that is easy to understand. You use reports, slides, dashboards, and visualisations to share your findings.
By the end of this module, you will be able to create a data report, use R Markdown, and present your work confidently. You will be a data communicator!
Chidi had spent weeks analysing data on school attendance and test scores. He had found interesting patterns, but his teacher asked him to present his findings to the class. He needed to communicate his results clearly.
He used R Markdown to create a report. He included a title, an introduction, his data, his analysis, and his conclusions. He added graphs and tables to make it visual. He also prepared a short presentation. When he presented, the class understood everything. They even asked good questions!
Chidi learned that good communication makes your hard work useful.
Definition: Data communication is the process of sharing your data findings with others in a clear and effective way.
Why it is important: If you cannot communicate your results, your analysis has no impact.
Simple explanation: It's like telling a story β you need a beginning, middle, and end.
Real-life example: A business analyst presents sales data to the CEO.
School example: A student presents a science project to the class.
Home example: You tell your family about your savings.
Nigerian example: A researcher presents findings on agriculture to farmers.
Illustration:
Data --> Analysis --> Insights --> Communication --> Impact
Mini summary: Data communication is sharing your findings to create impact.
Definition: Understand who you are communicating with and what they need to know.
Why it is important: Different audiences need different levels of detail.
Simple explanation: You talk differently to a friend than to a teacher.
Real-life example: You explain a game to a younger child differently than to a friend.
School example: You present your project to the teacher differently than to your classmates.
Home example: You explain your daily routine to a visitor.
Nigerian example: You present data to a community leader differently than to a government official.
Mini summary: Tailor your message to your audience.
A good report has:
Mini summary: Reports should be clear and structured.
Definition: R Markdown is a tool that combines R code and narrative text to create dynamic documents (HTML, PDF, Word).
Why it is important: It makes reproducible reporting easy.
Simple explanation: It's like a notebook where you write text and code together.
Code:
---
title: "My Data Report"
author: "Chidi"
date: "2026-07-04"
output: html_document
---
```{r}
summary(cars)
```
Mini summary: R Markdown combines code and text in one document.
Steps:
Mini summary: R Markdown is easy to use in RStudio.
Simple formatting:
# This is a heading This is *italic* and **bold**. - Bullet point
Mini summary: R Markdown uses simple formatting.
Definition: Code chunks are sections where you write R code. They are enclosed in ```{r} ... ```.
Why it is important: They run your analysis and show results in the report.
```{r}
library(ggplot2)
ggplot(mtcars, aes(x = hp, y = mpg)) +
geom_point()
```
Mini summary: Code chunks run R code in your report.
Definition: You can embed R code within text using `r code`.
Why it is important: It makes dynamic text, e.g., "The mean is `r mean(data$score)`".
The average score is `r mean(scores)`.
Mini summary: Inline code makes reports dynamic.
Definition: Good visualisations are clear and informative. They should tell a story.
Why it is important: A picture is worth a thousand words.
Tips:
ggplot(data, aes(x = study_time, y = score)) +
geom_point() +
labs(title = "Study Time vs Score",
x = "Study Time (hours)",
y = "Test Score") +
theme_minimal()
Mini summary: Clear visualisations enhance communication.
Definition: Tables organise and present numeric data clearly.
Why it is important: They provide detailed information.
Code:
library(knitr) kable(head(data), caption = "First few rows of data")
Mini summary: Tables present data in a structured way.
Definition: A narrative is the story you tell about your data. It connects the data to the real world.
Why it is important: It makes your report engaging and meaningful.
Simple explanation: Instead of just showing numbers, explain what they mean.
Example: "The data shows that students who studied more got higher scores. This suggests that study time is important for success."
Mini summary: A narrative makes data meaningful.
Definition: R Markdown can also create presentations (e.g., reveal.js, ioslides).
Why it is important: Slides are great for live presentations.
--- title: "My Presentation" output: ioslides_presentation ---
Mini summary: R Markdown creates slide presentations.
Definition: Dashboards are interactive reports with multiple panels.
Why it is important: They allow users to explore data themselves.
Mini summary: Dashboards are interactive reports.
Create a report on Nigerian education data: import, clean, analyse, and visualise. Include a narrative about the state of education.
Mini summary: Apply all skills to Nigerian data.
We learned to communicate data effectively through reports, visualisations, and narratives. Good communication makes data valuable.
+-----------------------+ | Title | +-----------------------+ | Introduction | +-----------------------+ | Data | +-----------------------+ | Analysis | +-----------------------+ | Visualisations | +-----------------------+ | Conclusions | +-----------------------+
| Tool | Use | Output |
|---|---|---|
| R Markdown | Reports | HTML, PDF, Word |
| flexdashboard | Dashboards | HTML |
| Shiny | Interactive apps | Web app |
| Slides | Presentations | HTML, PDF |
In this module, we learned about data communication β how to share your data findings effectively. We used R Markdown to create reports that combine code, text, and visualisations. We learned about structuring reports, writing a narrative, and presenting to an audience. Good communication makes data analysis valuable and impactful.
| Concept | Description |
|---|---|
| 1. R Markdown | A. Interactive report |
| 2. Dashboard | B. Dynamic reports |
| 3. Narrative | C. Story behind data |
| 4. Visualisation | D. Graph or chart |
| 5. Audience | E. People you communicate with |
Answers: 1-B, 2-A, 3-C, 4-D, 5-E
In groups, analyse a dataset (e.g., iris). Create a report in R Markdown including an introduction, analysis, visualisations, and conclusions. Present your report to the class.
Create an R Markdown report on a topic of your choice (e.g., your hobbies, school data). Include at least one visualisation and a narrative. Knit it to HTML and share it.
Title: "Nigerian Data Report"
Find a dataset about Nigeria (e.g., education, health, agriculture). Perform a complete analysis: import, clean, explore, visualise, and model. Create a comprehensive R Markdown report that tells a story about the data.
Using the economics dataset, create a report that shows trends in unemployment, population, and GDP. Include visualisations, summaries, and a narrative. Knit to HTML.
Find a complex dataset online. Create a dashboard using flexdashboard or Shiny that allows users to explore the data. Include filters and interactive elements.
Fill-in-the-Blank: 1. Data communication, 2. R Markdown, 3. narrative, 4. Code chunks, 5. Knit, 6. dashboard, 7. Visualisations, 8. audience, 9. title, 10. Slides.
True/False: 1F, 2T, 3F, 4T, 5T, 6F, 7F, 8T, 9F, 10T.
Multiple Choice: 1C, 2B, 3B, 4B, 5B, 6B, 7B, 8B, 9B, 10B, 11A, 12A, 13B, 14B, 15B.
In the next module, we will bring everything together in a capstone project. You will apply all the skills you have learned to a real-world data analysis project. Start thinking about a dataset and a question you want to answer.