Hello, young data explorer! You have been learning a lot about R, but there is another powerful tool for data analysis called SPSS. SPSS stands for Statistical Package for the Social Sciences. It is a software that helps you analyze data without writing code.
SPSS is used by many researchers, businesses, and governments around the world, including in Nigeria. It is popular because it has a point-and-click interface, which means you can click buttons and menus to do analysis, instead of typing commands. But don't worry โ you can also write syntax (code) in SPSS, just like in R.
In this guide, we will learn the basics of SPSS: how to enter data, clean it, explore it, and run some tests. By the end, you will be able to use SPSS for your own projects.
Chidi's school wanted to understand why some pupils were doing better than others. They conducted a survey and collected data on study hours, sleep, and test scores. The data was in a spreadsheet, but the teachers didn't know how to analyze it.
They called in a data analyst who used SPSS. The analyst imported the data, cleaned it, and ran some tests. In just a few minutes, they found that pupils who studied more and slept at least 7 hours had higher scores. The teachers used this information to improve study habits.
SPSS made the analysis fast and easy!
Definition: SPSS is a software for statistical analysis. It is used to clean, analyze, and visualize data.
Why it is important: It is user-friendly and widely used in research and business.
Simple explanation: It's like a calculator for data analysis, but much more powerful.
Real-life example: A market research company uses SPSS to analyze customer surveys.
School example: A teacher uses SPSS to analyze test scores.
Home example: You could use SPSS to analyze your family's expenses.
Nigerian example: The National Bureau of Statistics uses SPSS for surveys.
Mini summary: SPSS helps you analyze data without complex coding.
Definition: SPSS has two main views: Data View (where you see your data like a spreadsheet) and Variable View (where you define your variables).
Why it is important: You need to use both to work effectively.
Simple explanation: Data View is like a table of numbers; Variable View is where you give names and labels.
Illustration:
Data View: +------+------+-------+ | name | age | score | +------+------+-------+ | Ade | 10 | 85 | | Bola | 11 | 70 | +------+------+-------+ Variable View: +----------+--------+---------+ | Name | Type | Label | +----------+--------+---------+ | name | String | Pupil | | age | Numeric| Age | | score | Numeric| Score | +----------+--------+---------+
Mini summary: Data View shows the data; Variable View defines the variables.
Definition: You can type data directly into the Data View, or import from Excel, CSV, etc.
Why it is important: You need to get data into SPSS to analyze it.
Simple explanation: It's like typing into a spreadsheet.
Steps to enter data:
To import from Excel: File โ Open โ Data โ Choose your Excel file.
Mini summary: Enter data directly or import from files.
Definition: In Variable View, you define each variable's name, type (Numeric or String), label (description), and value labels (e.g., 1 = Male, 2 = Female).
Why it is important: This makes the data understandable and analysis easier.
Simple explanation: It's like giving your columns names and labels.
Example:
Name: Gender Type: Numeric Label: Gender of pupil Values: 1 = Male, 2 = Female
Mini summary: Variable View helps you define your data.
Definition: Cleaning means checking for missing values, duplicates, and errors.
Why it is important: Dirty data leads to wrong conclusions.
Simple explanation: It's like washing fruits before eating them.
How to clean:
Mini summary: Clean your data before analysis.
Definition: Descriptive statistics describe the data: frequencies, mean, median, mode, standard deviation.
Why it is important: It gives you a snapshot of your data.
How to do it:
Output example:
Score N 30 Mean 78.5 Median 80.0 Std. Dev 10.2 Minimum 60 Maximum 95
Mini summary: Descriptive statistics summarize your data.
Definition: Charts help you visualize your data.
Why it is important: Visuals are easier to understand than numbers.
How to create a bar chart:
How to create a histogram:
Mini summary: Charts make data visual.
Definition: The t-test compares the means of two groups (e.g., boys vs girls).
Why it is important: It tells you if the groups are significantly different.
How to do it:
Output: Look at the p-value (Sig.). If p < 0.05, the groups are different.
Mini summary: t-test compares two groups.
Definition: ANOVA compares means of three or more groups (e.g., different classes).
Why it is important: It tests if any group is different.
How to do it:
Output: Look at the p-value (Sig.). If p < 0.05, at least one group differs.
Mini summary: ANOVA compares three or more groups.
Definition: Chi-square tests if two categorical variables are related (e.g., gender and favourite subject).
Why it is important: It shows association between categories.
How to do it:
Output: Look at the p-value (Asymp. Sig.). If p < 0.05, the variables are related.
Mini summary: Chi-square tests relationships between categories.
Definition: Regression predicts one variable from another.
Why it is important: It helps understand relationships and make predictions.
How to do it:
Output: Look at R-squared and coefficients.
Mini summary: Regression predicts outcomes.
Definition: You can save your output (tables and charts) as files.
Why it is important: You can share your results with others.
How to export: File โ Export โ Choose format (Word, Excel, PDF).
Mini summary: Save your output for sharing.
Let's say we have data on pupils from three different schools in Lagos. We want to compare their test scores and see if there is a relationship between gender and favorite subject.
Mini summary: SPSS can analyze any Nigerian dataset.
Definition: SPSS syntax is the code that runs behind the scenes. You can write and save syntax to repeat analyses.
Why it is important: Syntax makes your work reproducible.
How to use syntax: File โ New โ Syntax. Type your commands and run them.
Example syntax:
FREQUENCIES VARIABLES=age score /STATISTICS=MEAN MEDIAN. T-TEST GROUPS=gender(1 2) /VARIABLES=score.
Mini summary: Syntax is the code version of SPSS.
SPSS is a powerful tool for data analysis. It has a point-and-click interface and a syntax option. You can clean data, explore it, test hypotheses, and make predictions. It is widely used in Nigeria and around the world.
Import Data --> Clean Data --> Explore Data --> Test Hypotheses --> Report Results
| Feature | SPSS | R |
|---|---|---|
| Interface | Point-and-click | Code-based |
| Learning curve | Easier | Steeper |
| Flexibility | Less | More |
| Cost | Paid (but free trial) | Free |
SPSS is a user-friendly software for data analysis. It allows you to clean, explore, test, and visualize data without writing complex code. You can also use syntax for reproducibility. SPSS is widely used in Nigeria and around the world. With the skills from this guide, you can analyze your own data and answer important questions.
| Concept | Description |
|---|---|
| 1. Data View | A. Defines variables |
| 2. Variable View | B. Shows the data |
| 3. t-test | C. Compare three+ groups |
| 4. ANOVA | D. Compare two groups |
| 5. Chi-square | E. Categorical test |
Answers: 1-B, 2-A, 3-D, 4-C, 5-E
In groups, create a small dataset (e.g., 10 rows) with variables like name, age, gender, and score. Enter it into SPSS. Define variables, clean it, and run descriptive statistics. Present your results.
Use the iris dataset (you can import it from a CSV). In SPSS, run a t-test comparing petal length between two species (e.g., setosa and versicolor). Interpret the output.
Title: "Analyzing Nigerian School Data with SPSS"
Find or create a dataset on Nigerian schools (e.g., test scores, attendance). Import it into SPSS, clean it, run descriptive statistics, perform at least two different tests, and write a short report.
Using the mtcars dataset (import as CSV), perform the following in SPSS:
Find a real dataset from Nigeria (e.g., NBS data). Import it into SPSS, perform a complete analysis (cleaning, descriptive, tests, regression), and write a comprehensive report.
Fill-in-the-Blank: 1. Statistical Package for the Social Sciences, 2. Data, 3. Variable, 4. mean, 5. t-test, 6. ANOVA, 7. Chi-square, 8. Regression, 9. p-value, 10. Syntax.
True/False: 1F, 2T, 3T, 4F, 5F, 6T, 7T, 8T, 9F, 10F.
Multiple Choice: 1B, 2B, 3B, 4A, 5C, 6B, 7B, 8A, 9B, 10B, 11A, 12B, 13D, 14B, 15A.
In the next module, we will bring everything together in a capstone project. You will apply all the skills you have learned โ in R and SPSS โ to a real-world data analysis project. Start thinking about a dataset and a question you want to answer.
Hello, young data explorer! Welcome to the world of SPSS. SPSS is a special computer program that helps people understand numbers and data. It is used by scientists, doctors, teachers, and even people in government to answer questions like "How many children like reading?" or "Which medicine works best?"
In this module, we will learn what SPSS is, how to open it, and how to look at data. By the end, you will be comfortable with the SPSS screen and ready to start your first analysis. Don't worry if you have never used a data program before โ we will go step by step, just like learning to ride a bicycle.
Chidi was a bright 10-year-old who loved numbers. His teacher, Mrs. Ade, asked him to help organize the class's test scores. She gave him a paper list with 30 names and scores. Chidi tried to find the average score, but it took him a long time with just a calculator. His fingers got tired, and he made some mistakes.
Then, Mrs. Ade showed him SPSS. She explained that SPSS is like a super-smart calculator that can handle hundreds of numbers at once. Chidi typed the scores into SPSS, and with just a few clicks, the computer showed the average, the highest, and the lowest scores. It even drew a colorful bar chart!
Chidi was amazed. He thought, "This is like magic!" From that day, he knew he wanted to learn more about SPSS. And now, you can learn too.
Definition: SPSS (which stands for Statistical Package for the Social Sciences) is a computer program that helps you analyze data. Data means numbers and facts that we collect.
Why it is important: SPSS makes it easy to organize, understand, and show data. It can handle thousands of numbers without making mistakes.
Simple explanation: Think of SPSS as a very smart calculator that can also draw pictures with your numbers.
Real-life example: A doctor uses SPSS to find out which medicine helps patients get better faster.
School example: A teacher uses SPSS to find the average test score of the whole class.
Home example: You could use SPSS to keep track of your pocket money and see how much you spend on sweets.
Nigerian example: The National Bureau of Statistics in Nigeria uses SPSS to count how many people live in each state.
Illustration:
Data (numbers) --> SPSS --> Results (answers)
(raw facts) (tool) (mean, charts, etc.)
Mini summary: SPSS is a tool that helps you understand numbers.
Definition: When you open SPSS, you will see a window with different parts: the menu bar, the toolbar, and the data area.
Why it is important: Knowing the parts of the window helps you find your way around.
Simple explanation: It is like looking at the dashboard of a car โ you need to know where the steering wheel and the pedals are.
How to open SPSS: Click the SPSS icon on your computer (it looks like a green and blue square). If you don't have it, ask your teacher to help install it.
Illustration:
+--------------------------------------------------+ | File Edit View Data Transform Analyze ... | <- Menu Bar +--------------------------------------------------+ | [Open] [Save] [Print] [Undo] ... | <- Toolbar +--------------------------------------------------+ | | | Data View (spreadsheet with rows and columns) | | | +--------------------------------------------------+ | Variable View (where you define variables) | +--------------------------------------------------+
Mini summary: The SPSS window has menus, buttons, and a big area for data.
Definition: SPSS has two main views: Data View (looks like a spreadsheet) and Variable View (where you define the columns).
Why it is important: You need to use both views to work with data.
Simple explanation: Data View is like a table with rows and columns. Variable View is like a form where you name each column.
Real-life example: Think of Data View as the pages of a notebook filled with numbers. Variable View is the index where you write what each column means.
School example: Data View has pupil names and scores. Variable View has labels like "Name" and "Score".
Home example: Data View has days of the week and hours spent playing. Variable View has "Day" and "Hours".
Nigerian example: Data View has states and populations. Variable View has "State" and "Population".
Illustration:
Data View: +-------+------+-------+ | Name | Age | Score | +-------+------+-------+ | Ade | 10 | 85 | | Bola | 11 | 70 | +-------+------+-------+ Variable View: +--------+--------+---------+---------+ | Name | Type | Label | Values | +--------+--------+---------+---------+ | Name | String | Pupil | None | | Age | Numeric| Age | None | | Score | Numeric| Score | None | +--------+--------+---------+---------+
Mini summary: Data View shows the data; Variable View gives details about the data.
Definition: In SPSS, rows are called cases (like one person or one thing). Columns are called variables (like age or score).
Why it is important: Understanding rows and columns helps you organize data correctly.
Simple explanation: Rows are the "who" or "what" you are studying. Columns are the "what you want to know" about them.
Real-life example: A doctor studies patients (rows) and records their age, weight, and blood pressure (columns).
School example: A teacher has pupils (rows) and records their names, test scores, and attendance (columns).
Home example: You record your family members (rows) and their ages, favorite foods, and hobbies (columns).
Nigerian example: A researcher records states (rows) and their population, area, and capital city (columns).
Illustration:
+------------------+-------------------+-------------------+ | CASES (rows) | VARIABLE 1 | VARIABLE 2 | +------------------+-------------------+-------------------+ | Pupil 1 | Name: Ade | Score: 85 | | Pupil 2 | Name: Bola | Score: 70 | +------------------+-------------------+-------------------+
Mini summary: Rows are cases, columns are variables.
Definition: You can type data directly into SPSS, just like you type into a spreadsheet.
Why it is important: This is how you create a new dataset.
Simple explanation: Click on a cell in Data View, type a number or word, and press Enter.
Steps:
Remember: Each row is one person or thing. Each column is a different piece of information.
Mini summary: Typing data in SPSS is like filling in a table.
Definition: In Variable View, you can give each column a name (like "Score") and a label (like "Final Test Score").
Why it is important: Names and labels make it clear what the data means.
Simple explanation: It's like putting name tags on boxes so you know what's inside.
How to do it:
Example:
Name: Score Type: Numeric Label: Score in Math Test
Mini summary: Variable View helps you name and describe your data.
Definition: Saving your work means storing it on your computer so you can open it later.
Why it is important: If you don't save, you might lose all your hard work!
Simple explanation: It's like saving a game โ you don't want to start over.
How to save:
Tip: SPSS files have a .sav extension.
Mini summary: Always save your SPSS file.
Definition: Opening a saved file brings your data back into SPSS.
Why it is important: You can continue your work anytime.
How to open:
Mini summary: Opening a file lets you continue your analysis.
Definition: When you run an analysis in SPSS, the results appear in a new window called the Output Viewer.
Why it is important: The Output Viewer shows your tables, charts, and statistics.
Simple explanation: It's like a report that SPSS writes for you.
How to view output: After running an analysis (like a mean or a chart), SPSS automatically opens the Output Viewer.
You can: Scroll through results, copy them, or save them.
Mini summary: The Output Viewer shows your analysis results.
Definition: The menu bar at the top of the SPSS window has options like File, Edit, View, Data, Transform, Analyze, and Graphs.
Why it is important: These menus contain all the tools you need.
Simple explanation: It's like a toolbox with many different tools.
Main menus:
Mini summary: The menus are your toolbox in SPSS.
Definition: The toolbar has icons (small pictures) that are shortcuts to common actions like Open, Save, and Print.
Why it is important: They help you work faster.
Simple explanation: Instead of clicking File then Save, you can just click the Save icon (the floppy disk).
Common icons:
Mini summary: The toolbar gives you quick shortcuts.
Definition: SPSS has a Help system that can explain things if you get stuck.
Why it is important: You don't need to memorize everything โ help is available.
How to get help:
Mini summary: Use Help when you are unsure.
Definition: Exiting SPSS means closing the program.
Why it is important: You should close programs properly to avoid losing data.
How to exit:
Mini summary: Always exit SPSS properly.
Let's imagine a school in Lagos. They collected data on 10 pupils: names, ages, and test scores. We can enter this data into SPSS.
Data View (example): +------+-----+-------+ | Name | Age | Score | +------+-----+-------+ | Ade | 10 | 85 | | Bola | 11 | 70 | | Chidi| 9 | 90 | | Dami | 12 | 65 | +------+-----+-------+
Then we can save the file as "School Data.sav".
Mini summary: SPSS works with any data from Nigeria.
We have learned what SPSS is, how to open it, and how to navigate the interface. We know the difference between Data View and Variable View, rows and columns, and how to enter and save data. You are now ready to use SPSS!
+------------------------------------------------------+ | File Edit View Data Transform Analyze Graphs | <- Menu Bar +------------------------------------------------------+ | [Open] [Save] [Print] [Undo] [Redo] ... | <- Toolbar +------------------------------------------------------+ | | | Data View (spreadsheet) | | +--------+--------+--------+ | | | Name | Age | Score | | | +--------+--------+--------+ | | | Ade | 10 | 85 | | | | Bola | 11 | 70 | | | +--------+--------+--------+ | | | +------------------------------------------------------+ | Data View | Variable View | (tabs at bottom) | +------------------------------------------------------+
| Feature | Data View | Variable View |
|---|---|---|
| What it shows | The actual data | Details about variables |
| Rows | Cases (people/things) | Variables |
| Columns | Variables | Properties (name, type, label) |
| Use | Enter and view data | Define and describe data |
In this first module, we have taken our first steps into the world of SPSS. We learned that SPSS is a powerful tool for analyzing data. We explored the two main views โ Data View and Variable View โ and learned about rows and columns (cases and variables). We also practiced entering data, defining variables, saving our work, and navigating the menus. You are now ready to move on to the next module, where we will learn how to clean data.
| Term | Description |
|---|---|
| 1. Case | A. A column in SPSS |
| 2. Variable | B. A description of a variable |
| 3. Label | C. A row in SPSS |
| 4. Data View | D. Shows results |
| 5. Output Viewer | E. Shows the data |
Answers: 1-C, 2-A, 3-B, 4-E, 5-D
In groups of three, create a small dataset of your favourite foods. Each person should have a name, favourite food, and rating (1-10). Enter the data into SPSS. Define variables and save the file. Present your data to the class.
Open SPSS. Create a dataset with 5 of your classmates. Include their names, ages, and favourite subject. Define variables and save the file as "Class Data.sav".
Title: "My First SPSS Dataset"
Create a dataset of 15 people. Include their names, age, and favourite colour. Enter it into SPSS, define variables, and save the file. Write a short paragraph describing your data.
Open the sample dataset "Employee Data" (if available in SPSS) or download a simple CSV file. Open it in SPSS, explore Data View and Variable View, and write down what you see.
Find a dataset about Nigeria (e.g., from the National Bureau of Statistics). Import it into SPSS and explore it. Write a short report on what the data contains.
Fill-in-the-Blank: 1. Statistical Package for the Social Sciences, 2. Data, 3. Variable, 4. cases, 5. variables, 6. label, 7. Output Viewer, 8. File, 9. toolbar, 10. .sav.
True/False: 1F, 2F, 3T, 4F, 5F, 6T, 7T, 8F, 9T, 10F.
Multiple Choice: 1B, 2A, 3B, 4A, 5B, 6B, 7C, 8A, 9B, 10A, 11D, 12A, 13C, 14A, 15B.
In the next module, we will learn how to clean data in SPSS โ fixing errors, handling missing values, and getting data ready for analysis. Great work completing Module 1! You are now an SPSS beginner.
Hello, young data explorer! In Module 1, we learned how to enter data into SPSS and how to save it. But what if the data has mistakes? What if some numbers are missing, or someone typed the same thing twice? This is called dirty data. Just like dirty clothes need washing, dirty data needs cleaning.
Data cleaning is a very important step. If your data is dirty, your answers will be wrong โ and we don't want that! In this module, you will learn how to find and fix problems in your data. You will learn to check for missing values, remove duplicates, fix spelling mistakes, and make sure everything is correct.
By the end of this module, you will be a data cleaning expert!
Chidi's teacher, Mrs. Ade, had a list of pupils and their test scores. But the list was a mess! Some scores were missing, one pupil's name was spelled "Ade" in one row and "AdE" in another, and two rows were exactly the same. Mrs. Ade was confused.
Chidi remembered learning about data cleaning. He opened the list in SPSS and started cleaning:
Now the list was clean and perfect. Mrs. Ade could calculate the class average without any mistakes. Chidi felt proud โ he had saved the day!
Definition: Dirty data is data that has errors, missing parts, duplicates, or is inconsistent.
Why it is important: Dirty data leads to wrong conclusions. It's like baking a cake with the wrong ingredients โ it won't taste good.
Simple explanation: Dirty data is like a messy room โ you need to tidy it up before you can find anything.
Real-life example: A shop has a list of customers with some addresses missing and some names spelled wrongly.
School example: A class register has some pupils' names written twice.
Home example: Your list of chores has some tasks repeated.
Nigerian example: A census dataset has some states missing population figures.
Illustration:
Dirty Data: Ade, 10, 85
AdE, 10, 85 (duplicate and spelling error)
Bola, 11, (missing score)
Mini summary: Dirty data has mistakes that need fixing.
Definition: The saying "garbage in, garbage out" means that if you put bad data into your analysis, you will get bad results.
Why it is important: Cleaning data ensures your answers are correct and trustworthy.
Simple explanation: If you use dirty water to make juice, the juice will be dirty too.
Real-life example: A doctor uses clean data to decide on the right medicine for a patient.
School example: A teacher uses clean scores to give the correct grades.
Home example: You need clean data to know how much money you have saved.
Nigerian example: The government uses clean data to plan budgets for schools.
Mini summary: Clean data = correct answers. Dirty data = wrong answers.
Definition: Missing values are empty cells in your data. In SPSS, they appear as dots (.) or blanks.
Why it is important: Missing values can affect your calculations (like averages).
Simple explanation: It's like having empty spaces in a puzzle.
How to find missing values: Use Analyze โ Descriptive Statistics โ Frequencies. SPSS will tell you how many values are missing.
How to replace missing values:
Real-life example: A survey respondent didn't answer a question โ you can fill it with the average response.
School example: A pupil was absent โ you can replace the score with the class average.
Home example: You forgot to record your chore time โ you can use the average time.
Nigerian example: A state's population data is missing โ you can estimate it.
Illustration:
Before: Ade, 10, 85
Bola, 11, .
Chidi, 9, 90
After (replace missing with mean): Bola, 11, 87.5
Mini summary: Missing values can be replaced or removed.
Definition: Sometimes it's best to simply remove the row (case) that has missing values.
Why it is important: If many values are missing, it might be better to drop that case.
How to do it:
Mini summary: You can delete rows with missing data.
Definition: Duplicates are rows that are exactly the same or have the same key information (like a name).
Why it is important: Duplicates can make your counts wrong (like counting a pupil twice).
How to find duplicates:
How to remove duplicates:
Real-life example: A mailing list has the same person twice โ you remove one.
School example: A pupil's name appears twice in the register โ you remove one.
Home example: You wrote the same chore twice โ you remove one.
Nigerian example: A state appears twice in a list โ you remove the duplicate.
Mini summary: Duplicate rows can be identified and removed.
Definition: Text cleaning means making all entries consistent โ e.g., "Ade" and "AdE" should be the same.
Why it is important: Inconsistent text can cause problems when you group or count.
How to fix text:
Example: "Ade" and "AdE" both become "ade".
Mini summary: Text can be standardised (e.g., all lowercase).
Definition: An outlier is a number that is very different from the others. For example, if everyone scored 70-80 and one person scored 100, that 100 might be an outlier.
Why it is important: Outliers can distort your analysis (like the average).
How to find outliers:
How to handle outliers:
Real-life example: A student's score is 10 when everyone else scored 90 โ it might be a data entry error.
School example: A test score of 150 when the maximum is 100.
Home example: You recorded 100 hours of screen time in one day.
Nigerian example: A state's population is recorded as 100 million when it should be 10 million.
Mini summary: Outliers are unusual numbers that need to be checked.
Definition: Some values are simply impossible โ like a person's age being 200, or a test score being -5.
Why it is important: These are usually data entry errors and must be fixed.
How to check:
Real-life example: A person's age is listed as 150 โ that's impossible.
School example: A test score of -10 โ that can't happen.
Home example: You spent -5 hours on homework.
Nigerian example: A state has -100,000 population.
Mini summary: Remove or correct impossible values.
Definition: When you have categories (like "Male" and "Female"), you want to make sure they are spelled the same way.
Why it is important: "Male", "male", and "MALE" should all be the same.
How to do it: Use Transform โ Recode into Same Variables or Transform โ Compute with LOWER.
Example: Replace "male" with 1 and "female" with 2 using value labels.
Mini summary: Categories must be consistent.
Definition: Sometimes the easiest way is to click on the cell and type the correct value.
Why it is important: For small datasets, manual fixing is quick.
How to do it:
Mini summary: You can manually edit data in Data View.
Here is a simple process to clean data:
Mini summary: Follow these steps to clean any dataset.
Let's clean a dataset from a Nigerian school.
Original data: Name Age Score Ade 10 85 AdE 10 85 (duplicate and spelling) Bola 11 . (missing) Chidi 9 90 Dami 150 65 (impossible age)
After cleaning:
Name Age Score ade 10 85 bola 11 87.5 (mean replacement) chidi 9 90 dami 12 65 (corrected age)
We removed the duplicate, fixed the spelling, replaced the missing score, and corrected the impossible age.
Mini summary: SPSS can clean any Nigerian dataset.
Definition: After cleaning, you should save your data as a new file so you don't lose the original.
How to do it: File โ Save As โ Give a new name (e.g., "Clean School Data.sav").
Why it is important: You keep the original raw data in case you need it.
Mini summary: Always save clean data as a new file.
Definition: Write down what you did to clean the data. This is called "documentation".
Why it is important: If someone else uses your data, they need to know what you changed.
How to document: You can write a simple note in a text file or use the SPSS Syntax editor.
Example: "Changed missing scores to mean (87.5). Removed duplicate rows. Corrected age from 150 to 12."
Mini summary: Document your cleaning steps.
We have learned how to find and fix missing values, duplicates, text errors, outliers, and impossible values. We also learned to save clean data and document our steps. Cleaning is the foundation of good analysis.
Dirty Data
|
V
Check Missing --> Fix/Remove
|
V
Check Duplicates --> Remove
|
V
Check Text --> Standardise
|
V
Check Outliers --> Fix/Remove
|
V
Clean Data
Before: Name Score Ade 85 Ade 85 (duplicate) After: Name Score Ade 85
| Problem | Method | SPSS Tool |
|---|---|---|
| Missing values | Replace or remove | Transform โ Replace Missing Values |
| Duplicates | Remove | Data โ Identify Duplicate Cases |
| Text errors | Standardise | Transform โ Compute (LOWER) |
| Outliers | Check and fix | Analyze โ Explore (boxplot) |
| Impossible values | Correct | Transform โ Recode |
In this module, we learned how to clean dirty data. We covered missing values, duplicates, text errors, outliers, and impossible values. We used SPSS tools like Transform, Data โ Identify Duplicate Cases, and Analyze โ Explore. We also learned the importance of documentation and saving clean data separately. Now your data is ready for analysis!
| Problem | Solution |
|---|---|
| 1. Missing values | A. Remove duplicates |
| 2. Duplicates | B. Replace or remove |
| 3. Text errors | C. Check boxplot |
| 4. Outliers | D. Standardise text |
Answers: 1-B, 2-A, 3-D, 4-C
In groups, create a dirty dataset with missing values, duplicates, and text errors. Exchange your dataset with another group. Clean their data using SPSS and present your cleaned version.
Find a dataset (or use the sample "Employee Data" in SPSS). Clean it by removing missing values, duplicates, and any errors. Save the clean data and write a short documentation.
Title: "Clean a Nigerian Dataset"
Find a small dataset about Nigeria (e.g., from a school, hospital, or business). It should have at least 20 rows and 5 columns. Clean the data using SPSS, document your steps, and save the clean data.
Open the SPSS sample dataset "cars.sav". Clean it by checking for missing values, duplicates, and outliers. Save the clean dataset as "cars_clean.sav".
Find a dirty dataset online (e.g., from Kaggle). Use SPSS to clean it completely. Write a detailed report on the cleaning process and the final clean dataset.
Fill-in-the-Blank: 1. Dirty, 2. correct, 3. missing value, 4. Duplicate, 5. outlier, 6. impossible, 7. standardise, 8. save, 9. documentation, 10. bad.
True/False: 1F, 2T, 3F, 4F, 5T, 6F, 7T, 8T, 9F, 10T.
Multiple Choice: 1B, 2B, 3A, 4B, 5A, 6B, 7A, 8A, 9A, 10A, 11C, 12A, 13A, 14B, 15D.
In the next module, we will learn how to explore data โ creating charts and summary tables to understand our data better. You will use your clean data to create beautiful visualisations and find interesting patterns.
Hello, young data explorer! In Module 2, we learned how to clean our data so it is nice and tidy. Now that our data is clean, it is time to explore it. Exploring data means looking at it from different angles to understand what it is telling us.
Think of it like being a detective. You have a case (your data), and you need to examine the clues (the numbers). You ask questions like: "What is the average?", "What is the highest?", "How spread out is the data?", and "Are there any patterns?"
In this module, we will use SPSS to create summary tables and charts that help us see our data clearly. By the end, you will be able to describe any dataset and find interesting insights.
Chidi's school conducted a survey. They asked 50 pupils about their favourite subject and how many hours they study each day. The data was clean, but no one had looked at it yet. The principal wanted to know: "What is the most popular subject?" and "How many hours do pupils study on average?"
Chidi opened the data in SPSS. He created a frequency table for favourite subject โ it showed that "Maths" was the most popular. He then calculated the average study time โ it was 2.5 hours per day. He also made a bar chart to show the results visually.
The principal was happy. Chidi had turned raw data into useful information. He learned that exploring data helps answer important questions.
Definition: Exploratory Data Analysis (EDA) is the process of looking at data to find patterns, trends, and interesting facts.
Why it is important: EDA helps you understand your data before you do any formal testing.
Simple explanation: It's like looking at a map before you start a journey โ you want to see where you are and where you might go.
Real-life example: A shop owner looks at sales data to see which products sell best.
School example: A teacher looks at test scores to see which topics are hardest.
Home example: You look at your spending to see where your money goes.
Nigerian example: A researcher looks at census data to see population trends.
Illustration:
Data --> Ask Questions --> Explore --> Find Insights --> Tell Story
Mini summary: EDA is the first step to understanding your data.
Definition: A frequency table shows how many times each category appears in your data.
Why it is important: It helps you see which categories are most common.
Simple explanation: It's like counting how many apples, oranges, and bananas you have in a fruit basket.
How to do it in SPSS:
Real-life example: Counting how many customers bought each product.
School example: Counting how many pupils like each subject.
Home example: Counting how many times you ate each meal.
Nigerian example: Counting the number of states in each region.
Output example:
Subject Frequency Percent Maths 20 40% English 15 30% Science 10 20% Art 5 10% Total 50 100%
Mini summary: Frequency tables count how often each value appears.
Definition: Descriptive statistics are numbers that summarise your data, like the average (mean), middle (median), and most common (mode).
Why it is important: They give you a quick snapshot of your data.
How to do it in SPSS:
Real-life example: Finding the average height of students in a class.
School example: Finding the average test score.
Home example: Finding the average amount of pocket money.
Nigerian example: Finding the average population of states.
Output example:
N = 50 Mean = 2.5 Median = 2.0 Std. Deviation = 1.2 Minimum = 1 Maximum = 6
Mini summary: Descriptive statistics summarise numeric data.
Definition:
Why they are important: They tell you the "typical" value.
Simple explanation: Mean is like sharing sweets equally. Median is the middle sweet. Mode is the sweet you have the most of.
Real-life example: A shop wants to know the average sale amount (mean).
School example: A teacher wants to know the middle test score (median).
Home example: You want to know your most common chore (mode).
Nigerian example: Finding the most common age group (mode).
How to find mode in SPSS: Use Analyze โ Descriptive Statistics โ Frequencies and look at the highest frequency.
Mini summary: Mean, median, and mode are different ways to find the "middle" or "typical" value.
Definition: Standard deviation tells you how much the numbers vary from the mean. A small standard deviation means numbers are close together; a large one means they are spread out.
Why it is important: It shows how consistent or varied your data is.
Simple explanation: If everyone in a class scores 70-80, the standard deviation is small. If scores range from 20 to 100, the standard deviation is large.
Real-life example: A factory wants to know if the weights of its products are consistent.
School example: A teacher wants to know if scores are similar or very different.
Home example: You want to know if your daily screen time is consistent.
Nigerian example: Checking if incomes across states are similar or different.
Mini summary: Standard deviation shows how spread out the data is.
Definition: A bar chart uses rectangular bars to show the count or percentage of each category.
Why it is important: It is an easy and visual way to compare categories.
How to do it in SPSS:
Real-life example: Comparing sales of different products.
School example: Comparing favourite subjects.
Home example: Comparing hours spent on different activities.
Nigerian example: Comparing populations of states.
Illustration:
Count
20 | ###
15 | ### ###
10 | ### ### ###
5 | ### ### ### ###
0 |__###__###__###__###___
Math Eng Sci Art
Mini summary: Bar charts compare categories.
Definition: A pie chart is a circle divided into slices, where each slice represents a percentage of the total.
Why it is important: It shows the proportion of each category.
How to do it in SPSS:
Real-life example: Showing how a budget is divided.
School example: Showing the percentage of pupils in each club.
Home example: Showing how you spend your time.
Nigerian example: Showing the percentage of exports by product.
Mini summary: Pie charts show parts of a whole.
Definition: A histogram is a bar chart for numeric data. It groups numbers into bins and shows how many fall into each bin.
Why it is important: It shows the shape of the distribution (e.g., normal, skewed).
How to do it in SPSS:
Real-life example: Showing the distribution of ages in a group.
School example: Showing the distribution of test scores.
Home example: Showing the distribution of hours spent on homework.
Nigerian example: Showing the distribution of incomes.
Illustration:
Frequency
8 | ###
6 | ### ###
4 | ### ### ###
2 | ### ### ### ###
0 |__###__###__###__###___
1-2 2-3 3-4 4-5
Study Time (hours)
Mini summary: Histograms show the distribution of numeric data.
Definition: A boxplot shows the median, quartiles (spread), and outliers of numeric data.
Why it is important: It quickly shows if there are outliers and how spread out the data is.
How to do it in SPSS:
Real-life example: Showing the spread of salaries in a company.
School example: Showing the spread of test scores.
Home example: Showing the spread of daily screen time.
Nigerian example: Showing the spread of state populations.
Illustration:
+-----+ outlier (o) | | | +--+--+ (box = middle 50%) | | | | | +--+--+ | | +-----+ (whiskers show min and max within range)
Mini summary: Boxplots show spread and outliers.
Definition: A scatter plot shows points for two numeric variables. It helps you see if they are related (e.g., more study time = higher scores).
Why it is important: It shows if two variables change together.
How to do it in SPSS:
Real-life example: Plotting height vs weight.
School example: Plotting study time vs test scores.
Home example: Plotting age vs screen time.
Nigerian example: Plotting education vs income.
Illustration:
Score
100 | .
80 | . .
60 | . .
40 | .
20 |.
0 |___._.___.___.___.___.___
0 2 4 6 8 10
Study Time
Mini summary: Scatter plots show relationships between variables.
Definition: The Output Viewer is the window where SPSS shows your tables and charts.
Why it is important: This is where you see the results of your exploration.
What you can do:
Mini summary: The Output Viewer shows your results.
Let's explore a dataset of Nigerian states with variables: state, region, population, and area.
Example output: The most populated region is South-West. The average population is 5 million.
Mini summary: SPSS can explore any Nigerian data.
Good exploration starts with good questions. Here are some questions you can ask:
Mini summary: Asking questions guides your exploration.
Definition: Documentation means writing down what you found.
Why it is important: It helps you remember and share your insights.
How to do it: Write a short report with:
Mini summary: Document your findings.
We have learned to use frequency tables, descriptive statistics, bar charts, pie charts, histograms, boxplots, and scatter plots. These tools help us understand our data and find insights.
Data --> Frequency Tables --> Descriptive Stats --> Charts --> Insights
Frequency
8 | ###
6 | ### ###
4 | ### ### ###
2 | ### ### ### ###
0 |__###__###__###__###___
1-2 2-3 3-4 4-5
| Chart | Type of Data | Use |
|---|---|---|
| Bar chart | Categorical | Compare categories |
| Pie chart | Categorical | Show parts of a whole |
| Histogram | Numeric | Show distribution |
| Boxplot | Numeric | Show spread and outliers |
| Scatter plot | Numeric | Show relationships |
In this module, we learned how to explore data using SPSS. We used frequency tables, descriptive statistics, and various charts to understand our data. We asked questions and found answers. Exploration is the foundation of all data analysis โ it helps us know our data and find interesting patterns.
| Term | Description |
|---|---|
| 1. Mean | A. Most frequent |
| 2. Median | B. Average |
| 3. Mode | C. Middle |
| 4. Bar chart | D. Shows distribution |
| 5. Histogram | E. Compares categories |
Answers: 1-B, 2-C, 3-A, 4-E, 5-D
In groups, collect data on favourite foods from your classmates. Enter it into SPSS, explore it (frequency tables, bar charts), and present your findings.
Use the iris dataset (import as CSV). Explore it: frequency tables for species, descriptive statistics for petal length, and create a histogram and boxplot.
Title: "Explore Nigerian States"
Find a dataset of Nigerian states (population, region, etc.). Explore it using SPSS: frequency tables, descriptive stats, and at least 3 charts. Write a short report.
Using the mtcars dataset, explore the variables mpg, hp, and cyl. Create frequency tables, descriptive statistics, histograms, and boxplots. Interpret the results.
Find a real dataset online (e.g., from Kaggle). Perform a complete EDA in SPSS: frequency tables, descriptive stats, and charts. Write a detailed report of your findings.
Fill-in-the-Blank: 1. EDA, 2. frequency, 3. mean, 4. median, 5. mode, 6. Standard deviation, 7. bar, 8. pie, 9. histogram, 10. boxplot.
True/False: 1F, 2T, 3F, 4F, 5T, 6T, 7F, 8T, 9T, 10F.
Multiple Choice: 1A, 2A, 3B, 4A, 5C, 6C, 7B, 8A, 9C, 10D, 11C, 12C, 13B, 14B, 15A.
In the next module, we will learn how to compare groups using statistical tests like t-tests and ANOVA. We will use the insights from our exploration to ask deeper questions. Great work exploring your data!
Hello, young data explorer! In Module 3, we learned how to explore our data using charts and summary statistics. We found out interesting things like averages and most common values. But what if we want to compare two groups? For example, do boys and girls have different test scores? Do pupils from different classes have different study times?
In this module, we will learn how to compare groups using SPSS. We will use t-tests (for two groups) and ANOVA (for more than two groups). These tests tell us if the differences we see are real or just due to chance.
By the end of this module, you will be able to answer questions like: "Are there differences between groups?" You will be a group comparison expert!
Chidi and his friend Bola argued about who was better at maths โ boys or girls. They decided to collect test scores from 20 boys and 20 girls. The boys' average was 78, and the girls' average was 82. The boys said, "See! Girls are better!" But the girls said, "That's just a small difference โ it could be luck."
They decided to use a t-test in SPSS. The test gave a p-value of 0.15. Since this was greater than 0.05, they concluded that the difference was not significant โ it could have happened by chance. They decided to be friends again.
Chidi learned that t-tests help you decide if differences are real or just luck.
Definition: Comparing groups means looking at two or more groups (like boys and girls) to see if they are different on some measure (like test scores).
Why it is important: It helps us answer questions like: "Does a new teaching method work better?", "Do older pupils score higher?", or "Are there differences between regions?"
Simple explanation: It's like comparing two teams to see which one is better.
Real-life example: A company wants to know if men and women have different salaries.
School example: A teacher wants to know if boys and girls have different test scores.
Home example: You want to know if you and your friend spend different amounts of time on homework.
Nigerian example: A researcher wants to know if urban and rural areas have different income levels.
Illustration:
Group 1 (Boys) vs Group 2 (Girls) Average Score: 78 Average Score: 82 Are they different? (t-test will tell us)
Mini summary: Comparing groups helps us see if they are different.
Definition: A t-test is a statistical test that compares the means (averages) of two groups to see if they are significantly different.
Why it is important: It tells us if the difference is real or just due to chance.
Simple explanation: It's like a referee that decides if a game was won fairly.
When to use: When you have one categorical variable with TWO groups (e.g., gender: male/female) and one numeric variable (e.g., test score).
Real-life example: Comparing the average height of men and women.
School example: Comparing test scores of pupils from two classes.
Home example: Comparing the time you and your sibling spend on chores.
Nigerian example: Comparing the average income of people in Lagos and Kano.
Illustration:
Group A: **** Group B: ***
(mean = 75) (mean = 69)
t-test checks if the gap is real.
Mini summary: The t-test compares two group averages.
Definition: The p-value is a number that tells you how likely it is that the difference between groups is due to chance.
Why it is important: It helps you decide if the difference is significant (real).
Simple explanation: If the p-value is less than 0.05, we say the difference is "significant" โ it's probably real.
Rule of thumb:
Real-life example: A p-value of 0.03 means there's only a 3% chance the difference is due to luck.
School example: A p-value of 0.07 means there's a 7% chance the difference is luck โ we might not conclude it's real.
Home example: If p < 0.05, you can be confident that you and your friend really differ.
Nigerian example: If p < 0.05, urban and rural incomes are truly different.
Mini summary: p < 0.05 = significant difference.
Definition: The independent t-test compares two independent groups (e.g., boys vs girls).
Why it is important: It's the most common t-test.
How to do it:
Output interpretation: Look at the "Sig. (2-tailed)" column. This is your p-value.
Example output:
t-test for Equality of Means t = 2.45, df = 38, Sig. (2-tailed) = 0.018 Mean Difference = 4.2
Interpretation: p = 0.018 < 0.05, so there is a significant difference.
Mini summary: Use Independent-Samples T Test in SPSS.
When you run a t-test, SPSS gives you several numbers. Here's what they mean:
How to decide:
Mini summary: Use the p-value to make your decision.
Definition: Assumptions are conditions that must be met for the test to be valid.
Why it is important: If assumptions are violated, the results may not be trustworthy.
Assumptions:
How to check in SPSS: Use Explore to check normality, and Levene's test (part of the t-test output) to check equal variances.
Mini summary: Check assumptions to ensure your t-test is valid.
Definition: ANOVA (Analysis of Variance) compares the means of three or more groups.
Why it is important: When you have more than two groups, you use ANOVA instead of a t-test.
Simple explanation: It's like a t-test for many groups.
When to use: One categorical variable with THREE or more groups (e.g., Class 5, Class 6, Class 7) and one numeric variable (e.g., score).
Real-life example: Comparing test scores of three different schools.
School example: Comparing study times of pupils from three classes.
Home example: Comparing the time spent on three different activities.
Nigerian example: Comparing incomes across three regions (North, South, East).
Illustration:
Group A: **** Group B: *** Group C: ***** (mean = 75) (mean = 69) (mean = 82) ANOVA checks if any group differs.
Mini summary: ANOVA compares three or more group averages.
Definition: One-way ANOVA is used when you have one independent variable (with three or more groups) and one dependent numeric variable.
How to do it:
Output interpretation: Look at the "Sig." column in the ANOVA table. This is the p-value.
Example output:
ANOVA Sum of Squares df Mean Square F Sig. Between Groups 245 2 122.5 4.23 0.021 Within Groups 520 27 19.3 Total 765 29
Interpretation: p = 0.021 < 0.05, so at least one group is different.
Mini summary: Use One-Way ANOVA in SPSS.
Definition: After ANOVA, if the result is significant, a post-hoc test tells you which specific groups are different.
Why it is important: ANOVA only tells you that at least one group is different, not which one(s).
How to do it: In the One-Way ANOVA dialog, click "Post Hoc" and choose a test (e.g., Tukey).
Output interpretation: Look at the table for pairs of groups. If the p-value for a pair is < 0.05, those two groups are significantly different.
Mini summary: Post-hoc tests tell you which groups differ.
Let's say we have test scores from three schools in Nigeria: School A, School B, and School C. We want to know if there is a difference in performance.
Example conclusion: "There is a significant difference between schools (p = 0.02). School A scored higher than School B and School C."
Mini summary: ANOVA works for any Nigerian data.
| Number of Groups | Test to Use |
|---|---|
| 2 groups | t-test |
| 3 or more groups | ANOVA |
Mini summary: Choose t-test for two groups, ANOVA for three or more.
Definition: When you report results, you should include:
Example: "A t-test was used to compare the scores of boys and girls. The p-value was 0.018, which is less than 0.05. Therefore, we conclude that there is a significant difference between boys and girls."
Mini summary: Report your findings clearly.
Mini summary: Avoid common mistakes.
Definition: A result can be statistically significant (p < 0.05) but not practically significant (the difference is very small).
Why it is important: Always consider if the difference matters in real life.
Example: A t-test shows a significant difference in test scores of 0.5 points โ that might not be meaningful.
Mini summary: Think about whether the difference is practically important.
We learned to compare groups using t-tests and ANOVA. We use the p-value to decide if differences are real. We also learned to interpret output and report results.
Group A: **** Group B: ***
(mean = 75) (mean = 69)
t-test checks if the gap is real.
Gap (difference)
|--------------|
75 69
Group A: **** Group B: *** Group C: ***** ANOVA checks if any group differs.
| Feature | t-test | ANOVA |
|---|---|---|
| Number of groups | 2 | 3 or more |
| Example | Boys vs Girls | Class 5, 6, 7 |
| SPSS menu | Compare Means โ T Test | Compare Means โ One-Way ANOVA |
| Post-hoc needed? | No | Yes (if significant) |
In this module, we learned how to compare groups using SPSS. We used the t-test for two groups and ANOVA for three or more groups. We learned about the p-value and how it helps us decide if differences are real. We also learned to interpret output and report our results. Comparing groups is a key skill in data analysis.
| Test | Number of Groups |
|---|---|
| 1. t-test | A. 3 or more |
| 2. ANOVA | B. 2 |
Answers: 1-B, 2-A
In groups, collect data on a numeric variable (e.g., height, score) from two or three groups (e.g., boys/girls, classes). Enter it into SPSS, run the appropriate test, and present your findings.
Use the iris dataset. Run a t-test comparing petal length between two species. Then, run ANOVA comparing petal length among all three species. Interpret the results.
Title: "Compare Nigerian Schools"
Find or create a dataset of test scores from three Nigerian schools. Use ANOVA to compare them. If significant, run post-hoc tests. Write a report.
Using the mtcars dataset, run a t-test comparing mpg between automatic and manual cars. Then, run ANOVA comparing mpg across different numbers of cylinders (4, 6, 8).
Find a real dataset online (e.g., from Kaggle). Perform t-tests and ANOVA on variables of your choice. Write a detailed report with interpretations.
Fill-in-the-Blank: 1. t-test, 2. ANOVA, 3. p-value, 4. significant, 5. post-hoc, 6. Assumptions, 7. Analyze, 8. Compare Means, 9. not significant, 10. ANOVA.
True/False: 1T, 2F, 3T, 4F, 5F, 6T, 7F, 8T, 9F, 10T.
Multiple Choice: 1B, 2A, 3A, 4A, 5A, 6A, 7A, 8A, 9B, 10A, 11B, 12A, 13B, 14C, 15C.
In the next module, we will learn about relationships between variables โ like correlation and regression. We will see how variables change together and how to predict one from another. Great job comparing groups!
Hello, young data explorer! In Module 4, we learned how to compare groups using t-tests and ANOVA. We answered questions like "Are boys and girls different?" But what if we want to know if two numbers are related? For example, does more study time lead to higher test scores? Does taller mean heavier?
In this module, we will learn about correlation and regression. Correlation tells us if two variables move together (like more study time = higher scores). Regression helps us predict one variable from another (like predicting test scores from study time).
By the end of this module, you will be able to see relationships in your data and make predictions. You will be a data predictor!
Chidi noticed that pupils who studied more got higher scores. But he wanted to know if this was really true. He collected data on study time and test scores from 30 pupils. He wanted to see if the two variables were related.
He used SPSS to calculate the correlation. The correlation was 0.75, which is a strong positive relationship. This meant that as study time increased, scores tended to increase too. He then used regression to predict a pupil's score based on their study time.
Chidi's teacher was impressed. He learned that relationships help us understand and predict the world.
Definition: Correlation measures the strength and direction of a relationship between two numeric variables.
Why it is important: It tells us if variables change together (e.g., more study = higher scores).
Simple explanation: It's like seeing if two friends always walk together โ if one goes, the other goes too.
Types of correlation:
Real-life example: Height and weight are positively correlated.
School example: Study time and test scores are positively correlated.
Home example: Hours spent playing and screen time are positively correlated.
Nigerian example: Education level and income are positively correlated.
Illustration:
Positive: * * * * * (upward slope)
* * * *
* * *
* *
*
Negative: *
* *
* *
* *
* *
* * (downward slope)
Zero: * * *
* * *
* * *
* * * (no clear pattern)
Mini summary: Correlation shows if and how two variables move together.
Definition: The correlation coefficient (r) is a number between -1 and 1 that measures the strength of the correlation.
Why it is important: It gives you a single number to describe the relationship.
Interpretation:
Simple explanation: r is like a score from -1 to 1 that tells you how strong the friendship is between two variables.
Example: r = 0.75 means a strong positive relationship. r = -0.60 means a moderate negative relationship.
Mini summary: r measures the strength of correlation.
Definition: Pearson correlation is the most common correlation measure.
How to do it:
Output interpretation: Look at the Pearson Correlation (r) and the Sig. (2-tailed) โ the p-value.
Example output:
Study Time Score Pearson Correlation 1.000 0.750 Sig. (2-tailed) . 0.001 N 30 30
Interpretation: r = 0.75, p = 0.001 < 0.05, so the correlation is significant and positive.
Mini summary: Use Bivariate Correlations in SPSS.
When you run correlation, SPSS gives you:
How to decide:
Mini summary: Use r and p-value to interpret correlation.
Definition: A scatter plot shows the relationship between two variables visually.
Why it is important: It helps you see the pattern before you calculate r.
How to do it in SPSS:
Illustration:
Score
100 | .
80 | . .
60 | . .
40 | .
20 |.
0 |___._.___.___.___.___.___
0 2 4 6 8 10
Study Time
Mini summary: Scatter plots show relationships visually.
Definition: Just because two variables are correlated does not mean one causes the other.
Why it is important: It's a common mistake to assume causation from correlation.
Simple explanation: Ice cream sales and shark attacks are correlated, but eating ice cream doesn't cause shark attacks โ both are caused by hot weather.
Real-life example: There is a correlation between shoe size and reading ability, but that's because both increase with age.
School example: Test scores and height might be correlated โ taller pupils are usually older, not taller because they are smarter.
Home example: You might see a correlation between your screen time and your mood, but it might be because you watch more TV when you're tired.
Nigerian example: There might be a correlation between wealth and literacy, but wealth doesn't cause literacy โ both are influenced by many factors.
Mini summary: Correlation does not mean causation.
Definition: Regression is a statistical method that models the relationship between variables and allows you to make predictions.
Why it is important: It helps you predict one variable from another (or several others).
Simple explanation: It's like finding the line that best fits the data points, so you can guess where new points will be.
Real-life example: A store predicts sales based on advertising spend.
School example: A teacher predicts test scores based on study time.
Home example: You predict your allowance based on chores.
Nigerian example: A farmer predicts crop yield based on rainfall.
Illustration:
y (predicted) ^ | / (line: y = a + b*x) | / | / |/__________________ x (predictor)
Mini summary: Regression predicts one variable from another.
Definition: Simple linear regression uses one predictor variable to predict one outcome variable.
How to do it:
Output interpretation: Look at the coefficients and R-squared.
Example output:
Coefficients: (Intercept) = 50.0 Study Time = 5.0 R-squared = 0.5625
Interpretation: Score = 50 + 5 * Study Time. R-squared = 0.56 means 56% of the variance in scores is explained by study time.
Mini summary: Use Linear Regression in SPSS.
Key numbers:
Example: If the equation is Score = 50 + 5 * Study Time, then:
Mini summary: Coefficients define the line; R-squared measures fit.
Definition: Once you have the regression equation, you can predict the outcome for any value of the predictor.
How to do it:
Example: If Score = 50 + 5 * Study Time, and a pupil studies 6 hours, then predicted score = 50 + 5*6 = 80.
In SPSS: After running regression, click "Save" and select "Unstandardized Predicted Values" to get predictions in your data.
Mini summary: Use the equation to make predictions.
Definition: R-squared (Rยฒ) is the proportion of variance in the outcome that is explained by the predictor(s).
Why it is important: It tells you how well your model predicts.
Interpretation: Rยฒ = 0.56 means 56% of the variation in scores is explained by study time. The higher the better.
Rule of thumb:
Mini summary: R-squared measures model fit.
Let's use data from a Nigerian school to predict test scores from study time.
Example equation: Score = 45 + 6 * Study Time. For 4 hours: Score = 45 + 24 = 69.
Mini summary: Regression works with Nigerian data.
Definition: Multiple regression uses two or more predictors to predict an outcome.
Why it is important: It gives a more accurate prediction.
How to do it: In the Linear Regression dialog, move more than one variable to "Independent(s)".
Example: Predict Score from Study Time and Sleep Hours.
Equation: Score = a + b1*Study Time + b2*Sleep Hours.
Mini summary: Multiple regression uses multiple predictors.
Regression assumes:
How to check: Use the "Plots" button in SPSS to create residual plots.
Mini summary: Check assumptions for a valid model.
We learned that correlation measures relationships, and regression predicts outcomes. We used SPSS to calculate correlations and run regression models. We also learned to interpret output and make predictions.
Score
100 | .
80 | . .
60 | . .
40 | .
20 |.
0 |___._.___.___.___.___.___
0 2 4 6 8 10
Study Time
| Feature | Correlation | Regression |
|---|---|---|
| Purpose | Measure relationship | Predict outcomes |
| Output | r (coefficient) | Equation, R-squared |
| Symmetry | Symmetric (x,y same) | Asymmetric (y depends on x) |
| Use | Explore relationships | Make predictions |
In this module, we learned about correlation and regression. Correlation measures the strength and direction of a relationship between two variables. The correlation coefficient (r) tells us how strong the relationship is. Regression goes a step further โ it allows us to predict one variable from another. We learned to run these in SPSS and interpret the output. We also learned the important rule: correlation does not imply causation.
| Term | Description |
|---|---|
| 1. Correlation | A. Predicts one variable from another |
| 2. Regression | B. Measures relationship |
| 3. r | C. Coefficient of correlation |
| 4. R-squared | D. Proportion of variance explained |
| 5. Scatter plot | E. Visualises relationships |
Answers: 1-B, 2-A, 3-C, 4-D, 5-E
In groups, collect data on two numeric variables (e.g., height and weight, or study time and scores). Enter into SPSS, calculate correlation, and run regression. Present your findings.
Use the mtcars dataset. Calculate the correlation between hp and mpg. Run a regression predicting mpg from hp. Interpret the output.
Title: "Predicting School Performance"
Find or create a dataset of Nigerian students with study time and test scores. Run a correlation and regression. Use the model to make predictions for new students. Write a report.
Using the airquality dataset, calculate the correlation between temperature and ozone levels. Run a regression predicting ozone from temperature. Interpret the output.
Find a real dataset online (e.g., from Kaggle). Perform correlation and regression analysis on two or more variables. Write a detailed report with predictions.
Fill-in-the-Blank: 1. Correlation, 2. 1, 3. increase, 4. decreases, 5. Regression, 6. x, 7. R-squared, 8. causation, 9. scatter, 10. Correlate.
True/False: 1F, 2T, 3T, 4T, 5F, 6F, 7T, 8T, 9F, 10T.
Multiple Choice: 1B, 2B, 3A, 4B, 5B, 6B, 7A, 8B, 9B, 10B, 11C, 12A, 13B, 14A, 15A.
In the next module, we will learn about data communication โ how to present your findings in reports and presentations. You will learn to combine all your skills to tell a data story. Great job exploring relationships!
Hello, young data explorer! In Module 5, we learned about relationships and regression. We found patterns and made predictions. But what good is all that work if you cannot share it with others? This module is about data communication โ how to tell a clear and compelling story with your data.
Data communication is like being a storyteller. You take your data, your analysis, and your insights, and you present them in a way that is easy to understand. You use reports, slides, dashboards, and visualisations to share your findings.
By the end of this module, you will be able to create a data report in SPSS, present your findings, and make an impact. You will be a data communicator!
Chidi had analysed data on study time and test scores. He found a strong positive relationship. His teacher asked him to present his findings to the class. Chidi knew he had to communicate his results clearly so everyone could understand.
He used SPSS to create charts and tables. He wrote a simple report with a title, introduction, findings, and conclusions. He prepared a short presentation. When he presented, he showed his scatter plot and explained what it meant. The class understood and even asked good questions.
Chidi learned that good communication makes your hard work useful and appreciated.
Definition: Data communication is the process of sharing your data findings with others in a clear and effective way.
Why it is important: If you cannot communicate your results, your analysis has no impact.
Simple explanation: It's like telling a story โ you need a beginning, middle, and end.
Real-life example: A business analyst presents sales data to the CEO.
School example: A student presents a science project to the class.
Home example: You tell your family about your savings.
Nigerian example: A researcher presents findings on agriculture to farmers.
Illustration:
Data --> Analysis --> Insights --> Communication --> Impact
Mini summary: Data communication is sharing your findings to create impact.
Definition: Understand who you are communicating with and what they need to know.
Why it is important: Different audiences need different levels of detail.
Simple explanation: You talk differently to a friend than to a teacher.
Real-life example: You explain a game to a younger child differently than to a friend.
School example: You present your project to the teacher differently than to your classmates.
Home example: You explain your daily routine to a visitor.
Nigerian example: You present data to a community leader differently than to a government official.
Mini summary: Tailor your message to your audience.
A good report has:
Mini summary: Reports should be clear and structured.
Definition: The Output Viewer is where SPSS shows your results (tables and charts).
Why it is important: This is where you see your analysis results.
How to use it:
Mini summary: The Output Viewer shows your results.
Definition: You can edit charts to make them clearer and more attractive.
Why it is important: A good chart tells a story quickly.
How to edit:
Example: You can change the colour of bars, add a title, or change the axis labels.
Mini summary: You can edit charts to make them better.
Definition: Exporting means saving your output as a file you can share.
Why it is important: You can include it in reports or presentations.
How to export:
Tip: Export as Word or PDF for reports, Excel for data tables.
Mini summary: Export output to share your results.
Definition: You can copy tables and charts from the Output Viewer and paste them into other documents.
Why it is important: It's a quick way to include results in reports.
How to do it:
Mini summary: Copy and paste to include results in other documents.
Definition: Good charts are easy to read and understand.
Why it is important: A clear chart conveys your message quickly.
Tips:
Example: A bar chart with "Subject" on the x-axis, "Count" on the y-axis, and a title like "Favourite Subjects".
Mini summary: Clear charts make your data easy to understand.
Definition: A narrative is the story you tell about your data. It connects the data to the real world.
Why it is important: It makes your report engaging and meaningful.
Simple explanation: Instead of just showing numbers, explain what they mean.
Example: "The data shows that students who studied more got higher scores. This suggests that study time is important for success."
Mini summary: A narrative makes data meaningful.
Definition: Tables organise and present numeric data clearly.
Why it is important: They provide detailed information in a structured way.
How to use: Include tables for means, frequencies, and other statistics.
Example:
+-------+--------+--------+ | Group | Mean | N | +-------+--------+--------+ | Boys | 75.0 | 20 | | Girls | 82.0 | 20 | +-------+--------+--------+
Mini summary: Tables present data in a structured way.
Definition: A presentation is a way to share your findings in a live setting.
Why it is important: It allows you to explain your work and answer questions.
Tips:
Mini summary: Presentations help you share your findings live.
Definition: You can create a report by saving your Output Viewer and adding narrative.
How to do it:
Mini summary: Combine SPSS output with narrative for a complete report.
Let's create a report on Nigerian education data.
Mini summary: You can create reports on Nigerian data.
Mini summary: Avoid common mistakes in communication.
We learned to communicate data effectively through reports, visualisations, and narratives. Good communication makes data valuable and impactful.
+-----------------------+ | Title | +-----------------------+ | Introduction | +-----------------------+ | Data | +-----------------------+ | Analysis | +-----------------------+ | Visualisations | +-----------------------+ | Conclusions | +-----------------------+
| Format | Use |
|---|---|
| Word (.docx) | Reports, documents |
| Excel (.xlsx) | Data tables |
| PDF (.pdf) | Sharing, printing |
| HTML | Web pages |
In this module, we learned about data communication โ how to share your data findings effectively. We used the SPSS Output Viewer to view and edit results, exported output to Word and PDF, and created clear charts and tables. We also learned about writing a narrative and tailoring our message to our audience. Good communication makes data analysis valuable and impactful.
| Term | Description |
|---|---|
| 1. Output Viewer | A. Story behind data |
| 2. Export | B. Results window |
| 3. Narrative | C. Save as file |
| 4. Visualisation | D. Chart or graph |
| 5. Report | E. Document with findings |
Answers: 1-B, 2-C, 3-A, 4-D, 5-E
In groups, analyse a dataset (e.g., the sample data in SPSS). Create a report in Word that includes output from SPSS, charts, and a narrative. Present your report to the class.
Create a report on a topic of your choice (e.g., your hobbies, school data). Use SPSS to analyse the data, export the output, and write a narrative. Share your report.
Title: "Nigerian Data Report"
Find a dataset about Nigeria (e.g., education, health, agriculture). Perform a complete analysis: import, clean, explore, and test. Create a comprehensive report with charts and narrative.
Using the economics dataset (available in SPSS or import), create a report that shows trends in unemployment, population, and GDP. Include visualisations, summaries, and a narrative. Export to Word.
Find a complex dataset online. Create a dashboard or report that tells a story about the data. Include multiple visualisations and a clear narrative. Present your work.
Fill-in-the-Blank: 1. Data communication, 2. Output Viewer, 3. Export, 4. narrative, 5. visualisation, 6. report, 7. audience, 8. Clear, 9. edit, 10. Tailor.
True/False: 1F, 2T, 3T, 4F, 5T, 6F, 7T, 8F, 9F, 10F.
Multiple Choice: 1B, 2A, 3B, 4B, 5B, 6B, 7B, 8B, 9B, 10A, 11B, 12A, 13B, 14A, 15B.
In the next module, we will bring everything together in a capstone project. You will apply all the skills you have learned to a real-world data analysis project. Start thinking about a dataset and a question you want to answer.
Hello, young data explorer! You have come a long way. You have learned how to clean data, explore it, compare groups, find relationships, and communicate your findings. Now it's time to bring everything together in a capstone project.
A capstone project is a final project that shows everything you have learned. You will choose a dataset, ask a question, analyse the data, and present your findings. It's like building a complete house after learning how to lay bricks, install windows, and paint walls.
In this module, we will guide you through the process step by step. You will apply all the skills from Modules 1 to 6. By the end, you will have a complete project that you can be proud of.
Chidi had learned all about SPSS. Now, his teacher asked him to do a capstone project. He had to choose a topic, collect data, analyse it, and present his findings. He was nervous but excited.
He chose to study the relationship between study time and test scores at his school. He collected data from 50 pupils. He cleaned the data, explored it, ran a correlation, and did a regression. He made charts and wrote a report. Finally, he presented his findings to the class.
Everyone was impressed. Chidi felt proud of his work. He realised that he had become a real data analyst.
Definition: A capstone project is a final project that demonstrates all the skills you have learned.
Why it is important: It shows you can apply your knowledge to a real problem.
Simple explanation: It's like a final exam, but more fun and creative.
Real-life example: An architect builds a model house as a final project.
School example: A student does a science fair project.
Home example: You build a birdhouse after learning woodworking.
Nigerian example: A researcher publishes a study on Nigerian agriculture.
Illustration:
Step 1: Choose Topic Step 2: Collect Data Step 3: Clean Data Step 4: Explore Data Step 5: Analyse Data Step 6: Report Findings
Mini summary: A capstone project shows everything you have learned.
Definition: Your topic is the subject you want to study. Your question is what you want to find out.
Why it is important: A good question guides your analysis.
Simple explanation: It's like choosing a destination before you start a journey.
Examples of questions:
Nigerian example: "Is there a relationship between education level and income in Nigeria?"
Mini summary: Choose a clear topic and question.
Definition: Data can come from surveys, experiments, or existing datasets.
Why it is important: You need data to answer your question.
Simple explanation: It's like gathering ingredients before cooking.
Where to find data:
Nigerian example: Download data from the Nigerian National Bureau of Statistics.
Mini summary: Find a dataset that fits your question.
Definition: Importing means bringing your data into SPSS.
How to do it:
Tip: If you have an Excel file, make sure the first row has column names.
Mini summary: Import your data into SPSS.
Definition: Cleaning means fixing errors, missing values, and duplicates.
Why it is important: Clean data leads to correct results.
Steps:
Mini summary: Clean your data before analysis.
Definition: Exploration means getting to know your data.
Steps:
Mini summary: Explore your data to understand it.
Definition: Compare groups to see if they differ.
Steps:
Mini summary: Use t-tests and ANOVA to compare groups.
Definition: Examine if variables are related.
Steps:
Mini summary: Use correlation and regression to examine relationships.
Definition: Interpretation means explaining what the numbers mean.
Steps:
Example: "There is a significant positive correlation between study time and test scores (r = 0.75, p < 0.05)."
Mini summary: Interpret your results clearly.
Definition: A report presents your findings in a structured way.
Structure:
Mini summary: Write a clear report with all sections.
Definition: Charts and tables make your results easy to understand.
Tips:
Mini summary: Use charts and tables to show your results.
Definition: Presenting means sharing your work with others.
Tips:
Mini summary: Present your project confidently.
Let's walk through a complete project on Nigerian education.
Mini summary: Apply all steps to Nigerian data.
Mini summary: Avoid common mistakes.
You have completed your capstone project! You have shown that you can handle the entire data analysis process. Take a moment to reflect on what you have learned and how you have grown.
Topic --> Data --> Clean --> Explore --> Analyze --> Report --> Present
| Phase | Activities | SPSS Tools |
|---|---|---|
| Data | Import, clean | File โ Open, Transform |
| Exploration | Descriptive stats, charts | Analyze, Graphs |
| Analysis | t-tests, ANOVA, correlation, regression | Compare Means, Correlate, Regression |
| Communication | Report, presentation | Output Viewer, Export |
In this final module, you completed your capstone project. You chose a topic, collected data, cleaned it, explored it, analysed it, and communicated your findings. You applied all the skills from Modules 1 to 6. You have shown that you are a capable data analyst. Congratulations!
| Phase | Activity |
|---|---|
| 1. Data | A. Charts and stats |
| 2. Explore | B. Import and clean |
| 3. Analyze | C. Report and present |
| 4. Communicate | D. t-tests, regression |
Answers: 1-B, 2-A, 3-D, 4-C
In groups, choose a topic, collect data, and complete a capstone project together. Each group member should have a role. Present your project to the class.
Choose a topic of your choice. Complete a capstone project from start to finish. Submit your report and present your findings.
Title: "My First Data Analysis Project"
Complete a full project on a topic of your choice. Use all the skills you have learned. Submit your report and present your findings.
Using the iris dataset, complete a project: clean the data, explore it, compare groups (species), and examine relationships (petal length vs petal width). Write a report.
Find a real dataset from Nigeria. Complete a comprehensive project and present your findings. Try to make a real-world recommendation based on your results.
Fill-in-the-Blank: 1. capstone, 2. research question, 3. Cleaning, 4. Exploration, 5. t-test, 6. ANOVA, 7. Correlation, 8. Regression, 9. report, 10. presentation.
True/False: 1F, 2T, 3F, 4T, 5F, 6F, 7T, 8F, 9F, 10T.
Multiple Choice: 1B, 2B, 3B, 4A, 5B, 6B, 7B, 8B, 9B, 10B, 11A, 12A, 13B, 14B, 15D.
Congratulations on completing the SPSS for Data Analysis course! You have gained valuable skills. To continue your journey, you can:
You have done an excellent job. Keep exploring, keep asking questions, and keep analysing data!