Paath.online blog
Pandas Tutorial for Beginners: Analyze a Dataset Step by Step
By Mohit Agarwal, Paath.online9 min read
Learn Pandas by working through a small student-marks dataset. You’ll load a CSV, inspect and filter rows, deal with missing values, group results, and turn a table into clear findings—one practical step at a time.
What is Pandas?
Pandas is a Python library for working with structured data, such as spreadsheets and CSV files. Its main table-like object is called a DataFrame. Each column can have a name, and you can select, filter, clean, and summarize rows using concise Python code.
This tutorial uses a tiny fictional marks dataset. The same steps apply to many beginner projects, such as analyzing expenses, survey answers, product lists, or sports results.
1. Install and import Pandas
Install Pandas in your active Python environment, then import it using the common alias pd:
python -m pip install pandas import pandas as pd
In a notebook, run the install command in its own cell if needed, then import Pandas in the next cell.
2. Load a CSV file and inspect it
Imagine a CSV file named marks.csv with columns for student, subject, and score. Pandas can read it into a DataFrame:
import pandas as pd
df = pd.read_csv("marks.csv")
print(df.head())
print(df.shape)
print(df.info())head() shows the first few rows, shape gives the row and column counts, and info() shows column names, data types, and non-missing values. Inspecting first helps catch unexpected column names or empty fields early.
3. Create a small DataFrame to follow along
If you do not have a CSV file yet, create this example directly:
import pandas as pd
df = pd.DataFrame({
"student": ["Asha", "Ravi", "Meera", "Kabir", "Zoya"],
"subject": ["Math", "Math", "Science", "Science", "Math"],
"score": [88, 72, 95, None, 81],
})
print(df)None represents a missing score. Real datasets often contain blanks, so learning to find them is an important part of data analysis.
4. Select columns and filter rows
Select one column by its name, or filter the table with a condition:
# Select one column print(df["student"]) # Keep rows with a score of at least 80 high_scores = df[df["score"] >= 80] print(high_scores)
The condition creates a True/False check for each row; Pandas keeps the rows where it is True. You can combine conditions with parentheses and & (and) or | (or).
5. Find and handle missing values
First count missing cells in each column. Then choose a treatment that makes sense for the data and the question:
print(df.isna().sum()) # For this simple example, remove rows without a score scored = df.dropna(subset=["score"]) # Alternatively, fill missing scores with a chosen value filled = df.copy() filled["score"] = filled["score"].fillna(0)
Dropping a row or filling a blank with zero can change your conclusion. Use a method justified by what the missing value means; zero is not automatically a safe replacement.
6. Calculate summaries and group results
Use summary methods to answer simple questions. Grouping lets you compare categories, such as subjects:
print(scored["score"].mean())
print(scored["score"].max())
average_by_subject = scored.groupby("subject")["score"].mean()
print(average_by_subject)Here the mean gives the average among available scores. Before sharing a result, check how many rows contributed to it and whether missing values could affect the comparison.
7. Save your cleaned data
When you want to reuse or share the cleaned table, export it to a new CSV:
scored.to_csv("cleaned_marks.csv", index=False)Setting index=False avoids adding the DataFrame row index as an extra column in the exported file.
A quick Pandas workflow to remember
- Load the data and check its shape, columns, and data types.
- Inspect a few rows and count missing values.
- Select, filter, or clean the rows needed for your question.
- Group and summarize, then explain what the result does and does not show.
- Save a cleaned copy when you need to reuse the result.
Continue with the NumPy and Pandas 20–22 session roadmap for a structured sequence, or review our Python roadmap for beginners if you are still getting comfortable with Python. The official Pandas getting-started tutorials are a useful reference for these operations.
What to learn next
Once you can explore and clean tabular data, try a small project: analyze a public dataset, state one question, show the code you used, and summarize the evidence. Later, these skills will help when you prepare data for a beginner machine-learning project.