Paath.online blog

Pandas Tutorial for Beginners: Analyze a Dataset Step by Step

By Mohit Agarwal, Paath.online9 min read

Learn Pandas by working through a small student-marks dataset. You’ll load a CSV, inspect and filter rows, deal with missing values, group results, and turn a table into clear findings—one practical step at a time.

What is Pandas?

Pandas is a Python library for working with structured data, such as spreadsheets and CSV files. Its main table-like object is called a DataFrame. Each column can have a name, and you can select, filter, clean, and summarize rows using concise Python code.

This tutorial uses a tiny fictional marks dataset. The same steps apply to many beginner projects, such as analyzing expenses, survey answers, product lists, or sports results.

1. Install and import Pandas

Install Pandas in your active Python environment, then import it using the common alias pd:

python -m pip install pandas

import pandas as pd

In a notebook, run the install command in its own cell if needed, then import Pandas in the next cell.

2. Load a CSV file and inspect it

Imagine a CSV file named marks.csv with columns for student, subject, and score. Pandas can read it into a DataFrame:

import pandas as pd

df = pd.read_csv("marks.csv")
print(df.head())
print(df.shape)
print(df.info())

head() shows the first few rows, shape gives the row and column counts, and info() shows column names, data types, and non-missing values. Inspecting first helps catch unexpected column names or empty fields early.

3. Create a small DataFrame to follow along

If you do not have a CSV file yet, create this example directly:

import pandas as pd

df = pd.DataFrame({
    "student": ["Asha", "Ravi", "Meera", "Kabir", "Zoya"],
    "subject": ["Math", "Math", "Science", "Science", "Math"],
    "score": [88, 72, 95, None, 81],
})

print(df)

None represents a missing score. Real datasets often contain blanks, so learning to find them is an important part of data analysis.

4. Select columns and filter rows

Select one column by its name, or filter the table with a condition:

# Select one column
print(df["student"])

# Keep rows with a score of at least 80
high_scores = df[df["score"] >= 80]
print(high_scores)

The condition creates a True/False check for each row; Pandas keeps the rows where it is True. You can combine conditions with parentheses and & (and) or | (or).

5. Find and handle missing values

First count missing cells in each column. Then choose a treatment that makes sense for the data and the question:

print(df.isna().sum())

# For this simple example, remove rows without a score
scored = df.dropna(subset=["score"])

# Alternatively, fill missing scores with a chosen value
filled = df.copy()
filled["score"] = filled["score"].fillna(0)

Dropping a row or filling a blank with zero can change your conclusion. Use a method justified by what the missing value means; zero is not automatically a safe replacement.

6. Calculate summaries and group results

Use summary methods to answer simple questions. Grouping lets you compare categories, such as subjects:

print(scored["score"].mean())
print(scored["score"].max())

average_by_subject = scored.groupby("subject")["score"].mean()
print(average_by_subject)

Here the mean gives the average among available scores. Before sharing a result, check how many rows contributed to it and whether missing values could affect the comparison.

7. Save your cleaned data

When you want to reuse or share the cleaned table, export it to a new CSV:

scored.to_csv("cleaned_marks.csv", index=False)

Setting index=False avoids adding the DataFrame row index as an extra column in the exported file.

A quick Pandas workflow to remember

  1. Load the data and check its shape, columns, and data types.
  2. Inspect a few rows and count missing values.
  3. Select, filter, or clean the rows needed for your question.
  4. Group and summarize, then explain what the result does and does not show.
  5. Save a cleaned copy when you need to reuse the result.

Continue with the NumPy and Pandas 20–22 session roadmap for a structured sequence, or review our Python roadmap for beginners if you are still getting comfortable with Python. The official Pandas getting-started tutorials are a useful reference for these operations.

What to learn next

Once you can explore and clean tabular data, try a small project: analyze a public dataset, state one question, show the code you used, and summarize the evidence. Later, these skills will help when you prepare data for a beginner machine-learning project.

Want help learning Python and data analysis step by step?

Build your Python, Pandas, and data skills with clear explanations and live 1:1 guidance.

Frequently asked questions

What is Pandas used for in Python?▾

Pandas helps Python users work with labelled, tabular data. Its DataFrame structure makes it easier to read CSV files, select and filter rows, handle missing values, group records, and calculate summaries.

Can I learn Pandas without NumPy?▾

Yes. You can begin using Pandas for common data-analysis tasks without studying NumPy first. Learning basic Python and lists is useful, and NumPy becomes helpful as you explore numerical arrays and more advanced analysis.

What should I learn before Pandas?▾

Start with Python variables, lists, dictionaries, functions, and importing packages. You do not need advanced Python or machine-learning knowledge to follow a beginner Pandas tutorial.

How do I install Pandas?▾

In a Python environment, install it with pip install pandas. In a notebook, you can run the installation command in a cell if Pandas is not already available.

Want hands-on help? Explore our Python classes and AI classes for beginners.

About the instructor

Mohit Agarwal teaches live Python and AI classes at Paath.online. Sessions focus on beginners and students: clear explanations, debugging practice, and project-based learning for school, university, and career goals.

Instruction is available in English or Hindi. Topics include Python fundamentals, NumPy & Pandas, machine learning basics, RAG, and applied AI workflows.