Programming

Python for Data Analysis — Complete Guide for Beginners 2026 (NumPy, Pandas, Matplotlib)

MITS Faculty 9 min read

Complete Python data analysis guide for beginners in India 2026. How to use NumPy, Pandas, and Matplotlib for data analysis from scratch — with practical examples for cleaning data, exploring datasets, and creating charts. For students at MITS Academy Amritsar and Jalandhar.

## Python for Data Analysis — Complete Guide for Beginners 2026

Python has become the dominant language for data analysis globally. With three libraries — NumPy, Pandas, and Matplotlib — you can clean, analyze, and visualize virtually any dataset. This guide explains each library and how to use them together for real data analysis projects.

---

Why Python for Data Analysis

Before Excel, data analysis meant manual work. After Excel, analysts could process thousands of rows. With Python, you can process millions of rows, automate repetitive analysis, build reproducible workflows, and extend into machine learning when ready.

Python vs Excel for data analysis:

TaskExcelPython (Pandas)
Dataset size~1M rows maxUnlimited (RAM-bound)
RepetitionManual every timeAutomate with script
Complex transformationsDifficult, error-proneClean, readable code
Machine learning integrationPlugins/add-onsNative Python ML libraries
Version controlDifficultGit-compatible

For datasets under 50,000 rows and simple analysis: Excel is fine. For larger datasets, complex transformations, automation, or any machine learning: Python wins.

---

Setting Up Python for Data Analysis

Step 1 — Install Anaconda: Anaconda is a Python distribution that comes with NumPy, Pandas, Matplotlib, and Jupyter Notebook pre-installed. Download free from anaconda.com.

Step 2 — Open Jupyter Notebook: Jupyter Notebook is an interactive environment where you write code in "cells" and see output immediately below each cell. Perfect for data analysis and exploration.

Open Anaconda Navigator → Launch Jupyter Notebook → browser opens → create new notebook.

Step 3 — Import libraries: Every data analysis notebook starts with these imports:

import numpy as np import pandas as pd import matplotlib.pyplot as plt import seaborn as sns

---

NumPy — Numerical Computing Foundation

NumPy (Numerical Python) is the foundation of all Python data science. It provides arrays and mathematical functions.

Why NumPy matters: Python's built-in lists are slow for numerical operations. NumPy arrays (ndarray) are 10–100x faster for mathematical operations on large datasets.

Core NumPy operations:

Creating arrays: a = np.array([1, 2, 3, 4, 5])

Creating ranges: np.arange(0, 10, 2) → array([0, 2, 4, 6, 8])

Statistical operations: np.mean([10, 20, 30, 40]) → 25.0 np.std([10, 20, 30, 40]) → 11.18 np.max([10, 20, 30, 40]) → 40 np.min([10, 20, 30, 40]) → 10

Array operations (element-wise, much faster than loops): a = np.array([1, 2, 3]) b = np.array([4, 5, 6]) a + b → array([5, 7, 9]) a * b → array([4, 10, 18])

2D arrays (matrices): matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]]) matrix.shape → (3, 3) matrix.T → transpose (flip rows and columns)

In practice: You rarely use NumPy directly for data analysis — Pandas uses NumPy under the hood. But understanding NumPy makes you a better Pandas user.

---

Pandas — The Core of Python Data Analysis

Pandas is what makes Python practical for data analysis. It provides two main structures:

Series: One-dimensional labeled array (like an Excel column) DataFrame: Two-dimensional labeled data structure (like an Excel spreadsheet)

Loading data:

From CSV: df = pd.read_csv('sales_data.csv')

From Excel: df = pd.read_excel('report.xlsx', sheet_name='Sheet1')

First steps with any new dataset:

df.head() # First 5 rows df.tail() # Last 5 rows df.shape # (rows, columns) tuple df.info() # Column names, data types, missing values df.describe() # Statistical summary (mean, std, min, max, quartiles)

Selecting data:

Select column: df['column_name'] Select multiple columns: df[['col1', 'col2']] Filter rows: df[df['age'] > 25] Multiple conditions: df[(df['city'] == 'Amritsar') and (df['age'] > 25)]

Cleaning data (most common tasks):

Check missing values: df.isnull().sum() Drop rows with missing values: df.dropna() Fill missing values: df['column'].fillna(0) or df['column'].fillna(df['column'].mean()) Remove duplicates: df.drop_duplicates() Rename columns: df.rename(columns={'old_name': 'new_name'}) Change data type: df['date'] = pd.to_datetime(df['date'])

Grouping and aggregation (the most important Pandas skill):

Group sales by city: df.groupby('city')['revenue'].sum()

Group and aggregate multiple metrics: df.groupby('product').agg({'revenue': 'sum', 'units': 'mean', 'orders': 'count'})

Sorting: df.sort_values('revenue', ascending=False)

Creating new columns: df['revenue_per_unit'] = df['revenue'] / df['units']

Merging datasets (like VLOOKUP in Excel, but better): merged = pd.merge(orders_df, customers_df, on='customer_id', how='left')

---

Matplotlib and Seaborn — Data Visualization

Matplotlib is the foundational plotting library. Seaborn is built on Matplotlib and makes statistical charts beautiful with less code.

Basic charts with Matplotlib:

Line chart: plt.figure(figsize=(10, 6)) plt.plot(df['month'], df['sales']) plt.title('Monthly Sales') plt.xlabel('Month') plt.ylabel('Sales (₹)') plt.show()

Bar chart: plt.bar(df['city'], df['revenue']) plt.title('Revenue by City') plt.show()

Histogram (distribution of a column): plt.hist(df['age'], bins=20) plt.title('Age Distribution') plt.show()

Scatter plot: plt.scatter(df['marketing_spend'], df['revenue']) plt.title('Marketing Spend vs Revenue') plt.show()

Seaborn (better defaults, statistical charts):

Box plot (shows distribution + outliers): sns.boxplot(x='city', y='revenue', data=df) plt.show()

Heatmap (correlation between variables): correlation = df.corr() sns.heatmap(correlation, annot=True, cmap='coolwarm') plt.show()

Count plot (categorical frequency): sns.countplot(x='category', data=df) plt.show()

---

A Complete Data Analysis Mini-Project

Here's the workflow for a complete analysis on a sample sales dataset:

Step 1 — Load and Inspect: df = pd.read_csv('sales.csv') df.head() df.info() df.describe()

Step 2 — Clean: df.isnull().sum() # Check missing values df = df.dropna() # Remove if < 5% missing df['date'] = pd.to_datetime(df['date'])

Step 3 — Explore: total_revenue = df['revenue'].sum() avg_order_value = df['revenue'].mean() top_products = df.groupby('product')['revenue'].sum().sort_values(ascending=False).head(10) revenue_by_city = df.groupby('city')['revenue'].sum()

Step 4 — Visualize: top_products.plot(kind='bar', figsize=(12, 6), title='Top 10 Products by Revenue') plt.tight_layout() plt.show()

revenue_by_city.plot(kind='pie', autopct='%1.1f%%', title='Revenue Share by City') plt.show()

Step 5 — Insights: This analysis reveals which products drive the most revenue and which cities have the highest sales concentration — actionable insights for a business decision.

---

What You Can Do After Learning These Three Libraries

With NumPy, Pandas, and Matplotlib, you can: - Analyze sales data and build monthly dashboards - Identify top-performing products, channels, or regions - Clean messy raw data for analysis or ML input - Automate weekly/monthly reports that previously took hours manually - Build Kaggle submissions for data science competitions - Prepare for Machine Learning (Scikit-learn uses Pandas DataFrames as input)

Portfolio projects using Pandas + Matplotlib:

  • •. E-commerce sales analysis (public Kaggle datasets)
  • •. COVID-19 India data visualization (open government data)
  • •. Indian cricket IPL performance analysis
  • •. Stock price analysis (Yahoo Finance data)
  • •. Real estate price analysis (Punjab property listings)

Each project is a GitHub portfolio piece that demonstrates your data analysis skills to employers.

---

Data Science Course in Amritsar →

Data Science Course in Jalandhar →

Book Free Demo →

Written by

MITS Faculty

Part of the MITS Academy faculty — an ISO 9001:2015 certified IT training institute in Amritsar, Jalandhar and Ludhiana that has placed 2,000+ students across TCS, Infosys, Wipro, HCL, Amazon and Accenture since 2014. Posts in Programming draw from the team's hands-on classroom and placement experience.

Free · No Commitment

Interested in Learning This?

Get a free demo class & career counselling — our expert will call you

Deep Dive

Explore MITS Academy's full Python Academy — curriculum, fees, placements & more

View Full Programme