## Python for Data Analysis — Complete Guide for Beginners 2026
Python has become the dominant language for data analysis globally. With three libraries — NumPy, Pandas, and Matplotlib — you can clean, analyze, and visualize virtually any dataset. This guide explains each library and how to use them together for real data analysis projects.
---
Why Python for Data Analysis
Before Excel, data analysis meant manual work. After Excel, analysts could process thousands of rows. With Python, you can process millions of rows, automate repetitive analysis, build reproducible workflows, and extend into machine learning when ready.
Python vs Excel for data analysis:
| Task | Excel | Python (Pandas) |
|---|---|---|
| Dataset size | ~1M rows max | Unlimited (RAM-bound) |
| Repetition | Manual every time | Automate with script |
| Complex transformations | Difficult, error-prone | Clean, readable code |
| Machine learning integration | Plugins/add-ons | Native Python ML libraries |
| Version control | Difficult | Git-compatible |
For datasets under 50,000 rows and simple analysis: Excel is fine. For larger datasets, complex transformations, automation, or any machine learning: Python wins.
---
Setting Up Python for Data Analysis
Step 1 — Install Anaconda: Anaconda is a Python distribution that comes with NumPy, Pandas, Matplotlib, and Jupyter Notebook pre-installed. Download free from anaconda.com.
Step 2 — Open Jupyter Notebook: Jupyter Notebook is an interactive environment where you write code in "cells" and see output immediately below each cell. Perfect for data analysis and exploration.
Open Anaconda Navigator → Launch Jupyter Notebook → browser opens → create new notebook.
Step 3 — Import libraries: Every data analysis notebook starts with these imports:
import numpy as np import pandas as pd import matplotlib.pyplot as plt import seaborn as sns
---
NumPy — Numerical Computing Foundation
NumPy (Numerical Python) is the foundation of all Python data science. It provides arrays and mathematical functions.
Why NumPy matters: Python's built-in lists are slow for numerical operations. NumPy arrays (ndarray) are 10–100x faster for mathematical operations on large datasets.
Core NumPy operations:
Creating arrays: a = np.array([1, 2, 3, 4, 5])
Creating ranges: np.arange(0, 10, 2) → array([0, 2, 4, 6, 8])
Statistical operations: np.mean([10, 20, 30, 40]) → 25.0 np.std([10, 20, 30, 40]) → 11.18 np.max([10, 20, 30, 40]) → 40 np.min([10, 20, 30, 40]) → 10
Array operations (element-wise, much faster than loops): a = np.array([1, 2, 3]) b = np.array([4, 5, 6]) a + b → array([5, 7, 9]) a * b → array([4, 10, 18])
2D arrays (matrices): matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]]) matrix.shape → (3, 3) matrix.T → transpose (flip rows and columns)
In practice: You rarely use NumPy directly for data analysis — Pandas uses NumPy under the hood. But understanding NumPy makes you a better Pandas user.
---
Pandas — The Core of Python Data Analysis
Pandas is what makes Python practical for data analysis. It provides two main structures:
Series: One-dimensional labeled array (like an Excel column) DataFrame: Two-dimensional labeled data structure (like an Excel spreadsheet)
Loading data:
From CSV: df = pd.read_csv('sales_data.csv')
From Excel: df = pd.read_excel('report.xlsx', sheet_name='Sheet1')
First steps with any new dataset:
df.head() # First 5 rows df.tail() # Last 5 rows df.shape # (rows, columns) tuple df.info() # Column names, data types, missing values df.describe() # Statistical summary (mean, std, min, max, quartiles)
Selecting data:
Select column: df['column_name'] Select multiple columns: df[['col1', 'col2']] Filter rows: df[df['age'] > 25] Multiple conditions: df[(df['city'] == 'Amritsar') and (df['age'] > 25)]
Cleaning data (most common tasks):
Check missing values: df.isnull().sum() Drop rows with missing values: df.dropna() Fill missing values: df['column'].fillna(0) or df['column'].fillna(df['column'].mean()) Remove duplicates: df.drop_duplicates() Rename columns: df.rename(columns={'old_name': 'new_name'}) Change data type: df['date'] = pd.to_datetime(df['date'])
Grouping and aggregation (the most important Pandas skill):
Group sales by city: df.groupby('city')['revenue'].sum()
Group and aggregate multiple metrics: df.groupby('product').agg({'revenue': 'sum', 'units': 'mean', 'orders': 'count'})
Sorting: df.sort_values('revenue', ascending=False)
Creating new columns: df['revenue_per_unit'] = df['revenue'] / df['units']
Merging datasets (like VLOOKUP in Excel, but better): merged = pd.merge(orders_df, customers_df, on='customer_id', how='left')
---
Matplotlib and Seaborn — Data Visualization
Matplotlib is the foundational plotting library. Seaborn is built on Matplotlib and makes statistical charts beautiful with less code.
Basic charts with Matplotlib:
Line chart: plt.figure(figsize=(10, 6)) plt.plot(df['month'], df['sales']) plt.title('Monthly Sales') plt.xlabel('Month') plt.ylabel('Sales (₹)') plt.show()
Bar chart: plt.bar(df['city'], df['revenue']) plt.title('Revenue by City') plt.show()
Histogram (distribution of a column): plt.hist(df['age'], bins=20) plt.title('Age Distribution') plt.show()
Scatter plot: plt.scatter(df['marketing_spend'], df['revenue']) plt.title('Marketing Spend vs Revenue') plt.show()
Seaborn (better defaults, statistical charts):
Box plot (shows distribution + outliers): sns.boxplot(x='city', y='revenue', data=df) plt.show()
Heatmap (correlation between variables): correlation = df.corr() sns.heatmap(correlation, annot=True, cmap='coolwarm') plt.show()
Count plot (categorical frequency): sns.countplot(x='category', data=df) plt.show()
---
A Complete Data Analysis Mini-Project
Here's the workflow for a complete analysis on a sample sales dataset:
Step 1 — Load and Inspect: df = pd.read_csv('sales.csv') df.head() df.info() df.describe()
Step 2 — Clean: df.isnull().sum() # Check missing values df = df.dropna() # Remove if < 5% missing df['date'] = pd.to_datetime(df['date'])
Step 3 — Explore: total_revenue = df['revenue'].sum() avg_order_value = df['revenue'].mean() top_products = df.groupby('product')['revenue'].sum().sort_values(ascending=False).head(10) revenue_by_city = df.groupby('city')['revenue'].sum()
Step 4 — Visualize: top_products.plot(kind='bar', figsize=(12, 6), title='Top 10 Products by Revenue') plt.tight_layout() plt.show()
revenue_by_city.plot(kind='pie', autopct='%1.1f%%', title='Revenue Share by City') plt.show()
Step 5 — Insights: This analysis reveals which products drive the most revenue and which cities have the highest sales concentration — actionable insights for a business decision.
---
What You Can Do After Learning These Three Libraries
With NumPy, Pandas, and Matplotlib, you can: - Analyze sales data and build monthly dashboards - Identify top-performing products, channels, or regions - Clean messy raw data for analysis or ML input - Automate weekly/monthly reports that previously took hours manually - Build Kaggle submissions for data science competitions - Prepare for Machine Learning (Scikit-learn uses Pandas DataFrames as input)
Portfolio projects using Pandas + Matplotlib:
- •. E-commerce sales analysis (public Kaggle datasets)
- •. COVID-19 India data visualization (open government data)
- •. Indian cricket IPL performance analysis
- •. Stock price analysis (Yahoo Finance data)
- •. Real estate price analysis (Punjab property listings)
Each project is a GitHub portfolio piece that demonstrates your data analysis skills to employers.
---
Data Science Course in Amritsar →