Data Engineering
Write SQL, build Python data pipelines, and process big data with Spark
You master SQL from joins and subqueries to CTEs and views, then build Python pipelines through real projects like a file-format converter and a file-to-database loader. The course moves into big data with Spark SQL on Databricks over GCP, covering Delta tables, transformations, joins and aggregations at scale.
per level · 3 levels · complete programme ₹33,000
- Duration
- 8 mo
Fees by level
Start at any level, or take the complete programme. The fee you pay for a level is locked for you.
| Level | What it covers | Duration | Fee |
|---|---|---|---|
| Beginner | Start from zero | 2 mo | ₹8,000 |
| Intermediate | Build working projects | 3 mo | ₹11,000 |
| Advanced | Get job-ready | 3 mo | ₹14,000 |
| Complete programme (all levels) | ₹33,000 | ||
All fees are in Indian Rupees and include applicable taxes. See our pricing & payment terms.
The complete curriculum
This is the entire syllabus — all 47 modules and 259 lessons, in the order you'll learn them. Nothing hidden.
↓ Download full curriculum (PDF)01 Introduction To Data Engineering 6 lessons ▶
- 1.1Introduction to Data engineering
- 1.2What is SQL
- 1.3Overview of Application Architecture and RDBMS
- 1.4Overview of Database Technologies
- 1.5Purpose Built Database
- 1.6RDBMS vs Database Technologies
02 Setting Up the System 3 lessons ▶
- 2.1Setting up VS code and Python
- 2.2Setting up MYSQL Server and PostgreSQL
- 2.3Setting up Jupyter Lab online
03 Introduction to SQL Queries 3 lessons ▶
- 3.1Creating the table in SQL
- 3.2Relationship Model In SQL
- 3.3Starting with SQL
04 Manipulation Commands in SQL 7 lessons ▶
- 4.1Insert Command
- 4.2Update part1
- 4.3Update part2
- 4.4Delete command
- 4.5Drop and Truncate
- 4.6Alter Part1
- 4.7Alter Part2
05 Retrieving Data in SQL 4 lessons ▶
- 5.1Retrieving data using Select and where
- 5.2Sorting and using Alias
- 5.3Limit and Select with Expression
- 5.4Logical Operators
06 Aggregation in SQL 2 lessons ▶
- 6.1Aggregation function part1
- 6.2Aggregation Function part2
07 Joins in SQL 7 lessons ▶
- 7.1Types of joins
- 7.2Inner join
- 7.3Left Join
- 7.4Right Join
- 7.5Full Outer Join
- 7.6Self join
- 7.7Cross join
08 Set Operators in SQL 2 lessons ▶
- 8.1Set Operator Part1mp4
- 8.2Set Operator Part2mp4
09 Subqueries and Case Statements 7 lessons ▶
- 9.1Subqueries Part4
- 9.2Subqueries Part1
- 9.3Subqueries Part2
- 9.4Subqueries Part3
- 9.5Case Statements Part1
- 9.6Case Statements Part2
- 9.7Case Statements Part3
10 CTEs and Views 4 lessons ▶
- 10.1CTEs Part1mp4
- 10.2CTEs Part2
- 10.3Views Part1
- 10.4View Part2
11 Advanced Topics 8 lessons ▶
- 11.1Window Function
- 11.2RowNumber(), Rank(), DenseRank()
- 11.3NTILE(), Lead(), Lag()
- 11.4Window Functions examples
- 11.5Stored Procedure
- 11.6Index
- 11.7Mini Project Part1
- 11.8Mini Project Part2
12 Introduction to Python Programming 5 lessons ▶
- 12.1Introduction to Python
- 12.2Variables and Data Types
- 12.3Input and Output
- 12.4Basic Operator
- 12.5Constants and Type casting
13 Data Structures in Python 8 lessons ▶
- 13.1Conditional Statements
- 13.2Loop Part1
- 13.3Loop Part2
- 13.4Loop Part3
- 13.5List
- 13.6Tuple and Dictionary
- 13.7Sets
- 13.8Mini project
14 Functions, Modules & File Handling 11 lessons ▶
- 14.1Functions Part1
- 14.2Functions Part2
- 14.3Functions Part3
- 14.4Project Calculator Part1
- 14.5Project Calculator Part2
- 14.6Modules Part1
- 14.7Modules part2
- 14.8Module Part3
- 14.9File Handling Part1
- 14.10File Handling Part2
- 14.11Error Handling
15 Python Collections for Data Engineering 10 lessons ▶
- 15.1File I_O using Python
- 15.2Read Data from CSV into python
- 15.3Python Collectionsmp4
- 15.4Processing Python List
- 15.5Lambda Functions
- 15.6Filter data in Python
- 15.7Manipulating List using set()mp4
- 15.8Sort() & Overview of json strings
- 15.9Reading data from json files
- 15.10Extracting column from json file
16 Data Processing using Pandas 10 lessons ▶
- 16.1Overview of pandas for Data Processing
- 16.2Pandas Dataframe functionalities
- 16.3Aggregation using pandas
- 16.4Joining data using pandas
- 16.5Summarizing Joins
- 16.6Aggregation on Join Results Part1
- 16.7Aggregation on Join Results Part2
- 16.8Aggregation on Join Results Part3
- 16.9Sort data using pandas dataframe
- 16.10Exporting data to JSON files
17 Project1: File Format Converter 12 lessons ▶
- 17.1Overview of the files to be processed
- 17.2Working with regular expressions
- 17.3Getting Column names from the data
- 17.4Generating File paths and writing dataframe
- 17.5Recap of the Project
- 17.6Writing pandas dataframe to JSON file
- 17.7Converting csv to json
- 17.8Modularize File Format Converter
- 17.9Wrapping up the Project
- 17.10Runtime Arguments part1
- 17.11Runtime Arguments part2
- 17.12Environment variables
18 Project2: File to Database Loader 5 lessons ▶
- 18.1Configuring sql to python
- 18.2Connecting to the database
- 18.3Validate pandas and SQL integration
- 18.4Write CSV data to tables in chunks
- 18.5Project ready for deployment
19 Troubleshoot and Debugging Python Issues 6 lessons ▶
- 19.1Troubleshooting and Debugging
- 19.2Troubleshooting Database connectivity
- 19.3Troubleshooting Credentials for Database connectivity
- 19.4Troubleshooting Errors in Python
- 19.5Overview of SDLC
- 19.6Unit Testing & Debugging
20 Overview of Performance Tunning in Python 9 lessons ▶
- 20.1Introduction to Performance of Python Applications
- 20.2Setting up Database Loader
- 20.3Cleaning the data and Performance Tunning
- 20.4Performance Tunning
- 20.5Loading Multi Files in Python
- 20.6Develop applications for Multiprocessing
- 20.7Multiprocessing using python
- 20.8Introduction to Parallel Processing
- 20.9Executing Parallel Processing
21 Getting started with GCP 5 lessons ▶
- 21.1Introduction to Getting started with GCP
- 21.2Signing Up for GCP Account
- 21.3Overview of Google cloud platform
- 21.4Overview of google cloud shell
- 21.5Overview of Analytics services on GCP
22 Overview of Big Data and Data Lakes 9 lessons ▶
- 22.1Database and Their Types
- 22.2Technologies for Different Databases
- 22.3Usecase of different databases
- 22.4Volume of different database
- 22.5Overview of Big data and its evaluation
- 22.6Data Lake using Hadoop
- 22.7Overview of Modern Data Lakes on Cloud
- 22.8Implementation of Modern Data Lakes
- 22.9Implementation + Advantages of Modern Data Lakes
23 Overview of Spark and Spark Architecture 8 lessons ▶
- 23.1Overview of Data Processing
- 23.2Setting up the Environment
- 23.3Code Examples of different Libraries
- 23.4Difference between libraries and Distributed Computing
- 23.5Example of Distributed Computing
- 23.6Overview of Apache Spark Documentation
- 23.7Spark Key Features and Infrastructure
- 23.8Overview of Spark clusters and Key terms
24 Setup Databricks Environment using GCP 5 lessons ▶
- 24.1Overview of Databricks on GCP
- 24.2Signing up for Databricks on GCP
- 24.3Resolving Workspace errors and creating cluster
- 24.4Setup Databricks
- 24.5Databricks Features and Architecture
25 Basic Transformation using Spark SQL 7 lessons ▶
- 25.1Processing Data in databricks using spark
- 25.2Getting started with Spark SQL
- 25.3Create views and retrieve them using Spark SQL
- 25.4Spark SQL Query to compute Daily Product Revenue
- 25.5Overview of pyspark and Reading json files
- 25.6Get Schema details using pyspark
- 25.7Writing a function using Pyspark Python
26 Create Delta tables using Spark SQL 4 lessons ▶
- 26.1Understanding delta tables and Database
- 26.2Creating Delta tables
- 26.3Copy data from external file
- 26.4Performing operations in delta tables
27 Pre-defined Functions in Spark SQL 29 lessons ▶
- 27.1Overview of Functions
- 27.2Describe Functions in Spark SQL
- 27.3Case Conversion Functions
- 27.4Extracting Data using Functions
- 27.5Trimming and Padding Functions
- 27.6Reverse and Concatenation functions
- 27.7Date Manipulation Functions
- 27.8Beginning Date or Time Functions
- 27.9Extracting information using date_format Part1
- 27.10Extracting information using date_format Part2
- 27.11Calendar Functions
- 27.12Dealing with Unix Timestamp
- 27.13Mathematical Functions Part1
- 27.14Mathematical Functions Part2
- 27.15Miscellaneous Functions Part1
- 27.16Miscellaneous Functions Part2
- 27.17Miscellaneous Functions Part3
- 27.18CASE Statements Part1
- 27.19CASE Statements Part2
- 27.20Word count Query part1
- 27.21Word count query Part2
- 27.22Practice Exercise1
- 27.23Practice Exercise2
- 27.24Practice Exercise3
- 27.25Practice Exercise4
- 27.26Practice Exercise5
- 27.27Practice Exercise6
- 27.28Practice Exercise7
- 27.29Practice Exercise8-
28 Setup Spark Tables for Basic Transformations 1 lesson ▶
- 28.1Introduction to transformation using Spark SQL
29 Filtering Data using Spark SQL 2 lessons ▶
- 29.1Filtering using Spark SQL part1
- 29.2Filtering using Spark SQL Part2
30 Aggregations using Spark SQL 2 lessons ▶
- 30.1Aggregation using Spark SQL part1
- 30.2Aggregation using Spark SQL part2
31 Joins using Spark SQL 2 lessons ▶
- 31.1Joins using Spark SQL part1
- 31.2Joins using Spark SQL part2
32 Sorting using Spark SQL 1 lesson ▶
- 32.1Sorting using Spark SQL
33 Copy Query Results into Spark Tables 3 lessons ▶
- 33.1Copy Query Results using CTAS and INSERT
- 33.2Design Pipleline using CTAS and INSERT
- 33.3Copy Query Results using MERGE
34 Ranking using Spark SQL Windowing Functions 1 lesson ▶
- 34.1Ranking using Spark SQL
35 Processing JSON like data using Spark SQL 4 lessons ▶
- 35.1Creating Tables with Array type Columns
- 35.2Creating and Dealing with Array of Struct Type columns
- 35.3Overview of Functions to Process Data in Spark SQL
- 35.4Processing Delimited strings using Spark SQL
36 Getting Started with Pyspark DataFrame APIs 4 lessons ▶
- 36.1Introduction and Setup pf Pyspark
- 36.2Process schema details in json using pyspark
- 36.3Transforming and Processing Schemas
- 36.4Convert CSV to Parquet with schema
37 Create spark data frames using Pyspark 4 lessons ▶
- 37.1Creating Spark DataFrame and Processing json-like data
- 37.2Overview of Data Processing and DataFrame concepts
- 37.3Advantages and Transformations of DataFrame
- 37.4using withColumn() and Writing Dataframe to Delta Files
38 Basic Transformation using Pyspark 6 lessons ▶
- 38.1Overview of Transformations and withColumn usecase
- 38.2Filtering Data using PySpark DataFrame
- 38.3Aggregations using Spark DataFrame
- 38.4Aggregation using Spark DataFrame APIs
- 38.5Sorting data using PySpark DataFrame
- 38.6Dealing with Nulls while sorting the data
39 Joining Data using Spark DataFrame 2 lessons ▶
- 39.1Inner Join over the Spark dataframe APIs
- 39.2Left and Right Join using Spark dataframe APIs
40 Ranking using Pyspark 4 lessons ▶
- 40.1Overview of ranking along with OrderBy functionality
- 40.2Filter Based on rank using Spark dataframe APIs
- 40.3Filter based on Rank per Partition using spark Dataframe
- 40.4Difference Between Rank() and Dense_Rank()
41 Integration of Spark SQL and Pyspark data frame 3 lessons ▶
- 41.1Introduction to Integration of Spark SQL and PySpark Dataframe APIs
- 41.2Run Spark SQL queries and Create Spark Metastore Tables using DataFrames
- 41.3Dataframe Transformations over Metastore Tables
42 ELT Data Pipelines using Databricks 3 lessons ▶
- 42.1Overview of Databricks Workflows
- 42.2Multi-Task Workflows using Databricks jobs
- 42.3Working on Multi-Task job
43 Performance Tuning of Spark-Catalyst Optimizer 4 lessons ▶
- 43.1Getting Started with Performance Tuning and Spark Catalyst Optimizer
- 43.2Overview of Plans using Spark UI
- 43.3Understanding of Spark Architecture with Example
- 43.4Filter, Broadcast Joins & Aggregations
44 Performance Tuning of Spark Cluster Configuration 3 lessons ▶
- 44.1Introduction to Databricks Cluster Configuration and its Types
- 44.2Setting up All-Purpose Cluster and Overview of Auto-Scaling
- 44.3Cluster Performance Tuning using Auto Scaling
45 Performance Tuning while inferring schema from csv or json files 3 lessons ▶
- 45.1Overview of Inferring Schema from CSV or JSON files
- 45.2Overview of CSV_JSON files
- 45.3Overhead of Inferring Schema
46 Performance Tuning using Columnar File Format and Partitioning Strategy 2 lessons ▶
- 46.1Introduction of Performance tuning while storing data and Side effects of CSV files in Data Lake
- 46.2Restructure CSV Data and performance Boosting
47 Introduction to Linux Commands for Data Engineering 9 lessons ▶
- 47.1Introduction to linux and Basic Navigation Commands
- 47.2File and Folder Management Commands
- 47.3Commands for Viewing and Editing Files
- 47.4Commands to make modification inside a Vi Editor
- 47.5Understanding of SED Command
- 47.6Understanding of GREP Command
- 47.7Understanding of AWK Command
- 47.8Understanding of Find Command
- 47.9Disk, Process & System Monitoring
Tools you'll use
Every course includes
Related courses
Advanced Excel with AI
Master Excel formulas, data management and AI-powered automation
Business Analytics
Combine Excel, SQL, Power BI and Python to turn business data into decisions
Data Analytics
Analyse data end-to-end with Excel formulas, pivot tables, SQL and Power BI