End-to-End
Data Engineering

Master modern data pipelines, warehousing, and orchestration with a heavy focus on Databricks Certification Prep.

Level: Beginner to Intermediate
Starts: October 5
Duration: 40 Days
View Pricing & Schedule
Python SQL Databricks PySpark Apache Airflow dbt Snowflake/BigQuery Docker

Course Syllabus

A. Core Python Programming

  • Data types, Variables, I/O
  • Functions, Scoping, Lambda, Map/Filter
  • Control Structures, Loops
  • Exception Handling & Logging

B. Data Structures

  • Lists, Tuples, Dictionaries, Sets
  • Comprehensions and Generators

C. Working with Files & APIs

  • Reading and writing text, CSV, JSON, Parquet
  • Consuming REST APIs using requests
  • Working with environment variables & config files

D. Essential Libraries

  • NumPy: arrays, slicing, broadcasting, vectorized operations
  • Pandas: Series, DataFrames, indexing, grouping, pivoting, reshaping

A. Basics

  • SELECT, WHERE, ORDER BY, DISTINCT, LIMIT, ALIAS

B. Intermediate SQL

  • JOINs (INNER, LEFT, RIGHT, FULL)
  • GROUP BY, HAVING
  • Subqueries, CTEs

C. Advanced SQL

  • Window functions (ROW_NUMBER, RANK, LAG, LEAD)
  • Aggregations, Nested queries, Date/Time functions
  • Query optimization basics (indexes, execution plans)

D. Practice

  • Real-world datasets (Sales, HR, E-commerce)

A. Data Modeling Fundamentals

  • OLTP vs OLAP, Normalization vs Denormalization
  • Star Schema, Snowflake Schema
  • Fact tables vs Dimension tables
  • Slowly Changing Dimensions (SCD Types 1/2)

B. Cloud Data Warehouses

  • Snowflake basics: architecture, virtual warehouses, loading data
  • BigQuery basics: datasets, partitioning, clustering
  • Redshift basics (overview & comparison)

C. File Formats & Storage

  • CSV, JSON, Parquet, Avro, ORC
  • Delta Lake format (bridges into Databricks module)
  • Partitioning strategies

A. ETL vs ELT Concepts

  • Batch vs streaming ingestion
  • Idempotency, incremental loads, backfills

B. Data Cleaning & Transformation

  • Handling missing values, duplicates, outliers at scale
  • Schema validation and enforcement
  • Data quality checks (Great Expectations basics)

C. dbt (Data Build Tool)

  • Models, sources, seeds
  • Materializations (view, table, incremental)
  • Tests and documentation
  • dbt + warehouse workflow (Snowflake/BigQuery)

A. Airflow Fundamentals

  • Architecture: Scheduler, Webserver, Executor, Metadata DB
  • DAGs, Tasks, Operators
  • Installing and running Airflow locally / with Docker

B. Building Pipelines

  • PythonOperator, BashOperator, and provider operators
  • Task dependencies, branching, trigger rules
  • Sensors, hooks, and XComs for passing data

C. Scheduling & Orchestration

  • Scheduling intervals, backfills, catchup
  • Retries, SLAs, alerting (email/Slack)
  • Dynamic DAG generation

D. Airflow in Practice

  • Connecting Airflow to databases, APIs, and cloud storage
  • Orchestrating a full ETL pipeline end-to-end
  • Airflow with Databricks and Spark jobs
  • Monitoring, logging, and debugging DAGs
  • Deploying Airflow (Docker Compose, Astronomer, managed services)

A. PySpark Basics

  • Setting up SparkSession, RDDs vs DataFrames

B. DataFrame API

  • Read/Write CSV, JSON, Parquet
  • Filter, Select, GroupBy, Join
  • Partitioning and performance tuning basics

C. PySpark SQL & Streaming

  • Creating temporary views and writing SQL queries on Spark DataFrames
  • Structured Streaming basics: Reading from Kafka, writing to sinks

Aligned to Databricks Certified Data Engineer Associate. Covers Lakeflow, Delta Lake, and Unity Catalog.

Exam Domain Weighting

DomainWeight
Databricks Intelligence Platform6%
Data Ingestion and Loading21%
Data Transformation and Modeling22%
Working with Lakeflow Jobs16%
Implementing CI/CD10%
Troubleshooting, Monitoring & Optimization10%
Governance and Security15%

A. Databricks Intelligence Platform

  • Lakehouse architecture (Delta/Parquet)
  • Workspace components (Notebooks, Repos, Clusters, Jobs, DBFS/Volumes)
  • Compute types (All-Purpose, Job Clusters, SQL Warehouses, Serverless)

B. Data Ingestion and Loading

  • Auto Loader (incremental ingestion, schema inference/evolution)
  • COPY INTO (SQL-based batch loading)
  • Lakeflow Connect & Structured Streaming fundamentals
  • Relational entities & Unity Catalog namespace

C. Data Transformation and Modeling

  • Delta Lake internals (ACID, transaction log, time travel)
  • MERGE INTO, SCDs, Schema enforcement vs evolution
  • Performance optimizations (OPTIMIZE, ZORDER BY, Liquid Clustering, VACUUM)
  • Medallion architecture (Bronze → Silver → Gold)
  • ELT with Spark SQL and Python

D. Working with Lakeflow Jobs

  • Lakeflow Jobs (DAG-based orchestration)
  • Lakeflow Declarative Pipelines (formerly DLT)
  • Data quality Expectations & AUTO CDC

E. Implementing CI/CD

  • Databricks Repos / Git Folders
  • Databricks Asset Bundles (DABs) & Databricks CLI
  • Typical flow: Git → PR tests → Staging → Prod

F. Troubleshooting, Monitoring & Optimization

  • Reading the Spark UI (stages, tasks, skew, spill)
  • Cluster/job event logs
  • Common fixes (compaction, salting, AQE, broadcast joins)

G. Governance and Security (Unity Catalog)

  • Three-level namespace (catalog.schema.table)
  • Centralized access control (GRANT/REVOKE)
  • Data lineage, Volumes, and Audit logging

H. Databricks + Airflow Integration

  • Triggering Lakeflow Jobs/pipelines from Airflow
  • End-to-end lakehouse pipeline orchestration
  • Messaging system concepts (Kafka basics: topics, producers, consumers)
  • Batch vs real-time architectures
  • Simple streaming pipeline: Kafka → Spark Structured Streaming → Delta table

A. Git & GitHub

  • Version control basics for data pipelines, Branching strategy

B. Docker

  • Containerizing pipelines and jobs, Docker Compose for local setups

C. CI/CD Basics

  • GitHub Actions for testing/deploying, Automated testing (pytest, dbt tests)

D. Cloud Fundamentals (Overview)

  • AWS (S3, IAM), GCP (Cloud Storage, BigQuery), Azure (Data Lake)
  • When to use managed services vs self-hosted tools

Beginner Level

  • Batch ETL Pipeline: Extract from API/CSV → clean with Pandas → load into SQL database
  • SQL-based sales/HR data analysis with automated reporting

Intermediate Level

  • End-to-End Airflow Pipeline: Ingest → transform → load into cloud warehouse (Snowflake/BigQuery)
  • dbt-based transformation layer on top of a warehouse, orchestrated by Airflow

Advanced Level

  • Lakehouse Pipeline: Airflow + Databricks + Delta Lake, orchestrating multi-stage transformations
  • Real-time Pipeline: Kafka + Spark Structured Streaming + Delta table, monitored via Airflow
  • Fully Containerized Pipeline: CI/CD (Docker + GitHub Actions) deploying to cloud warehouse

Course Scope & Tools

Tools & Platforms

IDEs: VS Code, Jupyter, Databricks Notebooks
Version Control: Git, GitHub
Orchestration: Apache Airflow
Processing: PySpark, Databricks
Warehousing: Snowflake / BigQuery
Transformation: dbt
Streaming: Kafka (intro)
Containerization/CI-CD: Docker, GitHub Actions
Data Platforms: Kaggle, UCI ML Repo, synthetic datasets

Notes on Scope

This track intentionally excludes deep ML modeling (regression/classification/clustering) and BI-tool-heavy dashboarding (Power BI/Tableau), since those belong to the Data Analyst / Data Scientist tracks. The focus here is pipelines, orchestration, storage, and infrastructure — the core of what a Data Engineer is hired to build and maintain.

Upcoming Live Batch Details

₹7,000 All-Inclusive
Start Date
October 5
Duration
40 Days
Timings
8:00 PM – 9:00 PM IST
Live Classes
Mon, Wed, Fri
Reserve Your Seat Now