🔥 New Batch — 16 September 2026! AWS 8:00–9:30 AM & Azure 6:00–8:00 AM IST — Limited Seats! Reserve Your Seat →
Batches & Courses
🗓️ Upcoming Batches ⬇️ AWS Syllabus PDF ⬇️ Azure Syllabus PDF 🔥 Databricks Training ☁️ AWS Data Engineering 🔷 Azure Data Engineering 🌐 GCP Data Engineering 🔄 Apache Airflow 🤖 Generative AI ❄️ Snowflake + dbt 📊 Big Data Training
✅ Student Reviews 📬 Contact Us
📞 +91-8500002025 📞 +91-9247159150 💬 WhatsApp Trainer Venu 🚀 Reserve My Seat
New batch — Wednesday, 16 September 2026

Azure Data Engineering Syllabus
10 Weeks · 50 Topics · 5 Projects

The complete week-by-week curriculum for the Databricks + Azure Data Engineering program — built around PySpark + Azure + ADF + Databricks + Streaming + Fabric + AI/DevOps. Every weekday topic below is a live, hands-on class with Trainer Venu.

🕗 6:00 AM – 8:00 AM IST
📅 Monday – Friday
🎥 Recordings after every class
🎓 Weekend masterclasses
🛠️ 5 major projects
1200+Professionals trained
4.9 ★Google rating · 200+ reviews
50+Hiring companies
14+ yrsTrainer experience
10 weeks50 live topics + 5 projects
📋 Course Overview

What this program covers

Build strong PySpark foundations first, then connect them to cloud-native data engineering, lakehouse design, streaming, orchestration, governance, testing, performance, CI/CD and AI-assisted engineering productivity.

Course focusPySpark + Azure + ADF + Databricks + Streaming + Fabric + AI/DevOps
Structure10 weeks, 50 focused weekday topics, weekend masterclasses and 5 major projects
Learning styleConcepts + live coding + architecture + troubleshooting + project implementation
Ideal forData engineers, ETL developers, Spark developers and professionals moving into Azure/Databricks data engineering
Delivery modelInstructor-led sessions + guided labs + projects + weekend masterclasses
Batch timing6:00 AM – 8:00 AM IST, Monday – Friday
🎯 Learning Outcomes

What you will be able to do

Build ingestion and orchestration pipelines using ADF, ADLS and Databricks.

Implement Bronze/Silver/Gold architectures with Delta Lake and Unity Catalog.

Build streaming pipelines using Kafka/Event Hubs and Structured Streaming.

Design metadata-driven ADF frameworks and CI/CD with Databricks Asset Bundles.

Apply governance, security, testing, monitoring and Spark performance tuning.

Understand where Microsoft Fabric complements or replaces parts of an Azure data platform.

📚 Detailed Weekly Curriculum

Week-by-week breakdown — all 50 topics

Each numbered session is a focused class with demonstrations and coding exercises. Weekend sessions are used for AI-assisted engineering, metadata-driven design, architecture and project integration. Click any week to expand.

WEEK01PySpark & Databricks Foundations10 Hours · 5 topics + weekend masterclass
01PySpark & Databricks Introduction
Spark architecture: driver, executors and cluster managerJobs, stages, tasks and lazy evaluationDatabricks workspace, notebooks and computeDataFrame-based engineering workflow
02Databricks Date & Common Functions
Date/time parsingString and numeric functionsConditional logicArrays, maps and structs
03Databricks Window Functions
row_number, rank and dense_ranklag and leadPartitions and orderingRunning totals and top-N use cases
04Data Cleaning with Regular Expressions
Null handlingDuplicate handlingRegex extraction/replacementData standardization and validation
05spark-submit Internals
Submission lifecycleDriver and executor configurationCores/memory/parallelismReading logs and Spark UI evidence
🎓 Weekend Masterclass / Lab — Claude Code + Prompt Engineering Special Class
  • Generate and review PySpark
  • Debug Spark errors
  • Prompt patterns for ETL/SQL
  • AI-assisted coding with validation
WEEK02Advanced Spark Data Processing10 Hours · 5 topics + weekend masterclass
06Spark JSON Processing
Nested JSONStructTypeArrays/structsexplode and schema handling
07Spark JDBC Processing
OracleMySQLSQL ServerParallel JDBC reads and pushdown
08XML, CSV & Pandas Data Processing
XML patternsCSV optionsPandas vs PySparkPandas API on Spark overview
09Parquet & ORC
Columnar storageCompressionPredicate pushdownColumn pruning and partition design
10Delta Lake vs Apache Iceberg
ACID and table metadataTime travel and schema evolutionOpen-table format trade-offsWhen to choose Delta or Iceberg
🎓 Weekend Masterclass / Lab — Spark File-Format & Performance Lab
  • Convert raw files to Parquet/Delta
  • Inspect execution plans
  • Compare scan behavior
  • Practice partition-pruning scenarios
WEEK03Azure Storage & Azure Data Factory10 Hours · 5 topics + weekend masterclass
11Azure Introduction, Blob Storage & ADLS Gen2
Subscriptions/resource groupsStorage accounts and containersBlob vs ADLS Gen2Hierarchical namespace, access tiers and lake folder design
12Azure Data Factory Introduction & Internals
ADF architecturePipelines, activities and datasetsLinked servicesIntegration Runtime basics
13Core ADF Activities
Copy ActivityForEachLookupGet MetadataIf Condition and dependency control
14Databricks Notebook + Lookup + ForEach
ADF to Databricks orchestrationPassing parametersDynamic notebook executionLoop-driven processing
15Parameterization & Mapping Data Flows
Pipeline parametersDataset parametersDynamic content/expression languageData Flow transformations and performance considerations
🎓 Weekend Masterclass / Lab — ADF with Claude Code — JSON-Driven Development
  • Generate pipeline JSON safely
  • Understand ADF ARM/resource structure
  • Parameterize reusable pipelines
  • Validate AI-generated JSON before deployment
WEEK04ADF Advanced, Security & End-to-End Project10 Hours · 5 topics + weekend masterclass
16Schedulers, Key Vault & Triggers
Schedule/tumbling/event triggersAzure Key Vault integrationSecret referencesOperational scheduling patterns
17SCD Type 1, SCD Type 2 & Integration Runtime
SCD conceptsMapping/SQL approachesSelf-hosted IRAzure IR and network/connectivity considerations
18Databricks + ADLS Integration
ABFS accessManaged identity/service principal conceptsUnity Catalog storage credentialsSecure data access
19Data Governance & Security
Azure RBACEncryptionNetwork securityManaged identities, private endpoints and least privilege
20ADF End-to-End Pipeline Project using Databricks
Source ingestionADF orchestrationDatabricks transformationADLS Bronze/Silver/GoldMonitoring and rerun strategy
🎓 Weekend Masterclass / Lab — ADF Metadata-Driven Pipelines
  • Configuration tables/files
  • Dynamic source/target processing
  • Reusable Copy/Notebook activities
  • Centralized error handling and logging
WEEK05Delta Lake, Medallion & Databricks Optimization10 Hours · 5 topics + weekend masterclass
21Databricks Utilities
dbutils.fsdbutils.secretsdbutils.widgetsParameter-driven notebook design
22Delta Time Travel & Change Data Feed
Version historyRestore/read-as-ofCDF for incremental downstream processingMERGE and audit use cases
23Medallion Architecture on Azure
Bronze/Silver/Gold responsibilitiesADLS landing zonesData quality/quarantineIdempotent reprocessing
24Spark Declarative Pipelines (SDP)
Declarative pipeline conceptsDependency graphsExpectations/data qualityIncremental processing and operational benefits
25Databricks Optimization
OPTIMIZE/VACUUMPartitioning and file sizingZ-Ordering and Liquid ClusteringAQE, Photon, joins and skew
🎓 Weekend Masterclass / Lab — Lakehouse Federation + Delta Sharing
  • Federated-query use cases
  • Secure data sharing
  • Internal/external consumers
  • Governance considerations
WEEK06Databricks Engineering, Deployment & Governance10 Hours · 5 topics + weekend masterclass
26Databricks Asset Bundles (DABs)
Bundle structuredatabricks.ymlEnvironment targetsJobs/pipelines deployment
27Databricks Lakeflow
Lakeflow Connect overviewJobs and pipeline ecosystemOperational workflow patternsIngestion-to-orchestration integration
28Lakeflow Orchestration Advanced
Task dependenciesParametersRetries and notificationsMulti-task workflows and environment promotion
29Unity Catalog Introduction & Auto CDC
Metastore/catalog/schema/table hierarchyManaged vs external objectsLineageAuto CDC concepts for incremental changes
30Unity Catalog Advanced
Storage credentials/external locationsPrivilegesRow filters and column masksGoverned sharing and auditing
🎓 Weekend Masterclass / Lab — CI/CD with DABs
  • Git-based development
  • Dev/Test/Prod targets
  • GitHub Actions pipeline concept
  • Deployment validation and rollback discussion
WEEK07Streaming on Azure10 Hours · 5 topics + weekend masterclass
31Spark Structured Streaming Introduction
Streaming DataFramesMicro-batchesOutput modesState and checkpoints
32Kafka with Spark
Producer/consumer basicsTopics/partitions/offsetsSpark Kafka connectorReplay and failure handling
33NiFi + Kafka + Spark Streaming
NiFi ingestion patternsKafka transport layerSpark consumptionEnd-to-end observability points
34Azure Event Hubs Streaming Project
Event Hubs architecturePartitions and consumer groupsKafka-compatible endpoint conceptsDatabricks Structured Streaming integration
35Structured Streaming, Kafka & Auto Loader Internals
CheckpointsWatermarksExactly-once-oriented designAuto Loader schema inference/evolution
🎓 Weekend Masterclass / Lab — Real-Time Azure Streaming Project
  • Producer/Event Hub → Databricks
  • Bronze streaming table
  • Silver transformations
  • Gold aggregates + checkpoint/replay demonstration
WEEK08Airflow, Testing, Observability & Spark Performance10 Hours · 5 topics + weekend masterclass
36Apache Airflow Introduction
DAGs, tasks and operatorsScheduler/executorVariables/connectionsXCom basics
37Airflow Advanced Orchestration
TaskFlow APISensors and trigger rulesDynamic task mappingDatabricks/ADF integration patterns
38Testing & Observability
PySpark unit testsData-quality testsPipeline logging and metricsDatabricks/Azure Monitor concepts
39Performance Tuning Deep Dive
Explain plansPartitioningJoin strategiesMemory, shuffle and skew optimization
40Spark RDD Internals
RDD lineageNarrow vs wide transformationsDAG constructionFault tolerance and recomputation
🎓 Weekend Masterclass / Lab — Production Troubleshooting Workshop
  • Diagnose a slow Spark job
  • Trace pipeline failure across ADF/Databricks
  • Analyze logs and Spark UI
  • Prioritize performance fixes
WEEK09Microsoft Fabric & OneLake10 Hours · 5 topics + weekend masterclass
41Microsoft Fabric & OneLake Core Architecture
Fabric workloadsOneLake conceptsWorkspaces and capacitiesHow Fabric fits with Azure data platforms
42Fabric Lakehouse vs Synapse Data Warehouse
Lakehouse vs warehouseDelta tablesSQL analytics endpointWhen to choose each architecture
43Real-Time Intelligence & Data Activator
Real-time ingestion/analytics conceptsEventstreams/KQL conceptsTriggering actions from data changesOperational scenarios
44Semantic Models in Microsoft Fabric
Model relationshipsMeasures and star schema thinkingDirect Lake conceptsPerformance and governance considerations
45Prepare AI-Ready Analytics Data in Fabric
Curated data productsSemantic consistencySecurity/governancePreparing trusted data for AI/analytics consumers
🎓 Weekend Masterclass / Lab — Azure vs Fabric Architecture Workshop
  • ADF/Databricks/Fabric role comparison
  • Lakehouse vs warehouse selection
  • Migration and coexistence patterns
  • Cost and operating-model discussion
WEEK10AI-Assisted Engineering, CI/CD & Interview Preparation10 Hours · 5 topics + weekend masterclass
46Claude Code for Data Engineers
Repository-aware codingGenerate/debug PySparkGenerate ADF/Lakeflow examplesTesting AI output
47Prompt Engineering
Context engineeringStructured promptsDebugging promptsArchitecture prompts
48Codex & GitHub Copilot for Productivity
Code generationUnit testsDocumentationRefactoring and review
49GitHub Actions CI/CD
Git workflowBuild/testDatabricks Asset Bundle deploymentEnvironment/secrets management
50Interview Tips & Resume Preparation
PySpark and SQL scenariosADF/Databricks architectureStreaming/performance troubleshootingResume and project explanation
🎓 Weekend Masterclass / Lab — Mock Azure Data Engineer Architecture & Interview Clinic
  • Design an end-to-end Azure platform
  • Explain security and governance choices
  • Practice troubleshooting questions
  • Convert course projects into resume-ready experience statements
🛠️ Hands-On Projects

5 major projects you will build

Projects are deliberately aligned with the course sequence — you learn a concept, then implement it as part of a realistic, resume-ready pipeline.

Project 1 — ADF + Databricks Batch Pipeline
Source DB / Files → ADF → ADLS Bronze → Databricks → Silver/Gold Delta → Analytics
  • Linked services/datasets
  • Parameterization
  • Databricks orchestration
  • Medallion architecture
Project 2 — Metadata-Driven ADF Framework
Metadata Config → ADF Lookup/ForEach → Dynamic Copy/Notebook → Logging → Target
  • Reusable pipeline design
  • Dynamic expressions
  • Centralized configuration
  • Operational error handling
Project 3 — Azure Databricks Lakehouse
ADLS Landing → Auto Loader → Bronze → Silver → Gold → Unity Catalog → BI/Consumers
  • Incremental ingestion
  • Delta Lake
  • Data quality
  • Governance and optimization
Project 4 — Azure Real-Time Streaming
Producer / Kafka → Azure Event Hubs → Databricks Structured Streaming → Delta → Gold
  • Event Hubs
  • Kafka compatibility
  • Checkpointing
  • Watermarks and late-data handling
Project 5 — CI/CD & Production Deployment
Git → GitHub Actions → Databricks Asset Bundles → Dev/Test/Prod → Validation
  • Version control
  • Automated tests
  • Environment promotion
  • Deployment governance
🧰 Technology Stack

Tools and services covered

The curriculum focuses on data-engineering use cases. General cloud administration topics are covered only when they directly affect pipeline design, security, performance or operations.

AreaTechnologies / concepts
Core EngineeringPySpark, Spark SQL, JDBC, JSON, XML, CSV, Parquet, ORC, Delta Lake, Iceberg concepts
Azure StorageAzure Blob Storage, ADLS Gen2
Azure IntegrationAzure Data Factory, Integration Runtime, Triggers, Mapping Data Flows, Key Vault
Azure StreamingAzure Event Hubs, Apache Kafka, Structured Streaming, Auto Loader, NiFi
DatabricksDatabricks, Delta Lake, Unity Catalog, Lakeflow, Spark Declarative Pipelines, DABs
Governance & SecurityAzure RBAC, Managed Identity, Service Principal concepts, Encryption, Private networking, Unity Catalog
Orchestration & OpsADF, Airflow, Databricks Workflows/Lakeflow, Testing, Monitoring, Observability
Microsoft FabricOneLake, Fabric Lakehouse, Warehouse, Real-Time Intelligence, Semantic Models
DevOps & AIGit, GitHub, GitHub Actions, Claude Code, Codex, GitHub Copilot, Prompt Engineering
👨‍🏫 Your Trainer

Learn directly from Trainer Venu

Trainer Venu Katragadda — Databricks and data engineering trainer

Venu Katragadda

Founder, Sreyobhilashi IT · 14+ years in Big Data & Cloud Data Engineering

Venu has trained 1200+ working professionals on Spark, Databricks, AWS and Azure, and still teaches every session himself — no junior trainers, no recorded-only classes. Sessions are built around what actually breaks in production: skewed joins, failing streams, schema drift, cost blow-ups and the interview questions that follow.

PySpark & Spark internalsDatabricks & Delta LakeAWS Glue · EMR · Kinesis Azure ADF · FabricKafka & Structured StreamingAirflowCI/CD with DABsClaude Code for DE
⭐ Google Reviews

Reviews from our Azure & Databricks students

Real, verified reviews from data engineers who trained with Venu — on Databricks, AWS, Azure, PySpark and streaming.

4.9
★★★★★
Based on 200+ Google reviews
★★★★★

“Recently took Databricks classes with Venu to upskill in trending technologies, and the experience exceeded all expectations. While I initially sought guidance only on Databricks, Venu provided in-depth training across the entire ecosystem — AWS, Kafka, NiFi, Airflow and PySpark.”

AZ
Abhishek Zararia
Databricks · Cleared DE Professional Cert
✅ Verified Google Review
★★★★★

“I had a truly valuable experience with Venu's Spark training along with AWS & Azure Databricks Training. He is highly knowledgeable, and the sessions are very well structured with extensive hands-on coverage.”

DM
Dhevipriya Marimutbhu
AWS & Azure Databricks Training
✅ Verified Google Review
★★★★★

“This training has exceeded my expectations. Venu explains concepts clearly and uses hands-on examples that make the content easy to understand. I am learning a lot and would definitely recommend it.”

NL
Pataballa N V Lakshminarayana
Databricks Training
✅ Verified Google Review
★★★★★

“I recently completed the Data Engineering course on Databricks and AWS. Venu Sir delivers instruction at the next level, focusing on high-performance learning. He explains every concept clearly and thoroughly, with practical examples.”

MM
Mahaboob Mulla
Databricks & AWS Training
✅ Verified Google Review
★★★★★

“Venu sir has explained end to end streaming project, data cleansing, and parsing various source data. This has helped me in my project work. The explanation on Spark architecture and other key concepts helped me understand Spark deeply.”

NS
Nimisha Shah
Databricks Streaming
✅ Verified Google Review
★★★★★

“I had a truly valuable experience with Sreyobhilashi's AWS & Azure Databricks Training & Placement Program. Hands-on coverage of Spark, Kafka, Flink, NiFi, Airflow, Azure and Snowflake.”

NT
Naveen Kumar Tavva
AWS & Azure Databricks Training
✅ Verified Google Review

Next Azure batch starts Wednesday, 16 September 2026

Live online, 6:00 AM – 8:00 AM IST, Monday to Friday. Limited seats so every student gets doubt-clearing time and cloud lab support.

Chat with Venu
📞 Call 💬 WhatsApp 🚀 Reserve Seat