Exceptional Training for Data Professionals

Apache training courses
London UK & Live Online

EDF logo Capita logo Sky logo NHS logo RBS logo BBC logo CISCO logo

Welcome to the Apache course group.

This group contains JBI Training's Apache courses, covering the Apache technologies used in modern data engineering, big data processing, distributed computing, messaging, and enterprise application development. Whether you're learning a specific Apache project or building expertise across the Apache ecosystem, our courses provide practical, hands-on training using real-world scenarios.

Courses in this group cover technologies including Apache Spark, Apache Kafka, Hadoop, Hive, HBase, Cassandra, Flink, Camel, Maven, Tomcat, ActiveMQ, and other Apache frameworks commonly used for large-scale data processing, integration, and enterprise software development.

Browse the courses in this group to find the training that best matches your experience level and learning goals.

JBI Training offers four Apache courses covering the most widely used Apache data processing and streaming technologies. Available courses are Apache Spark Development (two days), Apache Spark 3 — Databricks Certified Associate Developer (five days), Apache Kafka Essentials (two days), and Apache Storm (two days). All courses are available as scheduled classroom sessions in London, as live online instructor-led training, or as customised onsite programmes for data engineering and platform teams.
Apache Spark is an open-source, distributed data processing engine designed for large-scale data analytics and transformation workloads. It processes data in memory across a cluster of machines, making it significantly faster than older batch processing frameworks such as Hadoop MapReduce for most workloads. Spark is used for large-scale ETL and data transformation pipelines, machine learning at scale using the MLlib library, graph processing, stream processing using Spark Structured Streaming, and interactive data analysis. It is the most widely adopted distributed data processing framework in the industry and is available on all major cloud platforms including Azure Databricks, AWS EMR, and Google Dataproc.
The Apache Spark Development course is a two-day practical introduction to Spark for data engineers and developers who need to build and run Spark workloads. It covers the Spark architecture, the DataFrame API, Spark SQL, data transformation and aggregation, reading and writing data in various formats, and an introduction to Structured Streaming. The Apache Spark 3 — Databricks Certified Associate Developer course is a comprehensive five-day programme that covers Spark 3 in full depth and prepares delegates for the Databricks Certified Associate Developer for Apache Spark certification examination. It includes advanced Spark topics, Databricks-specific features, performance tuning, and certification-focused preparation. The five-day course is suited to data engineers who want a thorough grounding in Spark 3 and a recognised professional credential.
Apache Kafka is an open-source distributed event streaming platform designed to handle high-throughput, fault-tolerant, real-time data streams. It acts as a highly scalable message broker that allows applications to publish, subscribe to, store, and process streams of events in real time. Kafka is widely used for building real-time data pipelines, event-driven microservices architectures, activity tracking, operational monitoring, and stream processing applications. JBI's two-day Apache Kafka Essentials course covers the Kafka architecture and core concepts — including topics, partitions, producers, consumers, and consumer groups — setting up and configuring Kafka, producing and consuming messages, Kafka Connect for integrating with external systems, Kafka Streams for stream processing, and operational and monitoring considerations for running Kafka in production.
Apache Storm is an open-source distributed real-time computation system designed for processing unbounded streams of data with very low latency. It processes individual events as they arrive, making it well-suited to use cases that require immediate, sub-second processing of each event — such as fraud detection, real-time alerting, and financial transaction processing. Spark Structured Streaming processes data in micro-batches, introducing a small amount of latency in exchange for higher throughput and easier integration with the rest of the Spark ecosystem. The choice between Storm and Spark Streaming depends on latency requirements, existing tooling, and the nature of the streaming workload. JBI's two-day Apache Storm course covers Storm's topology model, spouts and bolts, fault tolerance, state management, and practical stream processing use cases.
The Databricks Certified Associate Developer for Apache Spark is a professional certification that validates a developer's ability to use the Spark DataFrame API, Spark SQL, and Spark's core processing capabilities at an associate level. It is widely recognised in the data engineering community and is particularly relevant for professionals working in Azure Databricks, AWS, or Google Cloud environments. JBI's five-day Apache Spark 3 — Databricks Certified Associate Developer course is specifically designed to prepare delegates for this examination, covering the full scope of the certification syllabus with hands-on exercises, practice questions, and exam technique guidance alongside comprehensive technical content.
Yes. All Apache courses at JBI can be delivered as customised onsite or online programmes for corporate data engineering and platform teams. Content and exercises can be tailored to the team's existing data stack, cloud environment, and specific use cases — for example, a team using Azure Databricks can receive Spark training focused on the Databricks environment, or a team building an event-driven microservices architecture can receive Kafka training focused on their specific integration patterns. JBI has delivered data engineering and Apache ecosystem training for teams at organisations including the BBC, NHS, RBS, Sky, EDF, and Cisco.
Yes. The Apache ecosystem evolves continuously — with regular Spark releases introducing new features to the DataFrame API, Structured Streaming, and MLlib, and ongoing Kafka developments including updates to Kafka Streams, the KRaft consensus protocol replacing ZooKeeper, and new connector capabilities. JBI's Apache training content is continuously reviewed and updated to reflect the latest stable versions of Spark and Kafka, current Databricks platform features, and evolving best practices in data engineering and real-time streaming. Delegates learn skills that are current and directly applicable to the versions and tools used in professional data engineering environments today.

CONTACT
+44 (0)20 8446 7555

[email protected]

 

Copyright © 2026 JBI Training. All Rights Reserved.
JB International Training Ltd  -  Company Registration Number: 08458005
Registered Address: Wohl Enterprise Hub, 2B Redbourne Avenue, London, N3 2BS

Modern Slavery Statement & Corporate Policies | Terms & Conditions | Contact Us

POPULAR

AI training courses                                                                        CoPilot training course

Threat modelling training course   Python for data analysts training course

Power BI training course                                   Machine Learning training course

Spring Boot Microservices training course              Terraform training course

Data Storytelling training course                                               C++ training course

Power Automate training course                               Clean Code training course