IT Brief India - Technology news for CIOs & IT decision-makers
India
Cloudera adds GPU acceleration to Apache Spark 4.1 pipelines

Cloudera adds GPU acceleration to Apache Spark 4.1 pipelines

Thu, 20th Aug 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

Cloudera has added native GPU acceleration for Apache Spark 4.1 in Cloudera Data Engineering through a partnership with NVIDIA. The feature will be available as part of Cloudera Anywhere Cloud.

The move targets one of the costliest parts of large-scale data work as companies running AI and analytics projects seek faster ways to prepare data and control cloud spending. The integration uses the NVIDIA cuDF plug-in for Apache Spark, based on CUDA-X software libraries, and allows customers to speed up Spark jobs without rewriting PySpark or SQL code.

Apache Spark is widely used for extract, transform and load processes, as well as broader analytics workloads. On conventional CPU infrastructure, those jobs can run for hours, raising compute bills and slowing the handoff from raw data to machine learning and reporting-ready data sets.

Cloudera said customers using the new setup can see up to 4x acceleration on NVIDIA GPUs compared with traditional CPU-based systems. The software is also designed to work without manual driver configuration while preserving the governance and security controls already used across Cloudera's broader data platform.

The announcement reflects a wider market shift as vendors try to make AI projects less dependent on costly infrastructure expansion. In Cloudera's Great Re-Architecture Survey, 84% of respondents said AI workloads had increased infrastructure costs.

Hybrid focus

A central part of the rollout is support for hybrid environments rather than a single public cloud. Organisations will be able to run the accelerated Spark workloads across public cloud, private cloud, sovereign cloud and on-premises systems while keeping the same operating model, according to Cloudera.

That approach may appeal to large companies that cannot move all their data into one environment because of regulation, residency requirements or existing investments in on-site systems. By tying GPU acceleration to its data engineering service, Cloudera is aiming to reduce the need for bespoke tuning across separate environments.

The feature also expands NVIDIA's role in enterprise data stacks beyond AI model training and inference. GPU use in data engineering is drawing growing interest as businesses look for gains earlier in the AI pipeline, particularly in data cleaning and transformation, where large volumes of information must be processed before use in production systems.

Leo Brunnick, Chief Product Officer at Cloudera, said the main constraint is often data bottlenecks rather than model quality alone.

"For many organizations, AI isn't limited by models. It's limited by how quickly they can turn raw data into trusted, usable insights," said Leo Brunnick, Chief Product Officer at Cloudera. "Accelerating Spark inside Cloudera Data Engineering helps remove that bottleneck, allowing customers to move from data preparation to analytics and AI faster while keeping governance, security, and operational consistency at the center of their strategy."

Cost pressure

Cloud compute spending has become a more prominent concern as AI deployments expand from pilot projects to regular business use. Spark workloads are often among the largest recurring jobs in data estates, so even moderate runtime reductions can materially affect monthly infrastructure bills.

Cloudera is positioning the integration as a way to shorten those runtimes without requiring engineering teams to change established applications. For companies with substantial PySpark and SQL code bases, avoiding rewrites can matter as much as raw speed gains because it removes migration work and lowers operational risk.

NVIDIA said the partnership is intended to fit into existing enterprise practices rather than force changes in how teams manage their environments.

"The fastest path to accelerating AI deployments is the one that aligns with how enterprises already operate today," said Pat Lee, Vice President of Strategic Enterprise Partnerships at NVIDIA. "With NVIDIA AI infrastructure and CUDA-X libraries now native to Cloudera Data Engineering, enterprises can lower costs and dramatically speed up Apache Spark pipelines without changing a single line of PySpark or SQL code, turning business data into a foundation for AI."

For Cloudera, the launch adds another feature to its push to serve customers running mixed cloud and on-premises estates. For NVIDIA, it extends the use of its software libraries into mainstream data engineering workflows, where the business case for acceleration may be easier to measure through reduced runtime and lower compute use.

GPU acceleration for Apache Spark will be available in Cloudera Data Engineering as part of Cloudera Anywhere Cloud, with support aimed at organisations running Spark wherever their data resides.