Unlocking Big Data Processing

From Hours to Minutes with Dask and Coiled

Jul 23, 2025


In today’s data-driven environment, has your team ever struggled with:

  • Time-consuming manual data processing

  • Processing bottlenecks that impact productivity

  • Complex data operations that seem to take forever?

If so, you are not alone. So did we, at Siemens Global Business Services (GBS). Facing challenges with slow and manual data processing, the US Digital Consulting Services (DCS) team was able to turn that challenge into a solution by embracing innovation.

How? Two words, Dask and Coiled.

The result: A transformative leap in operational efficiency and scalability, demonstrating the immense potential of distributed computing within Siemens GBS.

The Challenge: Data Bottlenecks Slowing Progress

Needing to analyze a massive dataset of employee training records within the Siemens’ enterprise-wide learning platform, a customer came to GBS requesting assistance to streamline this task. Their existing process was not only inefficient but was overly complex with processing times stretching up to two hours per cycle. In addition to this, manual intervention was required to initiate data processing, creating operational risks and productivity drains.

Recognizing the growing need for speed and automation, it became clear that a new solution was critical to support daily operations without bottlenecks.

The Solution: Distributed Computing with Dask and Coiled

At first, our team explored larger servers and optimized Python scripts, achieving minor improvements. However, true transformation demanded a new approach: distributed computing

We evaluated several Amazon Web Services (AWS) tools, such as Glue (a serverless data integration and processing tool) and Elastic MapReduce (EMR) (a managed cluster platform for large-scale data processing in the cloud) with Apache Spark, (a distributed data processing engine designed for large-scale computing), but they required steep learning curves and major code rewrites.

Instead, Dask and Coiled presented a unique opportunity:

  • Dask, a flexible open-source library for parallel and distributed computing in Python , provides parallel computing within a familiar Pandas-like environment (Pandas is a popular Python library for data analysis). It allowed us to scale workflows from a single machine to a cluster with minimal code changes

  • Coiled, a cloud-based platform for managing and scaling Dask clusters, allows seamless deployment of Dask clusters in the cloud without needing to manage any underlying infrastructure. It handled environment setup, cluster provisioning, and resource optimization—letting us focus entirely on the data processing logic.

By integrating Dask and Coiled into our Apache Airflow environment, a platform for orchestrating data workflows, GBS was able to enhance the data processing capabilities. This implementation was cost-effective as Coiled offered 10,000 free Dask, coiled, Distributed computing Central Processing Unit (CPU) hours per month, which was enough to fully support our project needs without additional cost.

Implementation: Fast, Secure, and Seamless

During this process, we faced challenges such as managing Python package versions and ensuring compatibility across systems. Fortunately, Coiled’s responsive technical support helped us resolve these issues quickly and easily.

Efficiency Redefined

The results demonstrate substantial improvements across multiple aspects of data processing and resource management, setting new standards for performance and productivity. Below are the most notable quantitative achievements and lessons learned that will guide future developments.

Key Achievements and Lessons:

This project highlighted critical lessons for the future of data processing at Siemens:

  • Distributed computing is no longer optional for large datasets

  • Distributed computing is no longer optional for large datasets.

  • Flexibility and innovation drive real performance gains.

  • Given these insights, we are now actively reviewing other projects where distributed computing could bring similar transformational benefits.

    Broader Applications Across Siemens

    The successful use of Dask and Coiled is not limited to training records. This approach is highly adaptable to other big data processes across Siemens, offering significant benefits in cost, scalability, and speed across various industries and operational areas.

    We encourage teams handling large datasets to explore distributed computing early in the project lifecycle — it can redefine what is possible.

    By embracing modern technologies and innovative thinking, the DCS team not only solved a critical client challenge but also expanded our capabilities for future opportunities.

    If your team is seeking to optimize large-scale data processing, let GBS take care of that for you. We already did the heavy lifting, so you don’t have to. Explore tools like Apache Airflow, Dask, and Coiled — and feel free to contact us for support in bringing your data projects to the next level.

    Contact our team today & explore Access GBS

    Interested in reading more about data processing? Check out our other article, Meet the Business Solutions & Services Team.