03.05.2019

Paper presented at CLOSER 2019

Our paper "TASKWORK: A Cloud-aware Runtime System for Elastic Task-parallel HPC Applications" has been presented at the 9th International Conference on Cloud Computing and Services Science (CLOSER 2019).

TASKWORK: A Cloud-aware Runtime System for Elastic Task-parallel HPC Applications

Stefan Kehrer, Wolfgang Blochinger

Abstract: With the capability of employing virtually unlimited compute resources, the cloud evolved into an attractive execution environment for applications from the High Performance Computing (HPC) domain. By means of elastic scaling, compute resources can be provisioned and decommissioned at runtime. This gives rise to a new concept in HPC: Elasticity of parallel computations. However, it is still an open research question to which extent HPC applications can benefit from elastic scaling and how to leverage elasticity of parallel computations. In this paper, we discuss how to address these challenges for HPC applications with dynamic task parallelism and present TASKWORK, a cloud-aware runtime system based on our findings. TASKWORK enables the implementation of elastic HPC applications by means of higher-level development frameworks and solves corresponding coordination problems based on Apache ZooKeeper. For evaluation purposes, we discuss a development framework for parallel branch-and-bound based on TASKWORK, show how to implement an elastic HPC application, and report on measurements with respect to parallel efficiency and elastic scaling.

The paper has been published in the proceedings of the 9th International Conference on Cloud Computing and Services Science (CLOSER 2019).

Link:DOI