Fabian Lehmann

Ph.D. candidate

Humboldt-Universität zu Berlin

Biography

I am Fabian Lehmann, a Ph.D. candidate in computer science at the Knowledge Management in Bioinformatics Lab at the Humboldt-Universität zu Berlin. I get my funding through FONDA, a collaborative research center of the German Research Foundation (DFG).

Since my bachelor studies, I have been fascinated by any complex, distributed system. I love to understand and overcome their limits. In my Ph.D. research, I focus on workflow engines, improving the execution of distributed workflows while analyzing large amounts of data. In particular, my goal is to improve scheduling and data management. Therefore, I work closely with the Earth Observation Lab at the Humboldt-Universität zu Berlin to understand real-world requirements.

Interests

Distributed Systems
Scientific Workflows
Workflow Scheduling

Education

Master of Science in Information Systems Management, 2020

Thesis: Design and Implementation of a Processing Pipeline for High Resolution Blood Pressure Sensor Data

Technical University of Berlin
Bachelor of Science in Information Systems Management, 2019

Thesis: Performance-Benchmarking in Continuous-Integration-Processes

Technical University of Berlin
Abitur (comparable to A Levels), 2015

Hannah-Arendt-Gymnasium (Berlin)

Professional Experience

Ph.D. candidate (computer science)

Knowledge Management in Bioinformatics Lab (Humboldt-Universität zu Berlin)

Nov 2020 – Present Berlin, Germany

In my Ph.D. studies, I focus on improving the execution of large scientific workflows processing hundreds of gigabytes of data.

Student Assistent

DAI-Labor (Technical University of Berlin)

May 2018 – Oct 2020 Berlin, Germany

In my student job, we were working with time-series data in DIGINET-PS. For example, we predicted parking slot occupation.

GeoTripNet - Case Study

University of Oxford

Oct 2019 – Mar 2020 Oxford, England, United Kingdom

For the case study, we crawled restaurants' reviews on Google Maps to analyze the relations between different restaurants and examine gentrification in Berlin districts. One problem was to process and analyze the large amount of data in real-time.

Fog Computing Project

Einstein Center Digital Future

Apr 2019 – Sep 2020 Berlin, Germany

This project aimed to analyze SimRa’s bicycle rides. Therefore, we developed a distributed analysis pipeline. Moreover, we visualized the track information on an interactive web. We were able to classify risk hotspots for Berlin’s cyclists' tracks.

Application Systems Project

Conrad Connect

Oct 2017 – Mar 2018 Berlin, Germany

For Conrad Connect, we analyzed hundreds of gigabytes of IoT data. Moreover, I uncovered security vulnerabilities in their software.

Semester Term Work

Reflect IT Solutions GmbH

Mar 2016 – Apr 2016 & Sep 2016 – Oct 2016 Berlin, Germany

In my semester term work, I helped to develop the backend for a construction-progress-management system.

Gap work between school and studies

SPP Schüttauf und Persike Planungsgesellshaft mbH

May 2015 – Sep 2015 Berlin, Germany

Before I started my bachelor studies, I worked a few months, helping to manage a large construction project, gaining experience in dealing with different trades.

Computer skills

A small excerpt

JAVA

Python

Docker

Kubernetes

Spring Boot

Latex

SQL

React

JavaScript

Nextflow

Haskell

Excel

Software

Common Workflow Scheduler

Resource managers can enhance their scheduling capabilities by leveraging the Common Workflow Scheduler interface to receive workflow graph information from workflow systems. This enables the resource manager’s scheduler to make more advanced decisions.

Benchmark Evaluator

The Benchmark Evaluator is a plugin for the Jenkins automation server to load benchmark data and decide on the success of a build accordingly.

Publications

Florian Schintke, Khalid Belhajjame, Ninon De Mecquenem, David Frantz, Vanessa Emanuela Guarino, Marcus Hilbrich, Fabian Lehmann, Paolo Missier, Rebecca Sattler, Jan Arne Sparka, Daniel T. Speckhard, Hermann Stolte, Anh Duc Vu, Ulf Leser

August 1, 2024 Future Generation Computer Systems

Validity constraints for data analysis workflows

Porting a scientific data analysis workflow (DAW) to a cluster infrastructure, a new software stack, or even only a new dataset with some notably different properties is often challenging. Despite the structured definition of the steps (tasks) and their interdependencies during a complex data analysis in the DAW specification, relevant assumptions may remain unspecified and implicit. Such hidden assumptions often lead to crashing tasks without a reasonable error message, poor performance in general, non-terminating executions, or silent wrong results of the DAW, to name only a few possible consequences. Searching for the causes of such errors and drawbacks in a distributed compute cluster managed by a complex infrastructure stack, where DAWs for large datasets typically are executed, can be tedious and time-consuming. We propose validity constraints (VCs) as a new concept for DAW languages to alleviate this situation. A VC is a constraint specifying logical conditions that must be fulfilled at certain times for DAW executions to be valid. When defined together with a DAW, VCs help to improve the portability, adaptability, and reusability of DAWs by making implicit assumptions explicit. Once specified, VCs can be controlled automatically by the DAW infrastructure, and violations can lead to meaningful error messages and graceful behavior (e.g., termination or invocation of repair mechanisms). We provide a broad list of possible VCs, classify them along multiple dimensions, and compare them to similar concepts one can find in related fields. We also provide a proof-of-concept implementation for the workflow system Nextflow.

Mario Sänger, Ninon De Mecquenem, Katarzyna Ewa Lewińska, Vasilis Bountris, Fabian Lehmann, Ulf Leser, Thomas Kosch

June 1, 2024 GigaScience

A qualitative assessment of using ChatGPT as large language model for scientific workflow development

Scientific workflow systems are increasingly popular for expressing and executing complex data analysis pipelines over large datasets, as they offer reproducibility, dependability, and scalability of analyses by automatic parallelization on large compute clusters. However, implementing workflows is difficult due to the involvement of many black-box tools and the deep infrastructure stack necessary for their execution. Simultaneously, user-supporting tools are rare, and the number of available examples is much lower than in classical programming languages.To address these challenges, we investigate the efficiency of large language models (LLMs), specifically ChatGPT, to support users when dealing with scientific workflows. We performed 3 user studies in 2 scientific domains to evaluate ChatGPT for comprehending, adapting, and extending workflows. Our results indicate that LLMs efficiently interpret workflows but achieve lower performance for exchanging components or purposeful workflow extensions. We characterize their limitations in these challenging scenarios and suggest future research directions.Our results show a high accuracy for comprehending and explaining scientific workflows while achieving a reduced performance for modifying and extending workflow descriptions. These findings clearly illustrate the need for further research in this area.

Jonathan Bader, Fabian Lehmann, Lauritz Thamsen, Ulf Leser, Odej Kao

January 1, 2024 Future Generation Computer Systems

Lotaru: Locally Predicting Workflow Task Runtimes for Resource Management on Heterogeneous Infrastructures

Many resource management techniques for task scheduling, energy and carbon efficiency, and cost optimization in workflows rely on a-priori task runtime knowledge. Building runtime prediction models on historical data is often not feasible in practice as workflows, their input data, and the cluster infrastructure change. Online methods, on the other hand, which estimate task runtimes on specific machines while the workflow is running, have to cope with a lack of measurements during start-up. Frequently, scientific workflows are executed on heterogeneous infrastructures consisting of machines with different CPU, I/O, and memory configurations, further complicating predicting runtimes due to different task runtimes on different machine types.
This paper presents Lotaru, a method for locally predicting the runtimes of scientific workflow tasks before they are executed on heterogeneous compute clusters. Crucially, our approach does not rely on historical data and copes with a lack of training data during the start-up. To this end, we use microbenchmarks, reduce the input data to quickly profile the workflow locally, and predict a task’s runtime with a Bayesian linear regression based on the gathered data points from the local workflow execution and the microbenchmarks. Due to its Bayesian approach, Lotaru provides uncertainty estimates that can be used for advanced scheduling methods on distributed cluster infrastructures.
In our evaluation with five real-world scientific workflows, our method outperforms two state-of-the-art runtime prediction baselines and decreases the absolute prediction error by more than 12.5%. In a second set of experiments, the prediction performance of our method, using the predicted runtimes for state-of-the-art scheduling, carbon reduction, and cost prediction, enables results close to those achieved with perfect prior knowledge of runtimes.

David Frantz, Franz Schug, Dominik Wiedenhofer, André Baumgart, Doris Virág, Sam Cooper, Camila Gómez-Medina, Fabian Lehmann, Thomas Udelhoven, Sebastian van der Linden, Patrick Hostert, Helmut Haberl

December 1, 2023 Nature Communications

Unveiling Patterns in Human Dominated Landscapes through Mapping the Mass of US Built Structures

Built structures increasingly dominate the Earth’s landscapes; their surging mass is currently overtaking global biomass. We here assess built structures in the conterminous US by quantifying the mass of 14 stock-building materials in eight building types and nine types of mobility infrastructures. Our high-resolution maps reveal that built structures have become 2.6 times heavier than all plant biomass across the country and that most inhabited areas are mass-dominated by buildings or infrastructure. We analyze determinants of the material intensity and show that densely built settlements have substantially lower per-capita material stocks, while highest intensities are found in sparsely populated regions due to ubiquitous infrastructures. Out-migration aggravates already high intensities in rural areas as people leave while built structures remain – highlighting that quantifying the distribution of built-up mass at high resolution is an essential contribution to understanding the biophysical basis of societies, and to inform strategies to design more resource-efficient settlements and a sustainable circular economy.

Fabian Lehmann, Jonathan Bader, Lauritz Thamsen, Ulf Leser

November 12, 2023 SC-W ‘23: Proceedings of the SC ‘23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis

The Common Workflow Scheduler Interface: Status Quo and Future Plans

Nowadays, many scientific workflows from different domains, such as Remote Sensing, Astronomy, and Bioinformatics, are executed on large computing infrastructures managed by resource managers. Scientific workflow management systems (SWMS) support the workflow execution and communicate with the infrastructures' resource managers. However, the communication between SWMS and resource managers is complicated by a) inconsistent interfaces between SMWS and resource managers and b) the lack of support for workflow dependencies and workflow-specific properties.
To tackle these issues, we developed the Common Workflow Scheduler Interface (CWSI), a simple yet powerful interface to exchange workflow-related information between a SWMS and a resource manager, making the resource manager workflow-aware. The first prototype implementations show that the CWSI can reduce the makespan already with simple but workflow-aware strategies up to 25%. In this paper, we show how existing workflow resource management research can be integrated into the CWSI.

See all publications

Recent Talks

FORCE on Nextflow: Scalable Analysis of Earth Observation data on Commodity Clusters

Modern Earth Observation (EO) often analyses hundreds of gigabytes of data from thousands of satellite images. This data usually is …

Fabian Lehmann

November 1, 2021

Research Projects

FONDA

Foundations of Workflows for Large-Scale Scientific Data Analysis

Contact

fabian.lehmann@hu-berlin.de
+49 (0)30-2093-41285
Building 4, Room IV.426, Rudower Chaussee 25, Berlin, 12489