News-Detail

Big Data and Biodiversity: What Grows, How, Where, and Why?

  • Symbolbild für das Thema Big Data und Biodiversität zeigt Landstraße mit Feld und Blühstreifen, im Vordergrund ein Datennetz als Grafik © Werner/stock.adobe.com | bearbeitet durch hochschule anhalt

Using Data to Protect Biodiversity: An interdisciplinary team is working on this in the “BioTrain” research project. The latest news on climate change or biodiversity is increasingly based on large amounts of data—or data models. But “Big Data and Biodiversity” is still in its infancy. Researchers at Anhalt University of Applied Sciences are testing the extent to which machine learning methods can contribute to predictions in the life sciences—for example, as an early warning system for ecosystems or for forecasting agricultural yields. Prof. Dr. Korinna Bade explains the challenges on Anhalt University of Applied Sciences’ Climate Blog.

Professor Bade, what kind of data do you work with in "BioTrain"?

We’ll look at two different types of data that we intend to examine and then analyze using Data Science methods. The first set consists of data from agricultural field trials, all of which were conducted here in Bernburg-Strenzfeld. For example, the data includes information on how the field was managed, how it was fertilized, how it was tilled—with or without a plow—and, finally, the data also contains details on the yield achieved in the field. This applies to various crops such as corn, winter wheat, and others. In addition, data on the microbial communities in the field were collected, particularly on fungi and bacteria, which were identified through soil samples. And this data was collected over several years—that is, ten years or more.

Porträtbild einer Frau zeigt Prof. Korinna Bade von der Hochschule Anhalt

People often expect Data Science projects to deliver results at the push of a button. But the first step is to bridge the gap and foster a mutual understanding between different fields.

That sounds like a treasure trove of data...

In fact, there is a large amount of data available overall. Especially since we have a second data set labeled “Mobile Links.” This set primarily consists of movement data for animals in two study areas. These were horses on one plot and cattle on the other. We recorded how the animals moved across the plots over time and what impact this had on the vegetation. To this end, plant species were identified at specific points in time. We supplement this data with satellite imagery of the plots. While this data doesn’t provide details, it at least indicates whether an area is green—meaning it’s alive—or brown—meaning it’s been destroyed in terms of vegetation.

How can "big data" be turned into insights for biodiversity and Agriculture?

In addition to diversity, quality plays a major role in ensuring that data can be analyzed using automated machine learning methods at all. This is typically the first challenge a data scientist faces: preparing the data. This is also the case with “BioTrain.” Over the years, the data was collected through various research projects, which involved a great deal of manual work; that is, it wasn’t always collected from the same experimental fields, was sometimes collected at different times or with time gaps, and included different crops, for example.

What does this mean for the work of a data scientist?

The fact that we initially had to put a lot of work into data preparation. However, the lack of continuity in the data comes as no surprise to data scientists. While this makes it more difficult to develop a specific model, it doesn’t make it impossible. We’re currently looking at what the common denominator is, what can be predicted from the data, and what cannot. And if there’s a question from the application domain that the current data set can’t answer, we investigate how we might enrich the data with other sources or use simulation models to generate additional data. Incidentally, this is also a current research topic in machine learning: How can domain knowledge or simulation models be combined with data-driven machine learning methods to train good models? We’re currently working on such an approach for “BioTrain” as well.

  • Symbolbild Big Data und Biodiversität zeigt grün bewachsenes Versuchsfeld der Hochschule Anhalt © HS Anhalt
  • Symbolbild Big Data und Biodiversität zeigt Blühstreifen am Feld mit Wissenschaftlerin, die Daten auf einem Tablet aufnimmt © HS Anhalt

At Anhalt University of Applied Sciences’ experimental fields, researchers are studying various crops as well as measures to promote biodiversity. Over the course of several years, they have collected a wide range of data. The interdisciplinary team is investigating the extent to which this information can be analyzed using machine learning methods as part of the “BioTrain” project.

The overarching goal, however, is an early-warning system for vulnerable ecosystems. How do you plan to achieve that?

We have deliberately not yet taken the step toward making predictions. Before doing so, it is important to first gain insight into the relationships within the data. Biodiversity, in particular, involves a complex interplay of many different components. That’s why we’re initially focusing on analyzing how the individual factors might interact with one another and what aspects of that are reflected in the data. This means we’re currently working at the level of exploratory analysis, looking for statistically significant correlations—for example, between the animals’ movement data and vegetation, or even satellite data. The next steps will then emerge based on these findings. If we succeed in developing an understanding of how these individual factors interact and can substantiate this through the data and the models trained on it, we will have reached an important milestone.

That doesn't sound like something you can just get at the push of a button...

This is very often the expectation for Data Science projects. Yet when it comes to data analysis, the first and most extensive step is usually to first build an understanding of the data and raise awareness within the application domain about what is and isn’t possible with the data. But bringing these different worlds together through mutual understanding is explicitly one of the goals of “BioTrain.” While an early-warning system may be the ultimate goal, it is not, at first, a suitable target for a learning algorithm. It must be broken down into more specific subgoals.

You've also worked on projects in the fields of medicine and logistics, and now you're working on projects in Agriculture and the natural sciences. Is the approach always the same, or does it vary each time?

At the abstract level of computer science, the basic approach is essentially the same. There are the typical steps in the data analysis process that have been described: first, gaining an understanding of the data, followed later by modeling and evaluation. At the same time, you have to continually re-examine the specific data and make sense of it: What does the data mean? How should its quality be assessed? How does it relate to what you want to predict? This is important for selecting the right procedures and methods to use for modeling. In other words, we have an entire portfolio of machine learning methods. And I need to identify which methods are applicable to my problem. This also applies to evaluating the results: When is a result considered good? What criteria do I use to determine that? That, in turn, depends heavily on the domain from which the data originates. And for that, you need the expertise and experience of specialists. Constant coordination throughout the individual process steps is essential. For me, this is precisely what makes Data Science so exciting.

Prof. Dr. Korinna Bade...

... is the Degree Program Advisor for Data Science, a member of the Department Executive Committee, the Deputy Gender Equality Representative, and the director of the Anhalt Center for Data Science

... focuses her research on intelligent data analysis, data mining, machine learning, Data Science, AI, information retrieval, search engine technology, and promoting interest in STEM fields

... is currently conducting research on the following projects: DiLeLA—Digital Learning Labs for Anhalt; AI Engineering—an interdisciplinary, project-oriented bachelor's degree program with a focus on artificial intelligence and engineering; intoMINTgoesLSA; TRAINS Subproject 15—A multiple-unit train demonstrator with H₂ combustion and health monitoring, BioTrain—Training machine learning algorithms: New approaches to analyzing and predicting patterns and correlations in cross-scale biodiversity data

Learn more...

... regarding the interdisciplinary project “BioTrain”: What is the state of an ecosystem? Is there a need for action, and if so, what kind? “BioTrain” exists to provide data-driven answers to these questions. The overarching goal is to generate complex predictions using machine learning algorithms. On the one hand, the project examines the impact of grazing animals on various biological communities. On the other hand, it focuses on microbial communities in arable soils to ultimately enable sustainable soil management. The Federal Ministry of Education and Research is funding the project for a total of three years; researchers from various disciplines are involved: Data Science, agroecology, biodiversity, digital agriculture, and Ecotrophology. Details are available on the project website.

... on research and knowledge transfer in computer science at Anhalt University of Applied Sciences: “Facing the Future with Open Eyes”—this is the headline used by the Department of “Computer Science and Languages” at Anhalt University of Applied Sciences to describe its research projects. Key areas of focus include Data Science, machine learning, digital networking, and human-technology interaction. The Department is particularly committed to inspiring young people to pursue STEM subjects and training them in these fields.

 

Learn more about the use of remote sensing data on Anhalt University of Applied Sciences’ KlimaBlog in the section “Climate Adaptation for Municipalities & Organizations.”

The DFG project AgriRestore also relies on data-driven insights, as explained here on the Anhalt University of Applied Sciences Climate Blog by Prof. Dr. Christina Fischer: Flower Strips in the Agricultural Landscape: Using Research to Increase Acceptance

Redaktion

Claudia Aldinger

Prof. Dr. Korinna Bade

... can be reached for questions and inquiries by phone at +49 (0) 3496 67 3139 or by email at korinna.bade(at)hs-anhalt.de

... Recent publications: 
Holstein, K., Kozaeva, N., Bade, K.: Automated Assessment in the Practical Teaching of Machine Learning Methods, In: A. Greubel, S. Strickroth, and M. Striewe (Eds.): Sixth Workshop on “Automatic Evaluation of Programming Tasks,” Lecture Notes in Informatics (LNI), German Informatics Society, Bonn 2023, DOI: https://doi.org/10.18420/abp2023-10

Mikriukov, G., Schwalbe, G., Hellert, C., Bade, K.: Evaluating the Stability of Semantic Concept Representations in CNNs for Robust Explainability. In: Proceedings of XAI 2023 (Best Industry Paper Award), DOI: 10.1007/978-3-031-44067-0_26

... more about their research topics and responsibilities at Anhalt University of Applied Sciences at the end of the interview.

Cover image: Werner/stock.adobe.com