Niatsu

How to Build Better LCA Databases

Jakob Tresch

If you have ever seen a carbon equivalent on a product or read in a newspaper about it, you might wonder what is actually behind LCA databases. We have been building databases for over three years and I want to bring you the process and our learning behind it a bit closer.

Why even build an LCA database in the first place?

If you are running a sustainability team at a company and you are wondering why someone would even build their own database, if there is already so much available in existing databases, you probably made the right decision to not even consider it. However as we started to build Niatsu, we quickly realized that existing databases often have two problems:

  • They are not specific enough
  • You can not control them

With not specific enough I mean that one footprint might not be fully representative of your own sources. Think of finding a footprint for organic or regenerative produce in an existing database. Often the location coverage is lacking which means that in most databases you match a secondary data point purely on description and not on other attributes. Let alone organic or regenerative agriculture.

Second, you have to trust the maker of a dataset, that its values are correct. If you have full access to the inventory you might be able to modify it, but even there complexity often limits any quick productive results. We have worked with most databases, that are available in the food industry for environmental impacts, and while many provide a vast coverage, they lack the regional granularity and are limited to certain sectors.

Database or system

Today modern databases feel more like systems. A database consists of thousands of static data points stored somewhere. A system adapts to your needs and allows to be updated in the future. As we started building Niatsu, that was exactly what we realized needs to be changed. Instead of having a huge Excel Export with our database, emission factors should be easy to connect to a representation of the supply chain of a product. We do this by separating our database from the way digital twins of products are built within Niatsu. This means that the database can be a bit more flexible and is easier to be extended while benefitting all customers.

A more separated approach allows for new technologies such as AI can easily be integrated for search and processing customer information with the database.

Exactly this is what started to change our approach for building our own database. The Niatsu Database today is a database, but is powered by a lot more knowledge in the back that allows it to scale to more regions. All of this processing is invisible to our customers, but it is a crucial part of being able to serve better data at scale more quickly and also update it.

Carbon Footprint of 1 kilogram of wheat in the Ukraine

0.0

kgCO2e/kg

Wheat

Let’s look at a concrete example of what happens if you select a footprint for wheat in Ukraine. First we looked at satellite data to determine harvest time for this year and together with statistical data we model the yield of wheat in a specific region. This yield is then combined with fertilizer use, which depends on the local agricultural practices, pesticide application and land use change for this crop in that region of Ukraine. My colleague Elissa wrote more about our geospatial pipeline here.

Sources of truth

Every LCA has started at one point by gathering data and putting all of the information together. Step by step meticulously adding piece by piece. Modern database systems need to be prepared for the source of truth changing. In our case this means for example that we update every year yield values for crops to compute their respective footprint.

Our sources of truth contain agricultural production values, regional changes in production methodology (e.g in China changing from small scale farms to larger scale) or updated electricity emissions factors due to increases in renewable electricity.

The challenge with these sources is that they must be accurate, from a current year and most importantly available around the world. This limits the amount of sources we can use for our database but exactly in these decisions we realized, that simplifying your models can also help making them more comparable across the world. The goal is not to always have the best possible model but rather that we can give insights that allow to compare location A with location B.

The user needs to be able to modify

This limitation also helps us to provide datasets that have the potential to be easily modified according to a users specific production system and supply chain. This is in my opinion crucial as this allows us to move away from “secondary data” that has rarely something to do with your actual supply chain to emission factors that represent your actual product and its production. The modification completed by custom input data needs to be easy meanwhile remaining withing the methodological boundaries of all the other datapoints within Niatsu.

Our lessons

  • Separate Data from Business Models (ERPs, Digital Twins, Spending)
  • Plan for Sources of Truth to Change
  • Prefer Comparability over Precision
  • Make it modifiable

The next big database

It’s not just us that have realized this shift, even the leading LCA databases like Ecoinvent are moving towards providing more models at scale using the newest technologies. Ecoinvent 4.0 is expected towards the end of this year. I see a lot of AI environmental matching trying to tackle the market by combining a vast range of datasets to put together product carbon footprints and procurement assessments - this is one way to scale but without tackling the real need for better data. More data is not better data.

The next generation of databases will not only provide footprints, they will provide the infrastructure to tweak emission factors to more accurately represent actual supply chains. They will need to withstand higher compliance criteria as accounting compliance trickles down the supply chain. These databases need to provide the answers that regulations like EmpCo, CBAM and CSRD are trying to establish. A ground truth of what’s actually happening.

For us this means that we invested the last 18 months mostly into building better datasets and models that can be the foundation for the next years. I’m excited about what is ahead.