Causal Graph Discovery for Soil Carbon Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current process-based models for studying soil carbon changes are complex, require expert calibration, and struggle with accurately predicting the impact of weather and field management practices due to reliance on spurious correlations rather than cause-and-effect relationships, especially when real data is sparse.

Innovation Solution

A data-driven method that learns a causal graph from heterogeneous data sets, combining real and simulated data to identify true cause-and-effect relationships, allowing for improved modeling and decision-making without the need for domain expertise, using techniques like graph neural networks and variational auto-encoders to enhance process-based models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If process-based models are used to study soil carbon changes, then understanding of biological and natural systems is improved, but model complexity and requirement for expert calibration increase

Engineering Contradiction:
Improveaccuracy of soil carbon predictionVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces causal discovery algorithms as an intermediary between process-based models and data. This intermediary automatically learns causal relationships from heterogeneous data sources and uses them to guide and simplify process-based models, reducing the need for expert calibration while maintaining prediction accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual expert calibration processes with automated machine learning algorithms. Instead of requiring domain experts to manually adjust model parameters and structures, the system uses causal discovery algorithms to automatically learn from data, substituting the mechanical expert calibration process with an automated computational approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If process-based models rely on spurious correlations, then modeling is simplified, but prediction accuracy deteriorates

Engineering Contradiction:
Improveease of modelingVSAvoidprediction accuracy
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent substitutes correlation-based modeling approaches with causal discovery algorithms. Instead of relying on simple statistical correlations that are easier to compute but less reliable, the system uses advanced causal inference methods that automatically identify true cause-effect relationships, improving prediction accuracy while maintaining computational feasibility.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces causal discovery algorithms as an intermediary layer between raw data and process-based models. This intermediary automatically distinguishes between spurious correlations and true causal relationships, providing refined input to the modeling process that improves accuracy without requiring manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If real data is used for model calibration, then model accuracy is improved, but data sparsity and collection costs increase

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata availability
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent merges multiple data sources including real observational data, simulated data from process-based models, and data from various sensors. By combining these heterogeneous data sources, the system overcomes the sparsity of individual real data sources while maintaining accuracy through causal discovery algorithms that can distinguish between different data types and their relationships.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses causal discovery algorithms as an intermediary to effectively integrate sparse real data with simulated data. The algorithm learns causal relationships from the combined heterogeneous data, allowing the system to make accurate predictions even when real data is limited, by leveraging the complementary information from multiple sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of information

If domain expertise is required for model calibration, then model understanding is improved, but automation and scalability deteriorate

Engineering Contradiction:
Improvedomain knowledge utilizationVSAvoidautomation level
Core Design Contradiction:
Loss of informationVSExtent of automation

Solution Approach 1:

The patent enables the system to self-calibrate using automated causal discovery algorithms. Instead of requiring domain experts to manually calibrate models, the system automatically learns causal relationships from data and uses this knowledge to configure and improve process-based models, achieving both high automation and effective utilization of domain knowledge embedded in the data.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent substitutes manual expert calibration with automated machine learning systems. The causal discovery algorithms automatically perform the function that previously required domain expertise, replacing the manual calibration process while capturing and utilizing domain knowledge that is encoded in the relationships present in the data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240046073A1Data driven approaches to improve understanding of process-based models and decision making
Publication Date: 2024.02.08 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240046073A1 patent drawing
  • US20240046073A1 patent drawing
  • US20240046073A1 patent drawing

AI summary

This disclosure provides a data-driven and scalable method to discover cause-and-effect relationships in data from natural systems that include sparse data sets. This technique can learn a causal graph from heterogenous data sources by combining embeddings from real data and embeddings from simulated data generated by process-based models. The causal graph is used for what-if analysis in out-of-distribution settings. One application is understanding the factors that affect soil carbon. A causal model created by these techniques can be used to discover cause-and-effect relationships that affect soil carbon. This model has applications such as forecasting soil carbon for a future time point to help inform farm practices. Farm practices, like tilling, may be modified in response to predictions provided by the model.