A physical information neural network optimized soil carbon storage precise evaluation system
By optimizing the physical information neural network system and integrating multi-source heterogeneous data, a high-precision soil carbon storage assessment across time, space and region was achieved, solving the problems of insufficient assessment accuracy and applicability in existing technologies and providing scientific decision support and management tools.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
- Filing Date
- 2026-03-23
- Publication Date
- 2026-05-29
Smart Images

Figure CN122114383A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of artificial intelligence and earth science, and in particular to a precise soil carbon storage assessment system optimized by physical information neural network. Background Technology
[0002] Soil carbon storage is a key component of the global carbon cycle, significantly impacting climate change, ecosystem stability, and agricultural productivity. Accurate assessment and forecasting of soil carbon storage are crucial for developing sound carbon reduction policies and optimizing land management.
[0003] A published patent, "Method, Apparatus, Equipment, and Storage Medium for Assessing Soil Carbon Storage in Mangrove Wetlands" (Publication No.: CN119886874A), includes: acquiring remote sensing image data and classifying land cover in the area to be assessed; and measuring hyperspectral reflectance data of soil at different depths using a land cover spectrometer. This invention utilizes remote sensing image data to determine sample points, and combines the hyperspectral data measured on-site at these sample points with an assessment model to determine soil carbon storage at different depths, achieving a comprehensive and accurate assessment of carbon storage for the entire ecosystem.
[0004] The aforementioned published patents still have room for improvement in terms of the integrity of driving factors, deep mechanism modeling, model generalization and adaptation, spatial heterogeneity characterization, prediction uncertainty quantification, temporal dynamic characterization, and multi-library collaborative evaluation. They are difficult to achieve stable and high-precision soil carbon storage assessment in cross-temporal and cross-regional scenarios. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a precise soil carbon storage assessment system optimized by a physical information neural network.
[0006] To achieve the above objectives, the present invention employs the following technical solution: a soil carbon storage precision assessment system optimized by a physical information neural network, comprising a multi-source heterogeneous data intelligent fusion module for fusing multi-source heterogeneous data from satellite hyperspectral imagery, UAV lidar point clouds, microwave remote sensing, ground sensors, electromagnetic and thermal sensing, optical and acoustic monitoring, metagenomic sequencing, XRD mineralogy, hydrology, and DEM through multimodal spatiotemporal alignment, multi-physics field coupled sensing, bio-mineral map construction, and dynamic hydrological topology construction, outputting a consistent feature matrix, a hydrological driving factor matrix, and a carbon stability biological index; a hierarchical physical constraint neural network module for embedding multi-physics mixed PINN, enzyme kinetic time delay constraints, and an electrochemical-diffusion synergistic layer into a deep network to achieve unified modeling of surface and deep soil carbon cycle processes, outputting hierarchically consistent carbon flux and carbon pool prediction results; and a multi-scale three-dimensional carbon storage mapping module for topology-enhanced fractal kriging interpolation. The system utilizes a value-based, holographic soil mapping neural network and the Monte Carlo method for uncertainty decomposition to generate a 3D carbon storage voxel map on a centimeter-level grid, outputting confidence intervals, sensitivity heatmaps, and quality labels. A temporal evolution and digital twin module is used to achieve historical assimilation, future evolution projection, and stability risk assessment of soil carbon storage through a self-evolving carbon cycle cellular automata, a physical-transformer dual-channel state space, and multi-indicator stability early warning. An adaptive transfer learning module is used to achieve rapid adaptation and high-precision prediction of the system under different regions and soil types through cross-mechanism adaptive meta-learning, physically consistent federated distillation, and dynamic multi-scale constraint adjustment mechanisms. A decision support and visualization module is used to map the prediction results to a VR / AR interactive environment based on immersive twin interaction, a cross-modal interpretable framework, and blockchain-based trusted traceability visualization, outputting factor contribution ranking and a carbon storage traceability chain to assist researchers, policymakers, and ecological restoration managers in decision-making.
[0007] As a further description of the above technical solution: The multimodal spatiotemporal alignment in the multi-source heterogeneous data intelligent fusion module adopts a hierarchical autoregressive attention mechanism, using ground sensor sequences as the time reference, and performs temporal offset correction and spatial reprojection optimization on satellite imagery, UAV point clouds and microwave remote sensing data.
[0008] As a further description of the above technical solution: The multi-physics hybrid PINN in the hierarchical physical constraint neural network module introduces the Arrhenius temperature effect and water vapor transport equation in the surface layer, and adopts the Langmuir adsorption kinetic equation in the deep layer. It also achieves consistency between the surface and deep layer prediction results through cross-layer residual connections.
[0009] As a further description of the above technical solution: The enzyme kinetic time delay constraint in the hierarchical physical constraint neural network module uses the Michaelis-Menten kinetic model combined with time delay differential terms to describe the lag response of enzyme synthesis to environmental factors, thereby constraining the prediction of carbon decomposition rate.
[0010] As a further description of the above technical solution: The holographic soil map neural network in the multi-scale three-dimensional carbon storage mapping module constructs a three-dimensional multilayer map of granular layer, pore layer and capillary water layer based on CT or MRI scan data, and simulates the diffusion and consolidation of dissolved organic carbon through cross-layer information propagation mechanism.
[0011] As a further description of the above technical solution: The self-evolving carbon cycle cellular automata in the temporal evolution and digital twin module uses microorganisms, minerals, and organic carbon as intelligent agents. Through reinforcement learning to train interaction rules, it realizes a dynamic evolutionary pattern of carbon sequestration, mineralization, and migration.
[0012] As a further description of the above technical solution: The physical-transformer dual-channel state space in the temporal evolution and digital twin module consists of a data-driven channel and a physical equation channel. The data-driven channel encodes multi-source temporal observations, while the physical equation channel uses the state transition matrix to constrain the predicted trajectory and simulates the impact of different management measures on carbon storage under the counterfactual branch.
[0013] As a further description of the above technical solution: The adaptive transfer learning module introduces carbon conservation and energy conservation loss functions when aggregating models across regions through physical consistency federated distillation, thereby ensuring the physical consistency of models in different regions while protecting the privacy of monitoring data.
[0014] As a further description of the above technical solution: The cross-modal interpretable framework in the decision support and visualization module uses Shapley decomposition to extend to multimodal inputs, quantifies the marginal contributions of remote sensing bands, microbial genes, mineral types and environmental factors to carbon storage prediction, and outputs factor ranking and weight map.
[0015] As a further description of the above technical solution: The blockchain-based trusted traceability visualization module in the decision support and visualization module records the entire process of sampling, modeling, and prediction on the blockchain through hash locking to ensure the transparency and immutability of carbon sink data in scientific research, policy, and carbon trading scenarios.
[0016] The present invention has the following beneficial effects: 1. This invention firstly integrates data from more than ten different types and scales, including satellite remote sensing, UAVs, ground sensors, metagenomic sequencing, and XRD mineralogy, through multimodal spatiotemporal alignment, multiphysics coupled sensing, bio-mineral map construction, and dynamic hydrological topology construction. This overcomes the data heterogeneity problem of existing technologies, providing comprehensive and consistent input for subsequent high-precision modeling and enhancing the ability to comprehensively capture the driving factors of soil carbon cycle. Furthermore, by introducing multiphysics hybrid PINN, enzyme kinetic time-delay constraints, and an electrochemical-diffusion synergistic layer, key physical, chemical, and biodynamic equations of soil carbon cycle are embedded in a deep learning framework, making the model not only... It is data-driven and mechanism-constrained, significantly improving the model's interpretability, generalization ability, and prediction accuracy in complex environments. It achieves unified modeling and consistent prediction of carbon cycle processes in surface and deep soils, solving the problem that existing models cannot take into account carbon processes at different depths. Combining topologically enhanced fractal kriging interpolation, holographic soil map neural networks, and uncertainty decomposition Monte Carlo methods, it can generate three-dimensional carbon storage voxel maps at centimeter-level resolution, greatly improving the precision of the assessment. The system can output confidence intervals, sensitivity heatmaps, and quality labels, quantifying the uncertainty of the assessment results, enhancing the reliability and practicality of the assessment results, and providing data support for refined management.
[0017] 2. In this invention, through a self-evolving carbon cycle cellular automaton and a physical-transformer dual-channel state space, the system can realize the historical assimilation and future evolution projection of soil carbon storage. In particular, it simulates the impact of different management measures on carbon storage under a counterfactual branch, providing policymakers and ecological restoration managers with a powerful scenario analysis tool. This achieves a digital twin of the soil carbon cycle, significantly improving the scientific rigor and foresight of decision-making. The introduction of cross-mechanism adaptive meta-learning and physically consistent federated distillation technology enables the system to dynamically adapt to different regions and soil types, and through federated learning... While protecting privacy, the system aggregates multi-source data knowledge, significantly improving the model's generalization ability and regional applicability, and reducing the cost of model deployment and maintenance. Based on immersive twin interaction, a cross-modal interpretable framework, and blockchain-based trusted traceability visualization, the system intuitively maps complex prediction results to a VR / AR environment and provides factor contribution ranking and carbon storage traceability chains. Shapley decomposition expands and quantifies the contribution of each driving factor to the prediction, and blockchain technology ensures the transparency and immutability of carbon sink data, greatly assisting scientific decision-making and trust building in scenarios such as scientific research, policy making, and carbon trading. Attached Figure Description
[0018] Figure 1 This is a side view of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Reference Figure 1 This invention provides an embodiment of a soil carbon storage precision assessment system optimized by a physical information neural network. The system includes a multi-source heterogeneous data intelligent fusion module. This module uses multi-modal spatiotemporal alignment, multi-physics field coupled sensing, bio-mineral map construction, and dynamic hydrological topology construction to fuse multi-source heterogeneous data from satellite hyperspectral imagery, UAV lidar point clouds, microwave remote sensing, ground sensors, electromagnetic and thermal sensors, optical and acoustic monitoring, metagenomic sequencing, XRD mineralogy, hydrology, and DEM. The system outputs a consistent feature matrix, a hydrological driving factor matrix, and a carbon stability biological index. A hierarchical physical constraint neural network module embeds multi-physics mixed PINN, enzyme kinetic time-delay constraints, and an electrochemical-diffusion synergistic layer into a deep network to achieve unified modeling of surface and deep soil carbon cycle processes, outputting hierarchically consistent carbon flux and carbon pool prediction results. A multi-scale three-dimensional carbon storage mapping module uses topology-enhanced fractal kriging interpolation and holographic soil mapping. The soil map neural network and the Monte Carlo method for uncertainty decomposition generate a 3D carbon storage voxel map on a centimeter-level grid, outputting confidence intervals, sensitivity heatmaps, and quality labels. The temporal evolution and digital twin module is used to realize the historical assimilation, future evolution projection, and stability risk assessment of soil carbon storage through a self-evolving carbon cycle cellular automaton, a physical-transformer dual-channel state space, and multi-index stability early warning. The adaptive transfer learning module is used to achieve rapid adaptation and high-precision prediction of the system under different regions and soil types through cross-mechanism adaptive meta-learning, physical consistency federated distillation, and dynamic multi-scale constraint adjustment mechanisms. The decision support and visualization module is used to map the prediction results to a VR / AR interactive environment based on immersive twin interaction, a cross-modal interpretable framework, and blockchain trusted traceability visualization, outputting factor contribution ranking and carbon storage traceability chain to assist researchers, policymakers, and ecological restoration managers in decision-making.
[0021] The goal of the multi-source heterogeneous data intelligent fusion module is to achieve comprehensive, multi-dimensional, and multi-scale observation and fusion of driving factors affecting soil carbon storage, ensuring the consistency of multi-source heterogeneous data in time, space, and physical field dimensions. Through intelligent algorithms, it generates high-quality feature inputs for subsequent modeling, unifying data from satellites, UAVs, microwave remote sensing, ground sensors, electromagnetic and thermal sensors, optical and acoustic monitoring, metagenomic sequencing, XRD mineralogy, hydrology, and DEM into a temporally consistent, spatially aligned, and physically comparable feature system. This provides high-quality input for subsequent physical constraint modeling, temporal twin evolution, and decision support. Inputs include: satellite hyperspectral imagery, UAV lidar point clouds, surface soil moisture content retrieved from microwave remote sensing, ground sensor time series (temperature, air humidity, soil moisture content, salinity, electrical conductivity), electromagnetic conductivity sensor records, thermal flux plate records, optical color indices and vegetation phenotypes, acoustic reflectance and attenuation curves, metagenomic sequences and functional gene annotations, XRD mineral peaks, and DEM and watershed hydrological observations. Output: A unified, consistent feature matrix across modalities, where each row corresponds to the multimodal features of the same grid or sample point at the same time; synchronously outputting the hydrological driving factor matrix and the carbon stability biological index (BSI) raster or sample point sequence, along with uncertainty metrics and quality labels. Multimodal spatiotemporal alignment receives remote sensing data such as satellite hyperspectral imagery, UAV lidar point clouds, and microwave remote sensing data, as well as sequence data from ground sensors (e.g., soil temperature and humidity sensors, CO flux towers). A hierarchical autoregressive attention mechanism is employed, using the ground sensor sequence as the time reference, to perform temporal offset correction and spatial reprojection optimization on satellite imagery, UAV point clouds, and microwave remote sensing data. This mechanism captures the complex temporal dependencies between different data sources and precisely adjusts the acquisition time deviation of remote sensing data based on ground truth. Simultaneously, high-precision georegistration algorithms (e.g., image registration based on high-density control points) are used to resample and transform spatial data, ensuring that all data are aligned within a unified spatiotemporal framework, thus resolving the spatiotemporal inconsistency problem of multi-source data. The output is the aligned multimodal temporal features. Multi-physics coupled sensing integrates data from electromagnetic and thermal sensors (e.g., soil dielectric constant, thermal conductivity), optical and acoustic monitoring (e.g., vegetation index, soil color, acoustic characteristics of soil structure), and hydrological and DEM data (e.g., slope, aspect, water flow path). By constructing mutual information models and causal relationship diagrams among the multi-physics fields, the interactions between different physical fields (e.g., the influence of soil moisture on heat conduction, the influence of topography on water flow distribution) are perceived, generating an integrated physical field feature vector. Biological-mineral mapping processes metagenomic sequencing data (revealing soil microbial community structure and functional genes) and XRD mineralogy data (identifying soil mineral composition and crystal structure).Using bioinformatics tools (such as OTU clustering and functional gene prediction) and quantitative mineral analysis methods, we constructed carbon cycle functional maps of microbial communities and maps of mineral adsorption mechanisms for organic matter, generating carbon stability biological indices and mineral influence characteristics. Dynamic hydrological topology construction, based on DEM, rainfall, and evapotranspiration data, utilized hydrological models (such as TOPMODEL or SWAT models) to construct a three-dimensional dynamic topological structure of soil water flow and the dynamic changes in surface and deep soil moisture content. This helps to understand the impact of water flow paths on dissolved organic carbon migration and aggregate formation.
[0022] Multimodal spatiotemporal alignment is used to unify multimodal data from different observation platforms: satellite hyperspectral imagery acquires large-scale surface vegetation cover and soil spectral characteristics to characterize regional vegetation growth and soil properties. UAV lidar point clouds provide centimeter-level precision information on vegetation canopy height, topographic relief, and soil surface roughness, compensating for the limitations of satellite imagery at the local scale. Microwave remote sensing can stably acquire surface soil moisture dynamics in cloudy and humid environments, ensuring continuous observation capabilities under extreme climates. Ground sensor networks continuously collect time-series data on soil temperature, air humidity, soil moisture content, salinity, and conductivity to supplement the ground reality in remote sensing imagery. A hierarchical autoregressive attention mechanism is used to process the above multi-source data, automatically identifying and eliminating temporal delays and spatial misalignments between different modalities, ultimately generating a consistent feature matrix to ensure data fusion at the spatiotemporal scale. Different platforms have inconsistent sampling times, and there may even be lags. Hierarchical autoregressive attention identifies and compensates for time biases, aligning the same natural process on the same timeline. Using ground sensor sequences as a time base, windowed historical segments are constructed, mapping timestamps from satellites, UAVs, and microwave remote sensing to the nearest reference window, and allowing learnable time offsets. Autoregressive attention is used to predict aligned observations, while the time offset is output as an interpretable correction. Time attention formula: Offset formula: , Aligned at the reference time Cross-modal feature vectors; :by A set of time windows centered on the subject; Attention weights represent the historical moments within the window. For time Contributions; : Query and key vectors obtained from learnable linear layers; : From non-reference mode in offset Post-observational characteristics; Integer time offset, achieved by minimizing the time offset from the baseline sequence. The differences are determined; The L2 norm projects data with different spatial resolutions onto a unified grid, correcting offsets and distortions. Initial registration is performed using geographic coordinates and DEM control points, followed by fine registration using feature pairs (corners, edges, textures), and joint optimization of point clouds and imagery is used to correct minor distortions. Spatial reprojection error formula: Weighted fusion formula: , : Pixel or point cloud coordinates (homogeneous coordinates) of the source data. Geometric transformation matrices (translation, rotation, scaling, perspective). Imaging projection operator, : Coordinates after reprojection : Reference control points in the target coordinate system Single-point reprojection error Spatial alignment features after weighted fusion : No. The characteristics of a mode at this location, Modal weights: the smaller the error, the greater the weight. Temperature coefficient: controls the weights' sensitivity to error. Number of modalities, formula for total training or calibration loss: , Total loss during training or calibration. , Aligned features and baseline features Spatial gradient operator, used to penalize unreasonable spatial jumps. :Location The fusion characteristics Mapping from features to physically observable quantities (e.g., predicting leaf area index from spectra). The independent observation or experimental measurement value of this physical quantity. Weighting coefficient.
[0023] Multiphysics coupled sensing is used to establish a joint observation system encompassing optical, acoustic, electrical, and thermal dimensions. Electromagnetic conductivity sensors are used to measure soil ion concentration and mineral structure information, thereby revealing the soil's internal physicochemical properties. Apparent conductivity is used to infer ion concentration and mineral framework structure, providing constraints for organic carbon adsorption and stabilization. A linearized representation of the Archie relation is used to characterize the formula. , Effective conductivity, Porosity Water saturation The regression parameter, obtained through experimental calibration, links electrical conductivity with porosity and water content, indirectly reflecting ion channels and the mineral framework. The heat flux plate is used to determine the thermal diffusivity of soil, characterizing the energy consumption during carbon mineralization. Carbon mineralization involves energy exchange, and thermal diffusivity reflects microbial metabolism and soil hydrothermal processes. The formula for estimating thermal diffusivity is: , Thermal diffusivity Thermal conductivity Soil volume density Specific heat capacity at constant pressure Vertical heat flux Vertical temperature gradient. By optically monitoring surface vegetation and soil color, acoustically monitoring subsurface structural response, electrically monitoring ion distribution, and thermally monitoring temperature gradient, a comprehensive modeling system integrating light, sound, electricity, and heat is formed. This achieves full coverage of soil carbon cycle driving factors, and characterizes organic matter content, vegetation cover, and soil surface condition through spectral and color indices. Normalized Difference Vegetation Index (NDVI) formula: Example formula for earth tone index: , Near-infrared and red band reflectivity Visible light three-channel reflectance Vegetation index A simplified soil color index is used to characterize surface color shifts and identifies pore structure and water content variations through excitation and echo attenuation. The simplified attenuation model formula is as follows: , The propagation distance is Echo amplitude at time Initial incentive magnitude The effective attenuation coefficient is affected by pore structure and saturation. Propagation distance, weighted fusion formula based on physical consistency: , , The fundamental eigenvectors obtained through physical calculations. Electrical conductivity, thermal diffusivity, and acoustic attenuation are estimated or inverted using formulas D, E, and G. Vector concatenation operation : No. A nonlinear submap (e.g., a small feedforward network or kernel function). Learnable fusion weights Advanced features resulting from the fusion of light, sound, electricity, and heat.
[0024] Biological-mineral mapping was used to reveal the interactions between microbial communities and soil minerals. Metagenomic sequencing was used to extract functional genes related to carbon cycling. X-ray diffraction (XRD) experiments were used to determine the types of clay minerals in the soil, including kaolinite, montmorillonite, and illite, to reveal the structural characteristics of soil minerals. Based on graph attention networks, a bipartite gene-mineral graph was constructed by linking microbial functional genes and mineral types. The biological carbon stability index (BSI) was calculated using the graph structure to reflect the contribution of different microbial-mineral interactions to organic carbon sequestration. A set of functional genes directly related to carbon cycling (cellulase family genes, methane monooxygenase genes, lignin-degrading enzyme genes, and polymer degradation-related coenzyme genes) was extracted from metagenomic annotations. Clay mineral types (kaolinite, montmorillonite, illite, and chlorite) and their relative abundances were obtained from XRD. A bipartite gene-mineral graph was constructed, with genes as nodes on one side and minerals as nodes on the other side. The edge weights represent the intensity of gene interaction or co-occurrence on the mineral surface. The bipartite graph attention aggregation formula is as follows: , , :node Input features (gene expression abundance or functional score, mineral crystallinity or interlayer cation exchange capacity). : Node representation after attention aggregation :node The neighborhood group, Linear transformation matrix Attention parameter vector Vector concatenation From the neighbor arrive Attention weights Nonlinear activation function, carbon stability biological index formula: , Carbon stability biological index : Set of functional gene nodes Mineral node set : Gene-mineral interaction weights obtained through attention learning. : No. The stabilization tendency coefficient of each gene (scored by literature or experiments, which can be quantified as the strength of its positive or negative effect on carbon sequestration). : No. The adsorption-protection capacity coefficient of a mineral (which can be obtained by combining layer charge, specific surface area, and cation exchange capacity). Genes at the sampling site With minerals The co-occurrence intensity or probability of action (from co-occurrence analysis or binding degree experiment) indicates that the higher the BSI, the greater the contribution of the microbial-mineral mechanism to the stabilization of organic carbon in the soil.
[0025] Dynamic hydrological topology construction is used to characterize the impact of hydrological processes on soil carbon migration: Based on a digital elevation model (DEM) combined with watershed hydrological observation data, a dynamic hydrological map is generated to reflect topography and flow paths. Tidal fluctuations, groundwater level rises and falls, and surface runoff paths are simulated to characterize the effects of hydrodynamic processes on carbon migration and leaching. The output hydrological driving factor matrix includes: TWI (hydrological indices), flow direction indices, and soil permeability parameters. These matrices are used as model inputs to characterize the temporal and spatial control of hydrological conditions on the carbon cycle. Slope and runoff accumulation are calculated from the DEM to determine surface runoff paths. By integrating tides, groundwater levels, and rainfall, time-varying hydrodynamic processes are simulated to generate hydrological indices, flow direction codes, and permeability parameter rasteres. The formula for the topographic humidity index (TWI) is as follows: , Topographic humidity index : The area of runoff contribution per unit contour line length Surface slope angle: A larger TWI indicates more water catchment and a gentler slope, which is often wetter and affects carbon leaching and migration. Flow direction coding formula: , Pixel Where the water flows DEM elevation function Eight-neighbor directional displacement, selecting the direction with the largest elevation drop as the mainstream direction, Green-Ampt infiltration approximation formula: , , :time infiltration rate, : Saturated hydraulic conductivity Moisturizing suction head. : Saturated volumetric water content Initial volumetric moisture content, The cumulative infiltration rate determines the vertical and lateral transport capacity of dissolved organic carbon. The formula for the hydrological driving factor matrix is as follows: , Hydrological factor vectors on each grid-time pair, groundwater depth from observation wells or inversion, tidal amplitude for coastal and estuarine areas, and effective rainfall refers to net infiltration rainfall after removing primary losses.
[0026] The multi-source heterogeneous data intelligent fusion module forms a comprehensive observation system covering the entire spectrum from the surface to deep soil, from micro-genes to macro-landforms, and from physical properties to biological processes through multimodal spatiotemporal alignment, multi-physics field coupled sensing, bio-mineral map construction, and dynamic hydrological topology construction. The core innovation of this module lies in introducing a hierarchical autoregressive attention mechanism for spatiotemporal consistency fusion, establishing a light-sound-electric-thermal multi-physics field sensing framework, and combining metagenomics and mineralogy to form an interpretable carbon stability map. Ultimately, it outputs high-quality, multi-scale, and multimodal consistent features, providing a solid data foundation for subsequent physical constraint modeling, temporal twin evolution, and decision support. It defines a unified grid resolution (e.g., 10 meters or 30 meters) and time step (e.g., day or week), aligns and resamples all modalities, forming multimodal feature rows for each grid-time pair. The final feature concatenation formula is as follows: Standardized formula: , :time ,Location The original splicing characteristics, Features of spatiotemporally aligned remote sensing and ground observation fusion Advanced features of multiphysics fusion Biomineral Index Hydrological factors Standardization operator (zero mean, unit variance, or quantile scaling). The final consistency feature used for downstream modeling, and the bias correction formula for cross-calibration: , Calibration value of sensor A : Sensor B reading at the same location and time. The proportion and bias parameters are obtained by fitting the data onto overlapping observations using least squares. The uncertainty synthesis formula under the linear approximation is as follows: , A derived quantity calculated from multimodal observations. : The gradient vector of the input quantity : Input the covariance matrix of the observations, The variance estimation of the derived quantities is used to generate quality labels. A common projected coordinate system and unified timestamps are used. All data are first coarsely aligned and then finely registered. First, intra-modal spatiotemporal alignment is performed, followed by cross-modal fusion. Finally, hydrological and biological-mineral information is overlaid. For each grid-time pair, a feature vector is output. and quality labels, time offset Spatial reprojection error Hydrological factors Compared with BSI, physical quantity observation With loss Continuous calibration ensures consistency between remote sensing, ground, and physical systems.
[0027] The hierarchical physical constraint neural network module mainly consists of three parts: a multi-physical hybrid PINN, an enzyme kinetic time-delay constraint, and an electrochemical-diffusion synergistic layer. It embeds verifiable mechanisms of key soil carbon cycle processes—temperature effects, water and air transport, mineral adsorption / desorption, enzymatic reactions, redox reactions, and pore diffusion—into the neural network, making predictions both accurate and interpretable, and maintaining physical consistency between the surface and deep layers. Input (consistency features from the multi-source heterogeneous data intelligent fusion module): temperature time series. Volumetric moisture content Porosity saturated hydraulic conductivity Dissolved organic carbon concentration near the surface and deep layers Mineral type and capacity parameters, redox potential The model incorporates observations or constraints on fluxes and concentrations at the boundary, and initial carbon pool components. Outputs include various carbon fluxes and carbon pool quantities (litter decomposition rate, mineral-bound carbon change rate, dissolved phase migration flux) on grid-time pairs, constrained residuals for stratified consistency, and interpretable physical factor contribution decomposition. Multi-physics hybrid PINN is used to construct deep learning models of soil carbon cycling. It incorporates residual terms from physical equations into the traditional neural network loss function. Addressing the differences between surface and deep soil carbon cycling: In the surface carbon cycle model, the Arrhenius temperature effect and water-gas transport equations are introduced. These equations are encoded in the form of partial differential equations (PDEs) as physical constraints for PINN. In the deep carbon cycle model, the Langmuir adsorption kinetic equation is used. This reflects the important role of minerals in carbon sequestration in deep soil layers. Through cross-layer residual connections, a smooth transition and consistency between surface and deep prediction results are achieved, ensuring that the predicted carbon fluxes and carbon pools across the entire soil profile are physically coherent and coordinated. Enzyme kinetic time-lag constraints further refine the biokinetics of carbon decomposition. They employ variations in the Michaelis-Menten kinetic model, resulting in a lag in the carbon decomposition rate. This time-lag constraint is introduced into the PINN loss function through a regularization term, enabling the model to more accurately predict microbial-mediated carbon decomposition rates. Electrochemical-diffusion synergy focuses on ion transport and mass diffusion in the soil microenvironment. Simulations are performed on the diffusion of dissolved organic carbon and nutrient ions in soil pore water, considering the influence of redox potential on carbon transformation. This helps explain the mechanisms of carbon migration and transformation under different soil depths and moisture conditions.
[0028] In topsoil, the temperature sensitivity of litter decomposition is combined with water vapor transport processes. Specifically, by introducing an Arrhenius temperature effect model reflecting the temperature dependence of the decomposition rate, and combining it with the topsoil water vapor transport equation, the organic matter decomposition rate and water vapor exchange processes are constrained. The Arrhenius temperature effect constrains the organic matter decomposition rate in the topsoil, while the water vapor transport equation limits the influence of moisture and gas phase diffusion on the decomposition supply. Arrhenius temperature effect and humidity correction formula: , : Decompose the apparent rate constant, Reference temperature The rate constant under the condition, Apparent activation energy Gas constant Absolute temperature Reference absolute temperature Moisture correction factor (a monotonic function reflecting the promoting or inhibiting effect of moisture on decomposition; its specific form can be fitted using a parameterized function during training); mass conservation formula for the surface organic matter pool: , : The concentration or content of decomposable organic carbon on the surface. Litter input flux Leaching loss term (determined by hydrological and diffusion processes, subsequently coupled), simplified water transport formula: Simplified gas transport formula: , Volumetric moisture content Hydraulic conductivity related to moisture content Water potential (the potential function of the sum of gravitational potential and matrix potential). The concentration of gases in the gas phase that are related to decomposition (such as oxygen or carbon dioxide). : Effective diffusion coefficient in the gas phase (affected by water content and porosity). : Gas source and sink terms (determined by respiration, diffusion boundaries, etc.).
[0029] In deep soil layers, a kinetic constraint is constructed for the binding of organic carbon with minerals. Langmuir adsorption kinetics is used to characterize the adsorption and desorption behavior of organic matter on the surface of deep minerals, thus reflecting the long-term sequestration mechanism of deep soil organic carbon. Langmuir kinetics is used to describe the binding and release of dissolved organic carbon from the surface of deep minerals. Langmuir kinetic formula: , The amount of organic carbon adsorbed on the surface of a unit mass of soil. Dissolved organic carbon concentration in pore water Maximum adsorption capacity (determined by mineral type and specific surface area). Adsorption rate constant : Desorption rate constant. Conservation formula for the deep dissolved phase (ignoring convection and diffusion-reaction): , Effective diffusion coefficient of the dissolved phase Soil volumetric mass : Deep biotransformation consumption items (can be modulated by temperature, moisture and redox state, and subsequently linked with the electrochemical layer).
[0030] Information sharing is achieved between the surface and deep neural networks through cross-layer residual connections. This involves feeding back surface observation data and computational features to the deep prediction units, preventing the separation of prediction results between layers and ensuring overall consistency from the surface to the deep layers. The intermediate representations and key physical quantities of the surface network are mapped back to the deep network, avoiding layer-by-layer fragmentation. Residual transfer formula: Consistency penalty formula: , : Intermediate features of surface networks Dimensionality reduction or projection mapping to extract transferable information. Deep subnetworks , Predicting target quantities (such as flux or concentration) in the surface and deep layers. : Convert surface predictions into a mapping of deep comparable quantities (including scale and physical difference corrections). Inter-layer consistency loss.
[0031] A modeling method based on Michaelis-Menten kinetics is used to model the role of microbial secreted enzymes in the decomposition of organic matter. This modeling allows for the characterization of the nonlinear relationship between enzyme concentration and carbon decomposition rate. Michaelis-Menten kinetics describes the nonlinear relationship between enzyme-substrate reaction rate and substrate concentration, and this equation is used as a soft constraint in the network. Michaelis-Menten kinetic rate formula: , : Substrate decomposition rate per unit time : Concentration of decomposable substrates (which can be determined by...) derived), Michaelis constant (reflects the substrate concentration required to reach half-maximum rate). Maximum reaction rate, which is affected by enzyme concentration. The formula relating enzyme concentration to enzyme concentration is: , Catalytic constant (the number of substrates converted per unit time by a single enzyme molecule). The effective concentration or activity level of the enzyme.
[0032] Building upon this, a time lag constraint is introduced to describe the response delay between enzyme synthesis and environmental conditions (temperature, humidity, nutrient supply). This time lag design is intended to reflect that enzyme activity in real ecosystems does not respond instantly to environmental changes, but rather exhibits a lag effect, thereby improving the accuracy and stability of carbon decomposition rate predictions.
[0033] During training, the neural network simultaneously optimizes the Mie dynamics nonlinear mapping and time-delay characteristics, enabling the model to accurately capture the dynamic response relationship between environmental stimuli and carbon decomposition over time. Enzyme synthesis exhibits a lag in its response to temperature, humidity, and nutrient supply, characterized by a delay differential equation. The time-delay kinetic formula for enzyme concentration is: , , Synthetic gain coefficient : Environmental stimulus function, which includes temperature Moisture content Nutritional supply Mapped to synthetic intensity, The response time lag to temperature, humidity, and nutrients. : Enzyme inactivation or degradation rate constant, loss corresponding to enzyme constraint (incorporated into PINN formula): , The network predicts the decomposition rate; other notations are the same as above. During training... It can be used as a learnable parameter or a semi-empirical prior parameter.
[0034] In soil carbon cycling, redox potential (RPP) significantly influences the stability of organic carbon. Using the principle of RRP correlation, we modeled the carbon transformation trends under different soil environments (e.g., aerobic and anaerobic conditions) and embedded this model into the electrochemical feature layer of a neural network. The intensity of decomposition or transformation under aerobic and anaerobic conditions is modulated as an electrochemical characteristic layer constraint. Redox modulation function formula: , Between and The modulation factor, aerobic conditions are close to Anaerobic conditions are close to , Slope parameter controls transition sensitivity. Redox inflection point potential. Formula for the electrochemically modulated biotransformation term: , Net biological consumption rate of dissolved organic carbon. Consumption coefficient under aerobic conditions Consumption coefficient under anaerobic conditions Dissolved organic carbon concentration.
[0035] Simultaneously, considering the diffusion behavior of dissolved organic carbon in porous media, the diffusion mechanism model is embedded into the diffusion feature layer to capture the migration and distribution patterns of organic carbon in soil pores. The diffusion mechanism is explicitly embedded in the network, ensuring that predictions follow the control of diffusion by Fick diffusion and pore geometry. The dissolved phase diffusion equation is as follows: The symbols have the same meaning as above. The formula for the porosity function of the effective diffusion coefficient is as follows: , Molecular diffusion coefficient in free water Volumetric moisture content Total porosity is used to explicitly incorporate the "impedance effect" of pore geometry and water content on diffusion into the network constraint. The parameters can be localized.
[0036] In the neural network structure, electrochemical feature layers and diffusion feature layers are jointly embedded to form an electrochemical-diffusion synergistic constraint layer. This ensures that the model, during training and prediction, adheres to both the influence of redox reactions on carbon stability and the constraints of diffusion behavior on carbon migration. and As an explicit feature input, residuals are also incorporated into the loss to ensure that the prediction does not overestimate decomposition in the anaerobic zone and underestimate transformation in the aerobic zone. The formula for the residual loss of the synergistic layer is: The symbols have the same meaning as above. This design can effectively correct the bias of traditional neural networks in predicting carbon conservation in anaerobic zones, improve the model's adaptability to complex environments and its prediction accuracy, and the total loss function formula is as follows: , Errors compared to observed data (monitored terms for concentration, flux, water content, and Eh). The residuals of the surface equation (A1)–(A3) : Residuals of deep equations; Enzyme kinetic constraints Electrochemical-diffusion residuals : Table-deep consistency Boundary conditions and initial conditions penalties Weight coefficients for each part (adaptively adjustable during training), initial conditions: , , , Boundary conditions: Neumann boundary at the surface (flux known): , For external direction, For surface flux. Zero flux boundary at depth: Boundary and initial values are entered uniformly. Data points and physical points are sampled from the same grid in the time domain. Normalization, parallel embedding of surface subnetworks, deep subnetworks, and electrochemical-diffusion synergistic layers, using sinusoidal or Fourier features to improve spatial-temporal resolution, for and Perform automatic differentiation to obtain the derivatives and divergences of (A2)(A3)(A5)(B3)(C3). First minimize... Then gradually add Using a normalization strategy for unbalanced residuals to make The algorithm is dynamically adjusted during training, and the mass conservation error, energy consistency index, and interlayer consistency index are periodically calculated to prevent overfitting of data and violation of the mechanism.
[0037] The hierarchical physical constraint neural network module achieves comprehensive modeling of the soil carbon cycle from the surface to the deep layers by combining multi-physics hybrid PINN, enzyme kinetic time-delay constraints, and an electrochemical-diffusion synergistic layer. This module not only embeds key physical and biological processes such as temperature effects, water and air transport, mineral adsorption, enzyme reaction delays, electrochemical regulation, and diffusion migration into the neural network, but also ensures the model's consistency in the spatial dimension and its rationality in the temporal dimension through cross-layer residual connections and synergistic constraint mechanisms, thereby enhancing the interpretability and reliability of the prediction results.
[0038] The multi-scale 3D carbon storage mapping module consists of three parts: topology-enhanced fractal kriging interpolation, holographic soil map neural network, and uncertainty decomposition Monte Carlo. It predicts soil carbon storage on a centimeter-level 3D grid (voxels) and quantifies uncertainty. It integrates two types of uncertainty—pore-root topology, microscopic three-layer media structure, and data / mechanism—into the modeling, ensuring high resolution, continuity, and interpretability. Inputs: Unified features from the fusion module and physical constraint module (temperature, water content, Eh, mineral parameters, BSI, biological / mineral factors, hydrological factors), measured carbon samples (volume or mass carbon content), CT / MRI volumetric data (particles / pores / aqueous phase), root reconstruction point cloud or skeleton, and pore connectivity index. Outputs: 3D carbon storage voxel map (centimeter level), confidence interval upper and lower limits, sensitivity heatmap (intensity of influence on input factors), quality labels, and uncertainty source decomposition (data vs. mechanism).
[0039] Topology-enhanced fractal kriging interpolation receives layered, consistent carbon pool predictions and a consistent feature matrix. Traditional kriging interpolation has limitations when handling data with complex topological structures and fractal properties. By introducing topological information (hydrological topology constructed from dynamic hydrological topology) and fractal dimension as covariates in the kriging model, it can more accurately capture the spatial variability and anisotropy of soil carbon storage, achieving interpolation down to the centimeter-level voxel level from sparse points or regions. Holographic soil graph neural networks are an innovative neural network architecture for processing more microscale soil structure information. It constructs three-dimensional multilayer graphs of granular, pore, and capillary water layers based on CT or MRI scan data. These layers are represented as independent but interconnected graphs in the graph neural network, and the diffusion and consolidation processes of dissolved organic carbon in different media (granular surfaces, macropores, capillary water films) are simulated through cross-layer information propagation mechanisms (such as message passing networks). This allows the model to understand the impact of microstructure on carbon storage. Uncertainty decomposition Monte Carlo methods are used to quantify and decompose uncertainties in carbon storage assessment. This method employs multi-sample simulations of model parameters, input data, and random variables in the physical equations, combined with variance decomposition techniques (such as the Sobol index) to decompose the total uncertainty into contributions from different data sources, model modules, or parameters. This enables the system to output 3D carbon storage voxel maps on a centimeter-scale grid, along with confidence intervals (indicating the reliable range of the prediction results), sensitivity heatmaps (revealing which input factors have the greatest impact on the prediction results), and quality labels (assessing the reliability level of local predictions). Multi-sample simulation: Key inputs to the system are randomly sampled to generate a large number of input scenarios. Each scenario is fully fed into the entire evaluation system (data fusion, PINN, mapping), resulting in a series of 3D carbon storage voxel maps. Variance decomposition: The total variance of the final carbon storage prediction results is analyzed using nonparametric regression or methods based on the Sobol index. The Sobol index quantifies the contribution of each input factor (e.g., satellite image quality, microbial gene abundance, soil moisture content, specific parameters of PINN) to the total prediction uncertainty. Output: Based on these simulation results, the system generates confidence intervals (e.g., % confidence range) for each voxel, a sensitivity heatmap (visually showing which region's carbon storage is most sensitive to a specific input factor), and quality labels (e.g., labeling each voxel as "high confidence," "medium confidence," or "low confidence" based on data coverage and model uncertainty). This provides users with comprehensive information about the reliability of the assessment results. The holographic soil mapping neural network is specifically designed to process microscale soil structure data, taking CT or MRI scans of soil samples as input. These scans are processed using image segmentation techniques to generate the three-dimensional geometry of the soil's granular layer (composed of mineral and organic matter particles), pore layer (filled with air or water), and capillary water layer (a water film adhering to the particle surface).Graph Construction: The three-layer structure is abstracted into a multi-layer graph. Within each layer, adjacent voxels or regions are considered nodes, and their physical connections form edges. Graph connections are also established between different layers. Information Propagation: This graph neural network uses GNNs for information propagation. Within each layer, node features (such as local carbon density, porosity, and particle composition) are aggregated and updated among their neighbors. More importantly, through the cross-layer information propagation mechanism, the model simulates the complex process of dissolved organic carbon (DOC) diffusing from the capillary layer to the macroporous layer and then adsorbing into the granular layer. For example, granular layer nodes receive information about DOC concentration and flow rate from capillary layer nodes and update their adsorbed carbon characteristics. Output: Finally, the holographic soil graph neural network outputs the carbon distribution and sequestration state at the microscale. This information is integrated into a three-dimensional carbon storage voxel map, supplementing the details of macroscopic interpolation. Uncertainty Decomposition Monte Carlo Method: This method not only quantifies uncertainty but also focuses on decomposing the sources of uncertainty. Multi-Sample Simulation: Key inputs to the system are randomly sampled to generate a large number of input scenarios. Each scenario is fully fed into the entire assessment system (data fusion, PINN, mapping), resulting in a series of 3D carbon storage voxel maps. Variance decomposition: Using nonparametric regression or a Sobol index-based method, the total variance of the final carbon storage prediction is analyzed. The Sobol index quantifies the contribution of each input factor (e.g., satellite image quality, microbial gene abundance, soil moisture content, specific PINN parameters) to the total prediction uncertainty. Output: Based on these simulation results, the system generates confidence intervals for each voxel, sensitivity heatmaps (visually showing which region's carbon storage is most sensitive to specific input factors), and quality labels (marking each voxel as high, medium, or low confidence based on data coverage and model uncertainty). This provides users with comprehensive information about the reliability of the assessment results.
[0040] Topology-enhanced fractal kriging interpolation improves upon traditional fractal kriging by incorporating soil pore connectivity and root topology information into the fractal variogram, enabling the interpolation process to more accurately reflect the spatial characteristics of complex soil structures. Pore connectivity characterizes the continuity of gas and liquid channels in the soil, while root topology represents the influence of plant roots on carbon input and migration; together, they constitute a topology enhancement for spatial interpolation. In practical applications, this method effectively fills data gaps in highly heterogeneous regions, such as gravel layer distribution areas or regions with complex rhizosphere structures, thereby improving the continuity and reliability of three-dimensional carbon storage spatial prediction. The study area is discretized into a three-dimensional voxel grid, defining anisotropic scales (different correlation lengths in the horizontal and vertical directions). The anisotropic Mahalanobis distance formula is as follows: , The difference in three-dimensional coordinates between two points. Anisotropic scaling matrix (the principal axes are the squares of the horizontal and vertical correlation lengths). Measure the equivalent distance after anisotropy. The power-law scale structure is characterized using the fractal variogram, and spatial correlation is enhanced using "pore connectivity" and "root topology". Fractal variogram (Matern / power-law example) formula: , Fractal variogram Step-up / Nuclear Effect Structural variance : Relevant length scale, Fractal / roughness index controls the intensity of small-scale variations. Anisotropic distance calculated by (K1). Introducing a topology enhancement term improves reachability through increased porosity connectivity, and the root network provides directional input-migration paths. Topology enhancement weight function formula: , Topology enhancement weights between two points (0–1). The Sigmoid function stabilizes the weights between 0 and 1. : Pore connectivity similarity (e.g., connectivity index based on the maximum connected component or seepage path). Root topological similarity (e.g., the normalized length of the shortest path between two points on the root system skeleton or the degree of co-existence of the same root sequence). Hyperparameters balance the influence of two types of topology. Formula for the topology-enhanced fractal variogram: , : The mutation function after topological enhancement Topology reinforcement strength coefficient The remaining symbols are shown above; the more "connected" the topology, the better. The smaller the value, the larger the corresponding covariance, resulting in smoother and more consistent interpolation. An enhanced variogram is used to construct the covariance matrix and perform three-dimensional kriging estimation. Covariance formula: Kriging system formula: , Sample points covariance, Sample covariance matrix Sample point—the covariance vector of the location to be estimated. Kriging weights Lagrange multipliers : A vector consisting entirely of 1s. Estimation formula: Kriging variance formula: , Carbon storage estimation at the three-dimensional voxels to be estimated. Carbon values observed at sample points Ordinary Kriging variance. A robust loss is added during variogram fitting to suppress the influence of outliers; local windows are used to re-estimate variance in regions of abrupt connectivity changes. To improve local adaptability; for extremely sparse zones, auxiliary covariate cokriging expansion is used.
[0041] A holographic soil mapping neural network constructs multi-layered three-dimensional soil maps based on CT or MRI scan data to simulate the effects of different soil structural units on carbon diffusion and sequestration. This multi-layered three-dimensional map includes: a particle layer (describing the morphology and arrangement of soil mineral particles); a pore layer (characterizing the channel structure of air and water in the soil); and a capillary water layer (depicting the retention state of liquid water in soil capillary pores). The neural network models the interactions between the particle, pore, and capillary water layers through a cross-layer information propagation mechanism, thereby capturing the diffusion, migration, and sequestration patterns of carbon in different soil media. This design can precisely reflect the influence of microstructure on carbon behavior in three-dimensional space, achieving holographic and three-dimensional carbon storage mapping by segmenting three types of voxels from CT / MRI volumetric data and constructing a multi-layered map. The definition of the three-layered map and the edge set formula are as follows: , : Set of granular layer nodes (mineral particle voxels). : Pore layer node set (air / channel voxels) : Capillary water layer node set (water-filled or film water element). Same-layer connection, Cross-layer connections (interfaces, wetting contacts, etc.). Assigning measurable characteristics to nodes / edges in each layer: Granular layer: mineral type, specific surface area, surface charge, Pore layer: local pore size, shape factor, curvature, channel parameters; capillary water layer: local water content, film thickness, contact angle, capillary pressure; edge features: interfacial tension, number of effective diffusion channels, interfacial area. A graph neural network is used on the three-layer graph for cross-layer message passing to characterize the coupled process of dissolved organic carbon diffusion in pores, retention in the aqueous phase, and adsorption and consolidation on particle surfaces. Same-layer and cross-layer message aggregation formulas: , :node The Layer representation, : Set of neighbors on the same floor Cross-level neighbor set Attention weights (determined by features and edge attributes). Learnable mappings Nonlinear activation.
[0042] Discrete diffusion constraints are embedded in the pore-water layer; Langmuir consolidation constraints are embedded in the particle-water interface. The formula for pore / water phase discrete diffusion constraints (Fick's law in the figure) is as follows: , :node The concentration of dissolved organic carbon at (pores or water layer nodes) :side Effective diffusion rate (related to pore size, water content, and connectivity). Local source and sink terms (e.g., adsorption, consumption by microorganisms). Particle / water interface consolidation (Langmuir diagram): , Particle nodes The amount of adsorption on the surface, Node-dependent adsorption / desorption rate constants : Local maximum adsorption capacity, others are as above. Physical consistency loss (included in network training) formula: , The time derivative is approximated by automatic differentiation from the network; other parameters are the same as above. The predictions (concentration, solids content) on the three-layer map are integrated within voxels to obtain voxel-level carbon storage. When the voxels are coarser than the map resolution, in-voxel averaging and scale uniformity are performed; when the voxels are finer than the map resolution, guided super-resolution (using pore / water film geometry as prior) is used to refine to the centimeter level. Voxel-level carbon storage aggregation formula: , The carbon storage estimate of this voxel The set of nodes in the pores, water layer, and particle layer that fall into this voxel. : The volume of the infinitesimal element represented by the corresponding node : Aggregation normalization constant (normalized by voxel volume or weight).
[0043] Uncertainty decomposition Monte Carlo (UDD) performs separate dropout processing on physical constraints and observational data during forward prediction to quantify uncertainties from different sources. Physical dropout characterizes the uncertainties introduced by mechanistic equations during modeling, including errors in physical mechanisms such as temperature effect models and adsorption kinetic models. Data dropout quantifies noise and defects in observational data, such as remote sensing image errors and sensor measurement biases. The final output includes confidence intervals for the predicted values and a sensitivity heatmap. Confidence intervals provide the range of prediction reliability, while the sensitivity heatmap visually demonstrates the influence of different input factors on the prediction results, thus providing a basis for subsequent decision-making. Two parallel randomization methods are used in forward prediction: Physical dropout: randomly perturbs the parameters or constraint strengths of the mechanistic layer (temperature effect, Langmuir, diffusion / electrochemical constraints) to express incomplete understanding of the mechanism. Data dropout: randomly masks / perturbs the input observational features and sample observation values to express sensing errors and observational gaps. The formulas for the two types of randomization are as follows: , : No. The set of physical parameters / weights for the next sampling (e.g.) (disturbance) : The prior distribution or empirical perturbation distribution of physical parameters Original input features Data dropout / noise operators (random masking, noise addition, missing data simulation). Using outer-loop sampling physical parameters and inner-loop sampling data perturbations, the total variance is estimated and its sources are separated. Two-level Monte Carlo variance decomposition formula: , Predicting carbon reserves : The variance of the sampling of physical parameters Given physical parameters, the expected value of the data perturbation. Given physical parameters, the variance of the data perturbation. : Expected value of the physical parameter sampling. For each voxel statistical distribution, output confidence intervals and sensitivity heatmaps. Confidence interval formula: Pixel variance formula: , Voxel location Confidence level is The interval, Quantile estimation The total variance estimate for this voxel. Factor importance is given using gradient or Sobol approaches and rendered as a heatmap. Gradient sensitivity and normalization formula: , Input factor In voxels The relative sensitivity of the location The gradient is obtained by automatic differentiation. The denominator is the sum of the absolute values of the gradients of all factors used for normalization. Note: The Sobol first-order / total effect exponent can also be calculated in parallel for cross-validation.
[0044] The multi-scale 3D carbon storage mapping module achieves high-precision, holographic, and interpretable 3D mapping of soil carbon storage through the synergistic effect of topology-enhanced fractal kriging interpolation, holographic soil map neural network, and uncertainty decomposition Monte Carlo. Data preparation involves aligning measured carbon sampling points, fused environmental / physical factors, CT / MRI three-layer maps, root skeleton, and pore connectivity indices to the same voxel grid. High heterogeneity areas are marked in the gravel layer and rhizosphere as trigger zones for topology enhancement and local parameter reestimation. Topology-enhanced kriging: fitting the parameters of the basic fractal variogram. .calculate and set ,get Solving (K5) and (K6) yields the initial carbon field of centimeter-scale voxels and Holographic soil map neural network: obtained from CT / MRI three-segmentation. Construct node / edge features. Use physical consistency residuals as regularization and train jointly with sample point / cell experimental supervision; aggregate graph-level results to voxels to obtain refined features. Uncertainty decomposition Monte Carlo: Outer loop sampling Inner ring Do Cumulative (U2) estimation of total variance and source decomposition; generation of (U3) confidence intervals and (U4) sensitivity heatmaps. Consistency check and release: verification of mass conservation and physical residual thresholds on surface / deep stratified voxels; backtracking of sources (physical or data) for high-variance voxels and marking priority retesting areas; output of 3D carbon storage map, confidence intervals, sensitivity map and mass labels; topological index calibration: pore connectivity derived from CT / MRI connectivity analysis or mercury pressure / nitrogen adsorption back-calculation; root topology derived from root point cloud skeletonization, calculation of shortest path of stratified root sequence, normalization to Fractal parameter fitting: fitting using the sample semivariance function Local revaluation in highly heterogeneous sub-blocks. Three-layer graph segmentation: training / calibrating the voxel classifier to ensure particle / pore / aqueous phase separation quality; interface area is statistically derived from voxel contact surfaces. Diffusion / sustainability parameters: diffusivity. The formula is given by empirical equations for pore size and water content and locally calibrated. Priors derived from batch adsorption experiments or literature. Monte Carlo priors: Based on the above calibration confidence range settings; The model combines sensor accuracy and remote sensing error. Computation and storage: Block-based parallelism and sparse graph storage are used for centimeter-level 3D voxels; iterative solutions and approximate covariance (such as NNGP / local neighborhood) are used to accelerate the Kriging matrix. Validation: External validation is performed using profiles and sample points; a comprehensive evaluation is conducted based on in-profile mass conservation, physical residuals, and observation errors.
[0045] The temporal evolution and digital twin module consists of three parts: a self-evolving carbon cycle cellular automaton, a physics-Transformer dual-channel state space, and a multi-index stability early warning system. On a unified grid and timeline, it combines self-organizing evolution simulation, physical mechanisms, and temporal deep learning to continuously assimilate historical observations, predict future evolution, and provide stability early warnings, forming an interpretable and predictable digital twin. The self-evolving carbon cycle cellular automaton discretizes the soil space into a series of cells, with microorganisms, minerals, and organic carbon as agents. These agents interact within the cells according to pre-defined rules, taking into account the influence of environmental factors. These interaction rules are trained through reinforcement learning, enabling the cellular automaton to learn dynamic evolution patterns of carbon sequestration, mineralization, and migration from historical data, thereby simulating long-term carbon cycle processes. The physics-Transformer dual-channel state space is a hybrid model combining data-driven and physical mechanisms. It consists of a data-driven channel and a physical equation channel. The data-driven channel employs a Transformer architecture to encode multi-source time-series observation data (such as long-term environmental monitoring data and land use change data), capturing complex nonlinear time dependencies. The physical equation channel constrains the predicted trajectory with a state transition matrix (derived from physical laws and mechanistic models), ensuring that simulation results conform to fundamental physical principles such as energy and mass conservation. This module supports counterfactual branching, allowing users to input different management scenarios to simulate the impact of these measures on carbon storage, providing a basis for ecosystem management decisions. The multi-indicator stability early warning system, based on historical assimilation and future evolution projections, combined with preset thresholds and risk assessment models, performs multi-indicator assessments and early warnings of soil carbon storage stability. Indicators may include carbon storage change rate, carbon pool stability index, and carbon saturation. When these indicators reach warning thresholds, the system triggers an alert, indicating potential carbon loss risks. Ultimately, this achieves historical assimilation, future evolution projection, and stability risk assessment of soil carbon storage.
[0046] The self-evolving carbon cycle cellular automaton discretizes the soil system into grid points, each grid point serving as a basic unit occupied by multiple agents. These agents include three categories: individual microorganisms, mineral particles, and organic carbon pools. Each type of agent possesses specific behavioral rules and state parameters. Individual microorganisms simulate decomposition activities, involving enzyme secretion and substrate utilization; mineral particles represent soil structure and their adsorption and protection of organic carbon; and the organic carbon pool reflects the dynamic balance between litter input, mineralization, and stable carbon sequestration. Reinforcement learning methods are used to train the interaction rules between agents, including dynamic interactions in three directions: carbon sequestration, mineralization, and migration, thereby forming a carbon cycle evolution model consistent with ecological mechanisms. During operation, the system adaptively learns the interaction patterns under different environmental conditions, achieving self-organized evolutionary simulation of the carbon cycle, and ultimately outputting time-series carbon storage change results, discretizing the region into two-dimensional or three-dimensional grid points; each grid is a "cell" containing three types of agent states: microorganisms... mineral particles Organic carbon pool (Decomposable libraries and stable libraries). Cell state vector formula: , Cell At any moment state, : A library of decomposable organic carbons Stable carbon pool Effective microbial biomass or enzyme activity proxy variable, Mineral protection capacity index (related to maximum adsorption capacity). Volumetric moisture content :temperature, Redox potential. The carbon library of each cell is updated using a verifiable local mechanism, allowing exchange (migration) between adjacent cells. Local mass conservation and migration formula: , , : Fallen matter input, Mineralization decomposition loss (released as gaseous / dissolved phase). : Stabilization (adsorption / embedding) in a stable library. Cell The neighborhood, Migration flux between adjacent cells (varying with water content and pore connectivity). Destabilization loss of the stablebase (release under extreme disturbances). Decomposition formula: The formula for the local rate of consolidation: , Temperature response parameters and constants Moisture correction function Redox modulation function Langmuir adsorption / desorption and maximum capacity parameters. Neighborhood migration flux (discrete diffusion) formula: , , : Effective diffusion coefficient of a cell pair : Diffusion constant in free water Regarding the water content and porosity at the boundary, Connectivity factor (derived from pore network or root channels). The three types of processes (sequestration, mineralization, and migration) are treated as adjustable strategies; under the premise of satisfying physical constraints, the rules are optimized through reinforcement learning to achieve optimal long-term carbon storage and steady-state performance. Reward function formula: , Stable library increment, Emissions due to mineralization Penalties for violations of mass conservation and boundary conditions. : Variance of total carbon Weighting coefficients. Strategy update (actor-critic system) formula: , : Strategy in state downsampling action The probability, Advantage function (the increment of returns relative to the baseline); Regularity coefficient : KL constraint between the policy and the old policy, used for stable training.
[0047] The system employs a dual-channel state space design, combining deep learning with physical mechanisms to enhance prediction accuracy and interpretability. The first channel is the data-driven channel, using a Transformer structure to encode remote sensing time-series data and in-situ sensor observation data, capturing long-term temporal dependencies and cross-scale features from multi-source observations. The second channel is the physical constraint channel, constructing a state transition matrix using physical equations related to the carbon cycle, providing physical consistency constraints for the prediction process. Under this dual-channel interaction mechanism, the predicted trajectory output from the data channel is corrected by the physical channel, thus avoiding overfitting and physical biases that may occur with purely data-driven methods. The system design also includes a counterfactual branch to simulate carbon storage changes under different management measures. For example, three management schemes—vegetation restoration, fertilization management, and drainage control—can be input to deduce the corresponding temporal evolution trends of carbon storage. The remote sensing and in-situ time-series data are concatenated and input into the Transformer to extract cross-scale dependencies, outputting a data-driven state prediction. Multi-head attention (simplified) formula: , : Query, key-value matrix, The bond vector dimension is used to construct a state-space model obtained by discretizing the carbon cycle equation, providing physically consistent prior evolution. Mechanistic state-space formula: , : State vector (layered carbon library, dissolved phase, enzyme activity, water content proxy, etc.). External drivers and management measures (vegetation restoration, fertilization intensity, drainage valve position). : The state transition matrix obtained by discretizing the local mechanism (including Arrhenius, moisture, and adsorption module mappings). Input-state coupling matrix, : Observation mapping matrix (maps the state to observables). Process noise and observation noise. The predictions from the data channel are corrected by the physical channel to achieve a fused state that is both data-driven and physically-compliant. Correction fusion (Kalman correction) formula: , , Data channel prediction, Prior knowledge obtained through physical channel advancement Fusion gain, Two-channel uncertainty covariance estimation The fused state is trained using joint loss to balance data fitting, physical residuals, and counterfactual consistency. Joint loss formula: , : Observational predictions obtained from fused state mapping Weights for each item. "Intervene" in the input strategy to simulate future trajectories under different management plans. Counterfactual intervention formula: , Counterfactual situation, : Input sequence given strategy (vegetation restoration intensity, fertilization amount, drainage settings).
[0048] Multi-indicator stability early warning is used to monitor the stability and provide risk warnings for the carbon storage evolution trend of an ecosystem. The Lyapunov index is used to detect the stability characteristics of the system during dynamic evolution; a positive index indicates that the system may be trending towards instability. Shannon entropy is used to assess the disorder of carbon flow processes; increased entropy indicates more unpredictable carbon migration and transformation processes, suggesting increased system complexity. A robustness index for complex networks is introduced to quantify the ecosystem's ability to withstand external disturbances. This includes the ability to maintain connectivity under network node removal or external shocks and the sensitivity to critical node failure. Finally, a threshold discrimination mechanism is established through the comprehensive judgment of these three indicators. When the system exhibits high instability, high disorder, or low robustness, an early warning mechanism is triggered, outputting risk warning information to provide decision support for proactive carbon storage management. The local Lyapunov index is estimated on the fusion state trajectory to determine whether the disturbance is attenuating or amplifying. The formula for maximum Lyapunov index estimation (time window) is as follows: , : The infinitesimal perturbation vector applied to the state. Sampling interval The number of steps within the window. It indicates an unstable trend. It normalizes various carbon fluxes per unit time into a probability distribution, measuring the degree of disorder. Entropy index formula: , : No. Non-negative values of fluxes (such as mineralization, consolidation, leaching, and lateral migration). Flux share Number of flux categories Larger scales indicate that the process is more unpredictable and the system is more "disorderly".
[0049] Construct a carbon transport network (nodes are grid points or carbon pool units, edge weights are carbon exchange intensity), and evaluate connectivity and functional maintenance under random or targeted node failures. Robustness integral index formula: , Number of network nodes Number of nodes removed (from 0 to 1) (process) Remove The maximum size of the connected subgraph after a certain number of nodes. The larger the value, the more robust the network. Algebraic connectivity can also be monitored in parallel. Changes are made to supplement the judgment criteria. The three indicators are normalized and then weighted and summed; exceeding the threshold triggers an early warning and is categorized accordingly. Risk scoring and warning formula: , : Normalization function (such as min-max or logical function). Weights of the three indicators (as determined by historical examples). An alert will be triggered if the threshold for the classification is exceeded.
[0050] The temporal evolution and digital twin module achieves full-chain modeling from historical data-driven to future dynamic extrapolation through self-organizing simulation of self-evolving carbon cycle cellular automata, mechanistic constraint prediction of physical-transformer dual-channel state space, and comprehensive evaluation of multi-indicator stability early warning. The temporal evolution and digital twin module can provide refined dynamic evolution simulation of carbon storage in complex environments and provide early warning before the stability of the ecosystem declines, thus providing scientific basis and forward-looking support for carbon sink management and ecological restoration.
[0051] The adaptive transfer learning module achieves rapid modeling and high-precision prediction of soil carbon storage under different environments through the synergistic effect of three parts: cross-mechanism adaptive meta-learning, physically consistent federated distillation, and dynamic multi-scale constraint adjustment. In cross-regional (wetland, farmland, forest) and cross-soil type (sand, loam, clay) scenarios, it utilizes mechanistic decoupling + meta-learning + federated distillation + dynamic constraint weights to achieve a balance between rapid small-sample adaptation, privacy protection, physical consistency, and high-precision prediction. Input: Unified features output from the fusion module and the physical constraint module (climate driving factors include temperature time series and precipitation / water content; soil physical factors include porosity and particle size distribution; hydrological and electrochemical factors; existing regional model parameters and samples). Output: Rapidly converged model parameters for the new region, physically consistent prediction results, spatiotemporally adaptive constraint weight curves, and uncertainty labels.
[0052] Cross-mechanism adaptive meta-learning addresses the potential differences in carbon cycle mechanisms across different regions or soil types (e.g., differences in microbial communities and carbon decomposition rates between cold and tropical soils) using a meta-learning framework. First, a meta-model is trained on multiple source domains (different regions and soil types), enabling it to quickly learn a "learning strategy" that adapts to the model parameters. When encountering a new target domain, the meta-model can be rapidly fine-tuned using a small amount of target domain data, achieving adaptation to the carbon cycle mechanism of the new region, significantly reducing the amount of training data and time required for deployment in new regions. Physically consistent federated distillation addresses privacy concerns associated with cross-regional data sharing by employing a federated learning paradigm. Sub-models are trained locally in each region, and knowledge (rather than raw data) is aggregated into a global model using model distillation techniques. Crucially, carbon and energy conservation loss functions are introduced during cross-regional model aggregation to ensure that the aggregated global model still conforms to fundamental physical laws. This protects the privacy of monitoring data while guaranteeing physical consistency and robustness after aggregating models from different regions. The dynamic multi-scale constraint adjustment mechanism dynamically adjusts the weights and scales of physical constraints in PINN based on the characteristics of the target region. This dynamic adjustment mechanism enables the model to flexibly find the optimal balance between data-driven and mechanistic constraints according to the actual situation, further optimizing the model's accuracy and generalization ability. The meta-training phase of cross-mechanistic adaptive meta-learning involves the system first training on a dataset (source domain set) containing multiple different regions and soil types. In each training iteration, a task is randomly selected (i.e., a mini-batch of data is extracted from one source domain), a base learner is trained, and the loss is calculated on another mini-batch of data. The goal of the meta-learning algorithm is to optimize the model's initial parameters so that these initial parameters can quickly adapt and achieve good performance on any new task with a small gradient step size. The meta-test / adaptation phase involves fine-tuning the initial parameters learned by the meta-learner using only a small amount of data from that region when a new region or soil type (target domain) appears. Because the meta-model has learned a general "learning strategy," it can quickly adapt to the specific carbon cycle mechanism of the target domain even with a small amount of data, achieving high-accuracy predictions and solving the problem of traditional models requiring large amounts of data for retraining in new regions. Local Training for Physically Consistent Federated Distillation: Each region (client) independently trains a hierarchical physically constrained neural network model on its local dataset. During model training, in addition to the regular supervised learning loss, carbon conservation and energy conservation loss functions are introduced. For example, the carbon conservation loss ensures that the carbon flux and carbon pool changes predicted by the model are conserved over time. Knowledge Distillation: After the client model is trained, instead of uploading the raw data, the "knowledge" learned by the model (such as model parameters, feature vectors of intermediate layers, or soft labels predicted by higher-level teacher models for unlabeled data) is compressed and encoded.Global Aggregation: The central server collects distilled model knowledge from all clients. During aggregation, the central server introduces a global carbon and energy conservation loss function as a regularization term to guide the aggregation process, ensuring that the global model, while incorporating knowledge from different regions, still strictly adheres to fundamental physical laws. This ensures better physical consistency and accuracy of the model than simple parameter averaging federated learning. Dynamic Multi-Scale Constraint Adjustment Mechanism: Based on data density: In a certain region or at a certain depth, if available measured data is scarce, the weight of the corresponding physical equation constraint terms is increased, forcing the model to rely more on physical laws for prediction. Conversely, when data is abundant, the weight of physical constraints can be appropriately reduced, allowing the data-driven part to play a greater role. Based on Soil Heterogeneity Adjustment: If the features output by the multi-source heterogeneous data intelligent fusion module show that the soil type, mineral composition, or hydrological conditions of a certain region are highly heterogeneous, the system dynamically adjusts the physical constraint weights for specific physical processes (such as Langmuir adsorption and enzyme kinetics) to better adapt to the characteristics of the local environment. Scale-based: At the microscale (e.g., in holographic soil map neural networks), the weights of the electrochemical-diffusion synergistic layer and the enzyme kinetics time-delay constraint layer are higher; at the macroscale, the weights of the hydrological transport equation are higher. This dynamic adjustment mechanism enables the system to intelligently balance the advantages of data-driven and physical-driven approaches according to the actual application scenario and data characteristics, achieving high accuracy and high generalization ability of the model under various complex conditions.
[0053] Cross-mechanism adaptive meta-learning establishes a more adaptive learning mechanism by decoupling climate-driving factors from soil physical factors. Climate-driving factors include moisture and heat conditions, while soil physical factors include porosity and particle distribution. These two types of factors are separated in the modeling process to reduce overfitting caused by coupling. A model-independent meta-learning method is employed, training general initial parameters in the source region task, enabling the model to quickly adapt to different physical constraints with a small number of samples when facing new regions. This method can achieve rapid adjustment in diverse environments, such as ecosystems like wetlands, farmland, and forests, flexibly adapting to differences in climate and soil properties, thereby shortening the model's convergence time in new environments and improving prediction efficiency. Separating climate-driving factors and soil physical factors at the representation layer reduces overfitting caused by coupling and provides a controllable transferable component for cross-domain transfer. A dual-path encoder is established, dividing the input into a climate path and a soil path, resulting in a separable latent representation. Factor decoupling representation formula: , Climate-driven inputs (temporal characteristics of temperature series, precipitation, or volumetric water content). Soil physical inputs (porosity, particle size distribution, mineral parameters). : The corresponding encoder, : Two-channel encoder parameters, : The implicit representation after separation The concatenated joint representation serves as the input for downstream prediction. (In the source region set) The above treats each region as a task and uses Model Independent Meta-Learning (MAML) to train general initial parameters. The formula for fast intra-task adaptation (inner loop) is: , Model parameters to be learned (which may include...) (with prediction head parameters). In the task Adaptive parameters obtained after one or more gradient steps. Inner loop learning rate :Task The training loss (including data and physical terms). Meta-objective (outer loop) formula: , In the task Evaluate the adapted performance on the validation set, and optimize the outer loop. This allows it to adapt to unseen new tasks in just one or two steps. Decoupling regularization formula: Physical consistency decoupling regularization: , Mutual information estimation, penalizing information leakage in two-way representations. Two-way linear mapping at the last level. Frobenius norm Differentiable mappings from representations to physically observable quantities (such as carbon flux or coulombic energy). : The observed or assimilated value of the corresponding physical quantity Weighting coefficients. Task loss integration (used in M2 / M3) ): , : Monitoring error relative to the sample targets (carbon storage, flux); other signs are the same as above. Arrival in a new region. At that time, only a small number of samples are used for one to three steps of inner loop updates. A usable model can be obtained immediately; if the number of samples gradually increases, a small-batch external loop fine-tuning can be performed to achieve stable improvement.
[0054] Physically consistent federated distillation addresses data privacy and consistency issues in cross-regional modeling. Each region trains its local model independently, avoiding direct sharing of raw observation data and protecting the sensitivity of ecological monitoring data. In the global aggregation phase, instead of simple parameter averaging, a loss function based on physical conservation is introduced, including carbon and energy conservation constraints. The carbon conservation constraint ensures that carbon input, output, and storage conform to the balance of the carbon cycle; the energy conservation constraint ensures the consistency of soil temperature, moisture, and energy consumption processes. Finally, through a federated distillation mechanism, the knowledge learned from local models is distilled and integrated at the global level, forming a unified model that satisfies both physical consistency and avoids privacy leaks. This improves the reliability and interpretability of cross-regional predictions. Train the model locally using its data and physical constraints. The soft target and intermediate characterizations are recorded for subsequent distillation. Local loss: , :area Supervision error, :area Physical consistency residuals : Weights of physical terms, carbon conservation (discrete-time) formula: , System carbon storage (including the sum of all storage facilities). Carbon input (litter, root exudate), Carbon output (respiratory mineralization, runoff / leaching migration). Net loss or net gain in inter-liquidity transfer due to adsorption / desorption and dissolution phase exchange (sign defined by direction). Energy conservation (simplified heat balance) formula: , Changes in soil internal energy Net radiation : Sensitive heat flux, Latent heat flux ( For latent heat of vaporization, (for evaporation) Soil heat flux. In the federated phase, the residuals of carbon and energy conservation are incorporated into the consistency constraints to ensure that the cross-regional model adheres to basic conservation relationships. Each region uploads the knowledge required for distillation (soft prediction distribution, feature statistics, physical residual distribution), but not the original data or sensitive labels. Distillation objective (matching global students to local teachers) formula: , :area The teacher model provides logits / continuous outputs for common anchor points or synthetic samples. : Output corresponding to the global student model : softmax or temperature-normalized function (for regression, Gaussian matching can be used instead). Distillation temperature Kullback–Leibler divergence. Formula for the physically consistent federated objective: , Global student parameters :area The carbon conservation residual (the squared error or absolute error calculated from the carbon conservation). :area The energy conservation residual (calculated by energy conservation). Conservation of weights. Privacy and robustness: Differential privacy noise can be superimposed onto the uploaded statistics, and a secure aggregation protocol is used to prevent reverse engineering of single-area information.
[0055] Dynamic multi-scale constraint adjustment flexibly adjusts the weights of physical constraints at different temporal and spatial scales, achieving a balance between accuracy and flexibility. In the temporal dimension, the initial physical constraint weights are low to ensure the model has sufficient degrees of freedom for rapid learning with limited samples. As the sample size increases, the physical constraint weights are gradually strengthened to improve prediction stability and consistency. In the spatial dimension, for small-scale laboratory data, the constraint weights are set relatively loosely to better capture microscopic differences. For large-scale data such as remote sensing observations, physical constraints are gradually strengthened to ensure the overall prediction results conform to regional scale regularities. This dynamic adjustment mechanism maintains model flexibility under small sample conditions while providing high-precision predictions in large-scale scenarios, thus achieving a balance across application scenarios. In the early stages with small samples, the model has more degrees of freedom; as the sample size and iterations increase, physical constraints are gradually strengthened, balancing convergence speed and physical consistency. The temporal weighting strategy formula is as follows: , Training steps Physical constraint weights at time Minimum and maximum weights The growth rate and the speed of hardening need to be controlled. At the small scale of laboratory testing, microscopic differences need to be captured (relatively loose constraints), while at the large scale of remote sensing, regional conservation and regularity need to be ensured (relatively tighter constraints). Spatial weighting strategy formula: , Spatial scale The physical constraint weights at (measured in meters or feature-related length) Scale sensitivity coefficient; Same as above. The larger (closer to the regional scale). The closer Spatiotemporal joint weight formula: Total loss formula: , Physical weights after spatiotemporal fusion : The mixing coefficient of time and space weights Other regular expressions.
[0056] The adaptive transfer learning module achieves rapid adaptation to new regions through cross-mechanism adaptive meta-learning, ensures privacy and physical consistency in cross-regional modeling through physically consistent federated distillation, and achieves a balance between flexibility and accuracy in both time and space through dynamic multi-scale constraint adjustment. The design of the adaptive transfer learning module enables rapid deployment and updating of models in various ecological environments such as wetlands, forests, and farmland, significantly reducing the cost of cross-domain modeling and improving the reliability and generalizability of soil carbon storage prediction.
[0057] The decision support and visualization module achieves visualization, interpretability, and traceability of carbon storage simulation results through three components: immersive twin interaction, a cross-modal interpretable framework, and blockchain-based trusted traceability visualization. The cross-modal interpretable framework in the decision support and visualization module extends Shapley decomposition to multimodal inputs, quantifying the marginal contributions of remote sensing bands, microbial genes, mineral types, and environmental factors to carbon storage prediction, and outputting factor ranking and weight graphs. The blockchain-based trusted traceability visualization in the decision support and visualization module uses hash locking to put the entire process of sampling, modeling, and prediction on the blockchain, ensuring the transparency and immutability of carbon sink data in scientific research, policy, and carbon trading scenarios. The immersive twin interactive interface maps the prediction results, such as 3D carbon storage voxel maps and evolution trends, to a VR / AR interactive environment in real time, forming a virtual digital twin of soil carbon. Users can use natural interaction methods such as gestures and voice to roam in the virtual environment, analyze soil structure, view carbon storage distribution at different depths, and simulate management measures, gaining an immersive perceptual experience. The cross-modal interpretable framework adopts the Shapley Additive Explanations (SHAP) framework and extends it to multimodal inputs. This framework quantifies the marginal contributions of remote sensing bands (such as NDVI and EVI), microbial gene types, mineral types, and environmental factors (such as temperature and humidity) to carbon storage prediction. The system outputs a ranking of factor contributions and a weighted graph, intuitively showing which factors have the greatest impact on carbon storage, as well as their direction and intensity of influence, enhancing the model's interpretability and helping researchers better understand the carbon cycle mechanism. Policymakers can then formulate more targeted management strategies based on this information. The blockchain-based trusted traceability and visualization uses hash locking to record the entire process of soil sampling (sampling time, location, and method), data preprocessing, model parameters, prediction process, and the final carbon storage assessment results on the blockchain. This means that every key step is encrypted and recorded on an immutable blockchain. Users can trace the chain to find the source and generation process of any assessment result, ensuring the transparency, immutability and credibility of carbon sink data in scientific research, policy making and carbon trading scenarios, effectively solving the trust crisis of carbon sink data.
[0058] A twin sandbox environment combining virtual reality (VR) and augmented reality (AR) is constructed, mapping the predicted spatial distribution and dynamic evolution of 3D carbon storage onto an interactive 3D scene. Users can actively adjust key environmental and management parameters within the twin sandbox, including water level, vegetation type, and fertilization measures. The system calculates the impact of these actions on soil carbon storage in real time. This interactive mechanism transforms complex carbon cycle dynamics into intuitive visual feedback, helping users directly observe the causal relationship between management measures and carbon storage changes, thus supporting scientific management and policy formulation. Simultaneously, the system supports scenario comparison across multiple scenarios. For example, under the same regional conditions, it can simulate three scenarios: water level increase, vegetation restoration, and organic fertilizer application. Users can compare the carbon storage change trajectories within the twin environment. The predicted spatial distribution and dynamic evolution of soil carbon storage are mapped onto a three-dimensional interactive environment combining VR and AR. Inputs: Carbon storage prediction results (centimeter-level 3D voxel mesh), management parameters (vegetation, water level, fertilization), and physical constraint factors (temperature, moisture content, redox potential). Implementation: A 3D rendering engine (such as Unity / Unreal) loads a carbon storage voxel field, transforming the dynamic evolution trajectory into a visual animation. 3D voxel mapping function formula: , The visualization values (color, transparency, size, and other rendering attributes) of a voxel in the twin sandbox. : The predicted carbon reserves at the corresponding location Environmental parameters (temperature, humidity, moisture content, etc.). Mapping functions are used to convert numerical values into visual attributes. Users can adjust the water level in a VR / AR environment. Vegetation type Fertilization intensity The system calculates and predicts changes in carbon reserves in real time. The formula for the impact of management intervention on carbon stocks is: , In time Changes in carbon reserves. Water level height : Vegetation type code (e.g., forest = 1, farmland = 2, wetland = 3). Fertilization measures (input rate kg / ha·yr). Model parameters (temperature sensitivity coefficient, adsorption capacity, decomposition rate constant, etc.). Based on the carbon storage update function derived from physics-Transformer constraints and cellular automata, the model instantly calculates the inhibition of decomposition rate by anaerobic conditions when the user raises the water level, thus displaying the increase in carbon storage in VR. The system supports parallel operation of multiple scenarios and compares three carbon storage evolution curves in a twin environment to aid in the evaluation of the scheme's merits.
[0059] An interpretability analysis framework is constructed based on the Shapley interpretation method and extended to multimodal data input to ensure the transparency and understandability of prediction results. In this framework, input factors include four main categories: remote sensing image bands, microbial genetic characteristics, soil mineral types, and environmental factors. The system calculates the marginal contribution of each of these input factors to carbon storage prediction. The output is a factor ranking and contribution weight map, allowing users to clearly identify which factors play a major role in the prediction process. For example, in wetland scenarios, hydrological factors may have the highest weight, while in farmland scenarios, fertilization-related factors and mineral protection factors may contribute more. This mechanism not only improves the interpretability of model results but also helps decision-makers identify key regulatory factors, thus providing data support for developing targeted measures. Multimodal input factors include: remote sensing image bands (reflecting vegetation cover and soil spectral characteristics), microbial genetic characteristics, soil mineral types, and environmental factors. The Shapley value definition formula is as follows: , Input factor Marginal contribution to forecasting The set of all input factors. Remove factors a subset of Number of elements in a subset The model uses only the factor set. The Shapley score measures contribution by the difference in predictions with and without the factor, fairly distributing the influence of all input factors. To avoid differences in the dimensionality of factors across different modalities, a normalized Shapley weight is defined: Normalized contribution formula: , :factor The relative contribution (range 0–1). :factor The Shapley value, in wetland scenarios, represents the contribution of hydrological factors. Often the highest; in farmland scenarios, the fertilization intensity factor With mineral protection factors This may be more prominent. Visual output: Factor ranking table: shows factors from largest to smallest contribution; Contribution weight chart: shows the strength of each factor's influence on the prediction results in a bar chart or radar chart; Scenario comparison chart: shows the differences in the contribution of key factors in different ecosystems.
[0060] The entire process, from data collection and model building to prediction output, will be managed on the blockchain to ensure traceability at every stage.
[0061] In a blockchain system, core operations at each stage, including sampling records, model parameter updates, and prediction result generation, are stored using hash locking to ensure the immutability of data and models. At the visualization level, users can view the entire traceability chain from sampling to prediction through the interface, verifying the source and reliability of the results. In carbon trading and ecological assessment scenarios, this mechanism provides transparent and reliable carbon sink data, ensuring the fairness and scientific rigor of trading data. From data collection to model training, parameter updates, and prediction output, the entire process generates blocks and uploads them to the blockchain. Block hash calculation formula: , : No. The hash value of each block, : No. Data from the phase (such as sampling records, sensor data summaries). : No. Model parameters (or parameter summary) for each stage. : No. Summary of prediction results for the stage The hash value of the previous block. Irreversible hash function: Each block contains the hash of the previous block, forming a chain structure that cannot be tampered with. Traceability chain visualization: Timeline view: Displays the time nodes and hashes for each stage (sampling, training, prediction); Flowchart: Users can click to view the original summary of each stage (such as sampling number, sensor ID, model version); Verification interface: By recalculating and comparing hashes, users can verify whether the data and results have been tampered with. Applied research: Verifies whether model training is based on real data; Policy formulation: Ensures the credibility of ecological assessment results; Carbon trading: Provides verifiable carbon sink data to ensure fair and compliant trading.
[0062] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A soil carbon storage accuracy assessment system optimized by a physical information neural network, characterized in that: The system includes a multi-source heterogeneous data intelligent fusion module, which uses multimodal spatiotemporal alignment, multi-physics field coupled sensing, biological-mineral map construction, and dynamic hydrological topology construction to achieve the fusion of multi-source heterogeneous data from satellite hyperspectral imagery, UAV lidar point clouds, microwave remote sensing, ground sensors, electromagnetic and thermal sensing, optical and acoustic monitoring, metagenomic sequencing, XRD mineralogy, hydrology, and DEM, outputting a consistency feature matrix, a hydrological driving factor matrix, and a carbon stability biological index; and a hierarchical physical constraint neural network module, which embeds multi-physics hybrid PINN, enzyme kinetic time-delay constraints, and an electrochemical-diffusion synergistic layer into a deep network. This system achieves unified modeling of surface and deep soil carbon cycle processes, outputting consistent carbon flux and carbon pool prediction results across different layers. A multi-scale 3D carbon storage mapping module generates 3D carbon storage voxel maps on centimeter-level grids based on topologically enhanced fractal kriging interpolation, holographic soil map neural networks, and uncertainty decomposition Monte Carlo methods, outputting confidence intervals, sensitivity heatmaps, and quality labels. A temporal evolution and digital twin module utilizes a self-evolving carbon cycle cellular automata, a physical-transformer dual-channel state space, and multi-indicator stability early warning to achieve historical assimilation, future evolution projection, and stability risk assessment of soil carbon storage. The adaptive transfer learning module is used to achieve rapid adaptation and high-precision prediction of the system under different regions and soil types through cross-mechanism adaptive meta-learning, physical consistency federated distillation and dynamic multi-scale constraint adjustment mechanism; the decision support and visualization module is used to map the prediction results to the VR / AR interactive environment based on immersive twin interaction, cross-modal interpretable framework and blockchain trusted traceability visualization, output factor contribution ranking and carbon storage traceability chain to assist researchers, policymakers and ecological restoration managers in decision-making.
2. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The multimodal spatiotemporal alignment in the multi-source heterogeneous data intelligent fusion module adopts a hierarchical autoregressive attention mechanism, using ground sensor sequences as the time reference, and performs temporal offset correction and spatial reprojection optimization on satellite imagery, UAV point clouds and microwave remote sensing data.
3. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The multi-physics hybrid PINN in the hierarchical physical constraint neural network module introduces the Arrhenius temperature effect and water vapor transport equation in the surface layer, and adopts the Langmuir adsorption kinetic equation in the deep layer. It also achieves consistency between the surface and deep layer prediction results through cross-layer residual connections.
4. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The enzyme kinetic time delay constraint in the hierarchical physical constraint neural network module uses the Michaelis-Menten kinetic model combined with time delay differential terms to describe the lag response of enzyme synthesis to environmental factors, thereby constraining the prediction of carbon decomposition rate.
5. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The holographic soil map neural network in the multi-scale three-dimensional carbon storage mapping module constructs a three-dimensional multilayer map of granular layer, pore layer and capillary water layer based on CT or MRI scan data, and simulates the diffusion and consolidation of dissolved organic carbon through cross-layer information propagation mechanism.
6. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The self-evolving carbon cycle cellular automata in the temporal evolution and digital twin module uses microorganisms, minerals, and organic carbon as intelligent agents. Through reinforcement learning to train interaction rules, it realizes a dynamic evolutionary pattern of carbon sequestration, mineralization, and migration.
7. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The physical-transformer dual-channel state space in the temporal evolution and digital twin module consists of a data-driven channel and a physical equation channel. The data-driven channel encodes multi-source temporal observations, while the physical equation channel uses the state transition matrix to constrain the predicted trajectory and simulates the impact of different management measures on carbon storage under the counterfactual branch.
8. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The adaptive transfer learning module introduces carbon conservation and energy conservation loss functions when aggregating models across regions through physical consistency federated distillation, thereby ensuring the physical consistency of models in different regions while protecting the privacy of monitoring data.
9. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The cross-modal interpretable framework in the decision support and visualization module uses Shapley decomposition to extend to multimodal inputs, quantifies the marginal contributions of remote sensing bands, microbial genes, mineral types and environmental factors to carbon storage prediction, and outputs factor ranking and weight map.
10. The soil carbon storage precision assessment system optimized by physical information neural network according to claim 1, characterized in that: The blockchain-based trusted traceability visualization module in the decision support and visualization module records the entire process of sampling, modeling, and prediction on the blockchain through hash locking to ensure the transparency and immutability of carbon sink data in scientific research, policy, and carbon trading scenarios.