Machine learning based method and system for constructing a three-dimensional prediction model of soil organic matter
By employing a momentum contrast learning architecture with a multi-scale spatial pyramid network and a spectral-spatial-depth triple enhancement strategy, combined with a regularized projection network with differentiable physical constraints and an implicit neural representation model, the problem of predicting the three-dimensional continuum of soil organic matter under sparse soil profile data was solved, generating a high-resolution, physically consistent three-dimensional predictive volume field of soil organic matter.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI
- Filing Date
- 2026-04-27
- Publication Date
- 2026-06-16
AI Technical Summary
Existing technologies struggle to construct a high-resolution, physically consistent three-dimensional continuum prediction field of soil organic matter based on sparse soil profile data. Traditional methods fail to effectively capture the vertical continuity of soil properties, resulting in poor continuity and insufficient physical consistency in the prediction results in the depth direction.
A multi-scale spatial pyramid network is used for feature extraction and fusion. A momentum contrast learning architecture with a triple enhancement strategy of spectral-spatial-depth is used for self-supervised pre-training. A regularized projection network with differentiable physical constraints is used for structured modulation. A probability diffusion model with implicit neural representation as the decoder is used to generate sparse profile labels in reverse. The model is trained and validated by combining noise scheduling weighting and loss functions with volume rendering consistency, generating a continuous three-dimensional soil organic matter volume field.
An end-to-end machine learning solution was developed, from multi-source data acquisition to the generation of three-dimensional continuums. This solution can generate high-resolution, physically consistent three-dimensional predictive fields of soil organic matter, improving the accuracy and continuity of the prediction of the three-dimensional spatial distribution of soil organic matter.
Smart Images

Figure CN122223262A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital soil technology, and in particular to a method and system for constructing a three-dimensional prediction model of soil organic matter based on machine learning. Background Technology
[0002] Soil organic matter is a key indicator for evaluating soil quality and carbon storage, and accurate prediction of its spatial distribution is of great significance for agricultural and environmental management. Existing technologies have widely adopted machine learning methods to achieve two-dimensional spatial mapping of soil organic matter by integrating multi-source data such as remote sensing imagery and topography, effectively improving mapping efficiency.
[0003] Existing methods struggle to extend predictions from two-dimensional space to three-dimensional continuum modeling. The core challenge lies in constructing a physically meaningful, high-resolution three-dimensional prediction field based on sparse soil profile samples. Traditional approaches predict different depth layers independently or simply use depth as a feature input, failing to effectively model the vertical continuity of soil properties. When profile data is scarce, conventional supervised learning models struggle to capture robust vertical distribution patterns, resulting in poor continuity and insufficient physical consistency in the prediction results along the depth direction. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a method for constructing a three-dimensional prediction model of soil organic matter based on machine learning, which solves the problem that existing technologies are unable to construct a high-resolution three-dimensional continuum prediction field with physical consistency using sparse soil profile data.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for constructing a three-dimensional prediction model of soil organic matter based on machine learning, which includes: collecting and preprocessing multi-source data to obtain a preprocessed multi-source aligned dataset; extracting and fusing features through a multi-scale spatial pyramid network to obtain an initial fused feature cube. By performing self-supervised pre-training using a momentum contrastive learning architecture that combines spectral-spatial-depth triple enhancement strategies, and then performing structured modulation using a regularized projection network based on differentiable physical constraints, a modulated structured geographic-deep relational representation is obtained. Using the modulated structured geographic-deep association representation as a condition, a probability diffusion model with implicit neural representation as the decoder is used to inversely generate sparse profile labels, resulting in a continuous three-dimensional soil organic matter volume field. The model was trained and validated using a loss function that combines noise scheduling weighting and volume rendering consistency, resulting in optimized model parameters after training and validation. The optimized model parameters after training and validation are used to predict new regions, and the three-dimensional predictive volume field of soil organic matter at arbitrary resolution is obtained by decoding the implicit neural representation field.
[0007] As a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the preprocessed multi-source aligned dataset includes, Multispectral images, hyperspectral images, and lidar point clouds are acquired to obtain raw multi-source acquisition data, which are then corrected to obtain a preprocessed multi-source aligned dataset.
[0008] In a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the initially fused feature cube includes: A multi-scale spatial pyramid network is constructed, and features are interpreted through heterogeneous receptive field paths and cross-path gating mechanisms. Then, hierarchical encoding and cross-scale interaction are performed through bidirectional paths within the feature pyramid to obtain the multi-scale feature set after interaction. Non-parametric feature recombination based on mutual information and conditional entropy among multiple sources is performed on the multi-scale feature set after interaction. By maximizing inter-group mutual information and minimizing intra-group conditional entropy, feature self-organization fusion is driven to obtain the recombined multi-scale features. The recombined multi-scale features and the preprocessed multi-source aligned dataset are fused through residual connections to obtain the initial fused feature cube.
[0009] As a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the self-supervised pre-training includes: A spectral-spatial-depth triple enhancement strategy is applied to the initially fused feature cube. The spectral-spatial-depth triple enhancement strategy performs adaptive probabilistic masking on specific organic matter-sensitive bands in the spectral dimension, performs clipping within the boundaries of terrain units derived from the digital elevation model in the spatial dimension, and performs random masking and sequence rearrangement based on geostatistical laws in the depth dimension to generate feature view pairs after triple enhancement transformation. The feature view after triple enhancement transformation is input to the momentum contrast learning architecture. Features are extracted by an online encoder and a target encoder that updates momentum. A set of positive and negative sample pairs is constructed based on the proximity constraint of spatial location and the continuity constraint of depth sequence. The improved InfoNCE contrast loss, based on the set of positive and negative samples, adjusts the discriminative power of positive and negative samples through a temperature coefficient and drives the encoder to learn a representation that is invariant to spectral interference, spatial terrain transformation and depth sequence perturbation, thus obtaining a contrast loss value for parameter update. The online encoder parameters are optimized by backpropagation using contrastive loss values and the target encoder parameters are updated by momentum. The trained encoder processes the initial fused feature cube to obtain a geographic-deep association representation.
[0010] As a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the modulated structured geographic-deep association representation includes, Differentiable physical constraints are used to process geographic-deep association representations. By using the differentiable form of energy constraints on soil aggregate formation dynamics and organic matter-mineral complex stabilization processes, specific soil structure formation and carbon stabilization processes are encoded as constraint operators, resulting in a representation of embedded process constraints. The representation of the embedded process constraint is subjected to process-guided feature transformation. The representation is mapped and regularized by differentiable aggregate dynamics and energy operators to obtain the transformed representation. The physical constraint loss function is calculated based on the transformed representation. The physical constraint loss function includes the aggregate formation constraint loss and the stabilization energy constraint loss, and the physical constraint loss value is obtained. A regularized projection network with differentiable physical constraints is combined with physical constraint loss and representation reconstruction loss to jointly optimize network parameters, resulting in a modulated structured geographic-depth association representation that satisfies physical constraints.
[0011] As a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the continuous three-dimensional soil organic matter volume field includes, The probabilistic diffusion model, which uses implicit neural representation as the decoder, takes the modulated structured geographic-deep association representation as the global conditional vector. Using sparse profile labels as local anchor constraints, the modulated structured geographic-depth association representation is fused with sparse profile labels through a cross-scale conditional injection mechanism to initialize the weights of the conditional denoising network. In each iteration of the inverse denoising process, the conditional denoising network couples the current noise latent variable, the denoising time step embedding, and the modulated structured geographic-deep association representation. It dynamically corrects the noise prediction direction through a physically constrained attention mechanism to obtain the denoised implicit neural representation field parameters for the current step. After completing all inverse denoising iterations, the probability diffusion model using implicit neural representation as the decoder obtains the implicit neural representation field parameters and defines a continuous mapping function from three-dimensional spatial coordinates to soil organic matter concentration. The implicit neural representation field parameters are parsed by a decoder composed of a multilayer sensing mechanism. It receives arbitrary query coordinates and combines them with the modulated structured geographic-depth association representation. The soil organic matter value at the query coordinates is obtained through forward propagation, and a continuous three-dimensional soil organic matter volume field is synthesized.
[0012] As a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the optimized model parameters after training and validation include: By combining the noise scheduling weighting and the loss function of volume rendering consistency, the noise prediction error in the diffusion process is coupled with the depth profile error generated by the rendering of the three-dimensional volume field based on continuous soil organic matter, and a joint optimization loss for model optimization is constructed. The joint optimization loss drives parameter updates in the probabilistic diffusion model, which uses implicit neural representation as a decoder, through a backpropagation process. The accuracy and continuity of the model's generation of a continuous three-dimensional soil organic matter volume field are evaluated through a validation set. After iterative training and validation, the conditionalized denoising network and decoder with implicit neural representation as the decoder in the probability diffusion model with parameter optimization are saved, and the optimized model parameters after training and validation are obtained.
[0013] As a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the prediction of the new area includes: The probabilistic diffusion model with implicit neural representation as decoder is loaded with optimized model parameters after training and validation. Feature extraction, self-supervised contrastive learning and regularized projection are performed on the preprocessed multi-source aligned dataset of the new region to generate a modulated structured geographic-deep association representation of the new region.
[0014] As a preferred embodiment of the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in this invention, the three-dimensional prediction field of soil organic matter includes, The probabilistic diffusion model, which uses implicit neural representation as the decoder, runs an inverse generation process conditioned on the modulated structured geographic-deep association representation of the new region, and outputs the implicit neural representation field parameters of the new region. The probabilistic diffusion model, which uses implicit neural representation as a decoder, queries the implicit neural representation field parameters of the new region through the decoder to generate a three-dimensional predictive volume field of soil organic matter in the new region.
[0015] Secondly, the present invention provides a soil organic matter three-dimensional prediction model construction system based on machine learning, including a fusion module, which collects and preprocesses multi-source data to obtain a preprocessed multi-source aligned dataset, and performs feature extraction and fusion through a multi-scale spatial pyramid network to obtain an initial fused feature cube. The training module performs self-supervised pre-training using a momentum contrastive learning architecture that combines a spectral-spatial-depth triple enhancement strategy, and performs structured modulation using a regularized projection network based on differentiable physical constraints, resulting in a modulated structured geographic-deep relational representation. The reverse generation module, using the modulated structured geographic-deep association representation as a condition, utilizes a probability diffusion model with implicit neural representation as the decoder to reverse generate sparse profile labels, thereby obtaining a continuous three-dimensional soil organic matter volume field. The validation module uses a loss function that combines noise scheduling weighting and volume rendering consistency to train and validate the model, and obtains the optimized model parameters after training and validation. The prediction module uses the optimized model parameters after training and validation to predict new regions and obtains the three-dimensional prediction field of soil organic matter at arbitrary resolution through implicit neural representation field decoding.
[0016] The beneficial effects of this invention are as follows: A momentum contrastive learning architecture combining spectral, spatial, and depth enhancement strategies is used for self-supervised pre-training, and a regularized projection network based on differentiable physical constraints is used for structured modulation to obtain a modulated structured geographic-depth correlation representation. Using this modulated structured geographic-depth correlation representation as a condition, a probabilistic diffusion model with implicit neural representation as the decoder is used to inversely generate sparse profile labels, resulting in a continuous three-dimensional soil organic matter volume field. The model is then trained and validated using a loss function combining noise scheduling weighting and volume rendering consistency to obtain optimized model parameters. These optimized parameters are used to predict new regions, and the three-dimensional predicted volume field of soil organic matter at arbitrary resolution is obtained through decoding the implicit neural representation field. From multi-source data acquisition, feature fusion, self-supervised and physically constrained representation learning, to diffusion model-driven three-dimensional continuum generation, an end-to-end machine learning solution is formed. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning.
[0019] Figure 2 A schematic diagram of a system for constructing a three-dimensional prediction model of soil organic matter based on machine learning. Detailed Implementation
[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0021] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0022] Secondly, the term "one embodiment" or "example" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the invention. The appearance of an embodiment in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments.
[0023] Reference Figures 1-2 This is one embodiment of the present invention, which provides a method for constructing a three-dimensional prediction model of soil organic matter based on machine learning, including the following steps: S1. Collect and preprocess multi-source data to obtain a preprocessed multi-source aligned dataset. Then, use a multi-scale spatial pyramid network to extract and fuse features to obtain an initial fused feature cube.
[0024] S1.1 Acquire multispectral images, hyperspectral images, and lidar point clouds to obtain raw multi-source acquisition data, and perform corrections to obtain preprocessed multi-source aligned datasets.
[0025] Furthermore, for specific research areas, multispectral images, hyperspectral images, and lidar point cloud data of the land surface were collected to form raw multi-source acquisition data. The multispectral and hyperspectral images were radiometrically calibrated to convert digital quantization values into surface reflectance, and atmospheric correction was performed to eliminate the effects of atmospheric scattering and absorption. The lidar point cloud was filtered to remove noise points, and ground points and non-ground points were separated by a point cloud classification algorithm. A high-precision digital elevation model was generated based on the ground points. All radiometrically and geometrically corrected multispectral images, hyperspectral images, and digital elevation model data were uniformly resampled to the same spatial resolution and strictly registered to the same geographic coordinate system to form a preprocessed multi-source aligned dataset.
[0026] S1.2 Construct a multi-scale spatial pyramid network, interpret features through heterogeneous receptive field paths and cross-path gating mechanisms, and then perform hierarchical encoding and cross-scale interaction through bidirectional paths within the feature pyramid to obtain the multi-scale feature set after interaction.
[0027] Furthermore, a multi-scale spatial pyramid network is constructed, comprising three parallel feature extraction paths employing small-sized, medium-sized, and large-sized dilated convolutional kernels, respectively, forming heterogeneous receptive field paths to capture local texture details, feature outlines, and global contextual information from the image. A cross-path gating mechanism dynamically adjusts the information flow between paths, allowing the network to adaptively select and fuse features from different paths based on the input content. The extracted features are then input into a network with a multi-level pyramid structure. At each level of the pyramid, features transmit high-level semantic information through a top-down upsampling path and fine-grained spatial information through a bottom-up downsampling path, resulting in a multi-scale feature set.
[0028] S1.3. Perform non-parametric feature recombination based on mutual information and conditional entropy among multiple sources on the multi-scale feature set after interaction. By maximizing inter-group mutual information and minimizing intra-group conditional entropy, feature self-organization fusion is driven to obtain the recombined multi-scale features.
[0029] Furthermore, the multi-scale feature set after interaction is analyzed. Mutual information between feature groups from different sources is extracted to measure their statistical dependence, and conditional entropy of features within each group is extracted to assess their uncertainty or redundancy. Based on the principles of information theory, a non-parametric feature recombination operation is implemented. The objective function is to maximize the mutual information between different feature groups to promote cross-modal or cross-scale information sharing and collaboration; at the same time, the conditional entropy of features within a group is minimized to suppress feature redundancy and improve the consistency of features within a group, generating recombined multi-scale features.
[0030] S1.4 The recombined multi-scale features and the preprocessed multi-source aligned dataset are fused through residual connections to obtain the initial fused feature cube.
[0031] Furthermore, the reconstructed multi-scale features are concatenated with the preprocessed multi-source aligned dataset along the channel dimension to form an extended feature tensor. This extended feature tensor is then fused element-wise with the original input data of the preprocessed multi-source aligned dataset through a residual connection structure. The weight parameters in the residual connection can be learned through training to adaptively adjust the contribution ratio between the original low-level information and the deeply reconstructed information. This fusion operation combines the deeply reconstructed and filtered high-level semantic features with the detailed original input data to obtain the initial fused feature cube.
[0032] S2. Self-supervised pre-training is performed using a momentum contrastive learning architecture that combines a spectral-spatial-depth triple enhancement strategy. The structured projection network based on differentiable physical constraints is then used for structured modulation to obtain the modulated structured geographic-deep association representation.
[0033] S2.1 Apply a spectral-spatial-depth triple enhancement strategy to the initially fused feature cube. The spectral-spatial-depth triple enhancement strategy performs adaptive probabilistic masking on specific organic matter-sensitive bands in the spectral dimension, performs clipping within the boundaries of terrain units derived from the digital elevation model in the spatial dimension, and performs random masking and sequence rearrangement based on geostatistical laws in the depth dimension to generate feature view pairs after triple enhancement transformation.
[0034] Furthermore, in the spectral dimension, based on the known response curves of soil organic matter sensitive spectral bands, the corresponding feature channels in the initially fused feature cube are randomly zeroed out using adaptive probability to simulate scenarios where some spectral information is missing under different environmental conditions. In the spatial dimension, based on the digital elevation model in the preprocessed multi-source aligned dataset, basic topographic units such as ridges, valleys, and slopes are delineated through hydrological analysis or topographic classification methods. The spatial clipping operation in the enhancement transformation is strictly limited to within the boundary of the same topographic unit to ensure that the clipped area has consistent topographic significance. In the depth dimension, based on the vertical correlation structure described by the variogram model in soil profile geostatistics, the feature slice sequence in the depth direction of the initially fused feature cube is randomly masked, or the depth slice order is randomly rearranged according to statistically reasonable permutations to simulate incomplete sampling or layer perturbations. The above three enhancement transformations are applied independently or in combination to the same initially fused feature cube to generate multiple different transformed views, from which positive sample pairs are formed for comparative learning, generating a triple-enhanced feature view pair.
[0035] S2.2. The feature view after triple enhancement transformation is input to the momentum contrast learning architecture. Features are extracted by an online encoder and a target encoder for momentum update. A set of positive and negative sample pairs is constructed based on the proximity constraint of spatial location and the continuity constraint of depth sequence.
[0036] Furthermore, the feature view pairs after triple augmentation transformation are input into the momentum contrastive learning architecture. This architecture contains two encoder branches with the same structure but different parameter update methods: an online encoder and a target encoder. The two views of the feature view pairs after triple augmentation transformation are input into the online encoder and the target encoder, respectively, to extract the corresponding feature vectors. When constructing the positive and negative sample pair sets, the positive sample pairs consist of feature vectors from the same spatial location but after undergoing different triple augmentation transformations. The construction of the negative sample pair set is subject to two constraints: the spatial proximity constraint requires excluding samples with excessively close geographical coordinates from the negative sample set to avoid penalizing reasonable spatial autocorrelation; the depth sequence continuity constraint requires that samples with adjacent depth layers on the same vertical profile should not be considered strictly negative samples under certain conditions to maintain vertical correlation. Based on these constraints, qualified negative sample feature vectors are selected from the batch data and, together with the positive sample feature vectors, constitute the positive and negative sample pair set.
[0037] Specifically, domain knowledge of geospatial and vertical sequences is embedded in the contrastive learning sample construction process as a constraint. Because geographically neighboring points may have highly similar soil properties, treating them as negative pairs would force the model to separate representations that should be similar, thus disrupting spatial continuity learning. By using spatial proximity constraints, the generation of such geographical false negative sample pairs is actively avoided during the sample construction stage. This ensures that the model can learn spatially smooth feature representations that conform to the first law of geography, acknowledges the correlation of vertical properties of soil profiles, and does not completely treat different depth layers as independent samples, thereby guiding the model to learn the correlation patterns between layers.
[0038] S2.3. An improved InfoNCE contrast loss based on positive and negative sample pair set calculation is obtained by adjusting the discriminative power of positive and negative samples through temperature coefficient and driving the encoder to learn a representation that is invariant to spectral interference, spatial terrain transformation and depth sequence perturbation, thus obtaining a contrast loss value for parameter update.
[0039] Furthermore, an improved InfoNCE contrastive loss is calculated based on the constructed positive and negative sample pairs. This loss function calculates the similarity between the anchor sample features and the positive sample features, as well as the similarity between the anchor sample features and the features of each sample in the filtered negative sample set. A temperature coefficient is used to adjust the sharpness of the similarity distribution, controlling the model's focus on difficult negative samples. The summation term in the denominator of the loss calculation does not traverse all other samples in the batch, but only calculates for samples in the valid hard negative sample index set that meet the geographical filtering and similarity filtering conditions. Geographical filtering excludes samples that are too geographically close to the anchor sample, while similarity filtering selects difficult negative samples whose feature similarity to the anchor sample is close to that of positive samples. Through dynamic selective calculation, the contrastive loss function can focus on the difficult negative samples that are most helpful for the discrimination task during the optimization process, while avoiding unreasonable penalties for geographically neighboring samples. The calculated contrastive loss value quantifies the representation learning quality of the current batch of samples.
[0040] Specifically, the expression is: ; in, To compare the losses, This represents the number of anchor point samples in the batch. For anchor point sample index, anchor point Similarity to its positive samples, anchor point With the Similarity of negative samples anchor point The effective hard negative sample index set, For negative sample index, For temperature parameters; ; ; ; in, For anchor point samples Features anchor point Positive sample features For the first Feature representation of each negative sample For the sample and Geographical distance, The similarity difference threshold, The geographic proximity threshold; Hard negative sample set Construction: Loss calculation no longer iterates through all negative samples; instead, a refined set of hard negative samples is constructed through two-stage filtering. Geographic filtering: Exclude geographically proximate locations ( < Samples of ) are used to avoid penalizing spatial autocorrelation; Similarity filtering: retain samples with similarity to positive samples ( Hard negative samples; Dynamic selectivity calculation: loss denominator only applies to The samples in the sample are summed to avoid false negative penalties for geographically neighboring samples, and the focus is on hard negative samples with the most information. The computational cost changes dynamically with the number of effective negative samples. Spatial continuity of soil properties ( Threshold) and feature discrimination difficulty ( The threshold is directly encoded into the structure of the loss function, aligning the contrastive learning process with domain characteristics; S2.4. The online encoder parameters are optimized by backpropagation using the contrastive loss value and the target encoder parameters are updated by momentum. The trained encoder processes the initial fused feature cube to obtain the geographic-deep association representation.
[0041] Furthermore, gradients are calculated using the backpropagation algorithm, and the trainable parameters of the online encoder are updated using an optimizer. The parameters of the target encoder are not updated directly through backpropagation, but rather using a momentum update method; that is, the parameters of the target encoder are an exponential moving average of the historical values of the online encoder parameters. This momentum update mechanism makes the feature representation provided by the target encoder more stable, which is beneficial for the training dynamics of contrastive learning. After multiple rounds of iterative training until the model converges, the trained online encoder possesses the ability to extract features robust to spectral-spatial-depth disturbances. The initial fused feature cube to be processed is input into this trained online encoder, which performs forward propagation calculations on it, outputting a geographic-depth correlation representation.
[0042] Specifically, a stable training mechanism is closely integrated with a spectral-spatial-depth triple enhancement strategy, as well as sample construction and loss calculation incorporating domain knowledge. The geographic-depth relational representation generated by the trained encoder, for example, has undergone terrain-based cell-based pruning enhancement and geographic proximity protection; and has undergone depth sequence perturbation and continuity constraints. This representation can capture continuous patterns in the vertical profile, resulting in a strong feature representation that deeply integrates multi-source information, is invariant to various disturbances, and encodes prior knowledge of geospatial and vertical structures. S2.5. Regularized projection network processing of differentiable physical constraints for geographic-deep association representation: By using the differentiable form of energy constraints on soil aggregate formation dynamics and organic matter-mineral complex stabilization processes, specific soil structure formation and carbon stabilization processes are encoded as constraint operators, resulting in a representation of embedded process constraints.
[0043] Furthermore, a regularized projection network with differentiable physical constraints is constructed, using a geographic-depth correlation representation as input. The physical relationships concerning interparticle cementation and pore evolution in soil aggregate formation dynamics, as well as the energy balance relationships involved in the stabilization of organic matter-mineral complexes, are abstracted and formalized into differentiable mathematical operators. For example, the dynamic constraints of aggregate formation can be expressed as a function in the feature space representing the ratio of the gradient of the vector along the depth direction to the gradient along the horizontal direction, simulating the anisotropy of aggregate stability; the stabilization energy constraints can be expressed as boundary conditions that the vector magnitude or energy function must satisfy, simulating carbon conservation mechanisms. These physical relationships are encoded into a series of constraint operators, which serve as activation functions or regularization terms for specific layers within the regularized projection network with differentiable physical constraints. These operators perform preliminary transformation and filtering on the input geographic-depth correlation representation, resulting in a representation of embedded process constraints that incorporates the aforementioned soil process constraints.
[0044] S2.6. Perform process-guided feature transformation on the representation of the embedded process constraint, and map and regularize the representation through differentiable aggregate dynamics and energy operators to obtain the transformed representation.
[0045] Furthermore, the regularized projection network of differentiable physical constraints performs a process-guided feature transformation on the representation of the embedded process constraints. This transformation is achieved through specific differentiable layers in the network, which implement the specific computation of the aforementioned aggregation dynamics operator and energy operator. For example, the physical field gradient is approximated by calculating the difference between the representation of the embedded process constraints in different spatial directions, and the influence of the dynamic process on the features is simulated through a pre-defined function mapping, and this state is adjusted using a differentiable function. This involves performing a nonlinear mapping and distortion of the representation of the embedded process constraints in the feature space, inspired by physical laws, while applying regularization to control the amplitude and smoothness of the transformation, outputting a transformed representation shaped by the laws of the physical process.
[0046] Specifically, a feature space manifold transformation driven by physical laws is performed. The core idea is that the physical constraints embedded in the previous step are static and point-like, while the transformation implemented in this step through differentiable operators is dynamic and propagating. This simulates the propagation and effects of physical processes in space and time. For example, aggregate dynamics not only constrain the state of a single point but also implicitly involve spatial correlations of properties (such as stability propagation). The transformation implemented through differentiable operators can propagate the physical constraint information of a point through hierarchical computation of the network, influencing the representations of its neighboring points in the feature space, thereby inducing a spatially covariant pattern consistent with physical laws throughout the representation tensor. The transformation by energy operators guides the representation vector to move towards energy-stable submanifold regions in the feature space. By reorganizing and reshaping the feature space according to the embedded physical rules, the transformed representation not only contains the original data information but its organizational structure also metaphorically represents the dynamic and steady-state characteristics of soil processes, providing a deep structural guarantee for generating physically consistent predictions.
[0047] S2.7. Calculate the physical constraint loss function based on the transformed representation. The physical constraint loss function includes the aggregate formation constraint loss and the stabilization energy constraint loss, and obtain the physical constraint loss value.
[0048] Furthermore, a physical constraint loss function is calculated based on the transformed representation. The cluster formation constraint loss is constructed by comparing the differences in the magnitude of change of the transformed representation in the vertical and horizontal directions, for example, by calculating a measure of the difference between its gradient norm ratio and the ideal ratio. The stabilization energy constraint loss is constructed by evaluating the deviation between the energy state implied in the transformed representation and a pre-defined stable energy benchmark, and may be weighted in conjunction with indicators representing local stability. The interaction constraint loss is constructed by comparing the ratio of the distance between feature vectors at different spatial locations in the transformed representation to the actual physical distance between these locations, penalizing feature distributions that do not conform to spatial decay laws. These three loss values are calculated separately and then weighted and summed according to pre-defined or adaptive weights to obtain the total physical constraint loss value, which quantifies the degree to which the transformed representation violates pre-defined physical laws.
[0049] Specifically, the expression is:
[0050] in, For the total physical constraint loss, To stabilize energy confinement loss, To stabilize energy confinement loss, For interaction constraint loss, The weighting coefficients for the constraints on aggregate formation. The weighting coefficients are used to stabilize the energy constraints. These are the weighting coefficients of the interaction constraints; Aggregate formation constraint loss ; in, For the first The transformed representation vector of each effective spatial location For the gradient operator in the depth direction, For the gradient operator in the horizontal direction, It is a constant. The total number of effective spatial locations, For the index of effective spatial location, Stabilization energy confinement loss ; in, To represent vectors The energy function, The preset energy baseline value, This is an index representing the stability of a vector; Interaction constraint loss ; in, Position in 3D space and location The Euclidean distance between them For the first The transformed representation vector of each effective spatial location; This serves as the primary index for the effective spatial location within the three-dimensional soil mass field. This serves as a subordinate index for the effective spatial location within the three-dimensional soil mass field; S2.8. A regularized projection network with differentiable physical constraints is combined with the physical constraint loss value and the representation reconstruction loss to jointly optimize the network parameters, thereby obtaining a modulated structured geographic-deep association representation that satisfies physical constraints.
[0051] Furthermore, the training of the differentiable physical constraint regularized projection network is performed through joint optimization. The overall optimization objective consists of two parts: a physical constraint loss value and a representation reconstruction loss. The representation reconstruction loss measures the ability of the network's output modulated structured geographic-deep association representation to retain key information after the input geographic-deep association representation has undergone physical constraint transformation, thus preventing the physical constraint process from excessively distorting the original effective information. During training, the gradients of the physical constraint loss value and the representation reconstruction loss with respect to the parameters of the differentiable physical constraint regularized projection network are simultaneously calculated via backpropagation, and the network parameters are updated using the optimizer. Through iterative optimization, the differentiable physical constraint regularized projection network learns how to transform the input geographic-deep association representation so that the output representation can retain the effective information of the input data to the maximum extent while satisfying the preset soil physical law constraints to the maximum extent. When training converges, the differentiable physical constraint regularized projection network, after processing the geographic-deep association representation, obtains a modulated structured geographic-deep association representation that satisfies the physical law constraints.
[0052] S3. Using the modulated structured geographic-deep association representation as a condition, the probability diffusion model with implicit neural representation as the decoder is used to reverse generate sparse profile labels to obtain a continuous three-dimensional soil organic matter volume field.
[0053] S3.1 The probabilistic diffusion model with implicit neural representation as the decoder uses the modulated structured geographic-deep association representation as the global conditional vector.
[0054] Furthermore, the modulated structured geo-deep association representation is first processed by a conditional encoding network, typically a multilayer perceptron, which maps the modulated structured geo-deep association representation to a fixed-dimensional global conditional vector. This global conditional vector persists and is continuously used throughout the diffusion generation process, providing global, high-level guidance information on the geographic background and soil physical structure of the entire target area. The generation of the global conditional vector signifies the formal integration of the modulated structured geo-deep association representation into the probabilistic diffusion model, which uses implicit neural representations as the decoder.
[0055] Specifically, the structured representation, which integrates spectral, spatial, topographic, and depth information and is shaped by physical rules, encodes the regional-scale soil formation environment and property distribution patterns. Using this as a global conditional vector provides a priori blueprint or style guidance for the generation process of the diffusion model, ensuring that the generated 3D volume field maintains consistency with the regional geographical background and physical laws in its overall pattern. For example, it ensures that the generated results macroscopically conform to the topographic sequence distribution patterns or soil type spatial patterns of the region—something that cannot be achieved using raw data or ordinary features.
[0056] S3.2. Using sparse profile labels as local anchor point constraints, the modulated structured geographic-depth association representation is fused with sparse profile labels through a cross-scale conditional injection mechanism to initialize the weights of the conditional denoising network.
[0057] Furthermore, sparse profile labels, as a set of discrete soil organic matter observations with precise geographic locations and depths, are input into a probabilistic diffusion model using implicit neural representations as the decoder. These sparse profile labels are first converted into a format suitable for model processing, for example, by assigning a specific code to each label location. A cross-scale conditional injection mechanism is responsible for fusing global conditional vectors and sparse profile label information. This mechanism may employ operations such as cross-attention or feature concatenation to interact and align the local precision information contained in the sparse profile labels with the overall background information provided by the global conditional vectors at multiple network layers. Through cross-scale fusion, sparse profile labels act as strong constraints anchoring the values of the generation process at specific spatial locations and depths, while the global conditional vectors provide background knowledge on how regions should transition reasonably between these anchor points. The fused information is used to initialize or modulate the weights or biases of corresponding layers in the conditional denoising network, enabling the conditional denoising network to be aware of the global background and local observations from the initial state.
[0058] Specifically, the continuous spatial distribution pattern (global prior) learned from unlabeled data and carried by the modulated structured geographic-deep association representation is deeply fused with sparse but accurate field measurements (local anchors) at the feature level. This is equivalent to constructing an initial belief field for the model from the very beginning of the generation process, which conforms to large-scale geophysical laws and is strictly aligned with real data at key observation points. The fusion ensures that the generation process does not randomly explore in an unconstrained space, but rather takes place in a highly structured solution space jointly defined by knowledge (global representation) and data (local labels), greatly reducing the uncertainty of generation and guaranteeing the prediction accuracy at the observation points.
[0059] S3.3 Conditional denoising network couples the current noise latent variable, the denoising time step embedding, and the modulated structured geographic-deep association representation in each iteration of the inverse denoising process. It dynamically corrects the noise prediction direction through a physically constrained attention mechanism to obtain the denoised implicit neural representation field parameters for the current step.
[0060] Furthermore, the physically-constrained attention mechanism is a key component of the conditional denoising network. It dynamically acquires the importance weights of different parts of the modulated structured geographic-deep association representation for the current denoising step and different spatial locations. For example, in the early stages of denoising, the attention mechanism may focus more on information about macroscopic terrain units in the representation; in the later stages, it may focus more on information related to the vertical distribution pattern of organic matter. Through dynamic, content-aware attention weighting, the modulated structured geographic-deep association representation is effectively and selectively injected into the denoising computation to correct the predicted noise direction at each step, thereby guiding the denoising process towards a direction consistent with physical laws and geographical context. The output of this step is the updated, less noisy implicit neural representation field parameters.
[0061] Specifically, through a physically constrained attention mechanism, the modulated structured geographic-deep relational representation is no longer a static background but becomes a dynamic and intelligent navigator. This is achieved in two aspects: First, physical constraint guidance. The calculation of attention weights incorporates considerations of the aforementioned physical constraints. For example, when predicting the attributes of a point, the attention mechanism tends to focus on the representation information of other locations that are geographically close to that point, have similar terrain, and are strongly correlated under physical constraints. This naturally embeds spatial autocorrelation and physical consistency. Second, dynamism and content awareness. The denoising process progresses from coarse to fine, and the attention mechanism adaptively emphasizes different aspects of the representation at different stages. Early on, it uses macroscopic structural information to determine the general distribution, and later, it uses detailed information for refinement. The mechanism ensures that physical knowledge and geographic information are utilized most effectively throughout the generation process, guiding the generation of a highly physically consistent and spatially detailed 3D field.
[0062] S3.4. After completing all inverse denoising iterations, the probability diffusion model using implicit neural representation as the decoder obtains the implicit neural representation field parameters and defines a continuous mapping function from three-dimensional spatial coordinates to soil organic matter concentration.
[0063] Furthermore, at each step, the conditional denoising network predicts noise based on the current noise latent variable, the time step, and conditional information, and subtracts the predicted noise from the current noise latent variable to obtain the updated noise latent variable, which serves as the input for the next step. This process is repeated until all the preset number of inverse denoising iterations are completed. The output of the final step, i.e., the fully denoised latent variable, is defined as the implicit neural representation field parameter. The implicit neural representation field parameter is a high-dimensional tensor that uniquely parameters a continuous function whose input is three-dimensional spatial coordinates and whose output is the predicted soil organic matter concentration value. Therefore, the implicit neural representation field parameter itself does not directly store the value of each voxel, but rather encodes all the information required to generate the entire continuous three-dimensional volume field.
[0064] Specifically, it is inherently continuous, allowing queries at arbitrary spatial resolutions, breaking through the limitations of fixed grid resolutions and enabling truly arbitrary resolution predictions. Secondly, the implicit neural representation field parameters are typically more compact than voxel grids of the same precision, reducing the data dimensionality and complexity required for diffusion models. Most importantly, modeling the generation process as learning the parameters of a continuous function aligns better with the continuous spatial variation of soil properties. Learning these parameters through the diffusion process is equivalent to exploring and converging in the function space to a continuous function that best satisfies global conditions, local anchor points, data distribution, and physical constraints simultaneously. S3.5 The implicit neural representation field parameters are parsed by a decoder composed of a multilayer sensing mechanism. It receives arbitrary query coordinates and combines them with the modulated structured geographic-deep association representation. The soil organic matter value at the query coordinates is obtained through forward propagation, and a continuous three-dimensional soil organic matter volume field is synthesized.
[0065] Furthermore, during the query, the 3D spatial coordinates of any target and its corresponding modulated structured geographic-depth relational representation (usually contextual features of the local area where the coordinates are located) are simultaneously input into the decoder. The decoder performs nonlinear transformation and fusion of the input coordinates and conditional features through its multi-layer feedforward network. In the final layer of the decoder, a linear mapping or constrained activation function is used to output a scalar value, which is the predicted soil organic matter concentration at the query coordinates. By densely querying coordinates and obtaining predicted values across the entire target 3D spatial range, or through vectorized computation, a continuous 3D soil organic matter volume field covering the entire region can be synthesized.
[0066] Specifically, the decoder's input includes spatial coordinates, which makes it represent a continuous field. Second, the decoder again combines the modulated structured geographic-deep relational representation. This means that when generating specific values, it not only relies on the global function information encoded by the implicit neural representation field parameters, but also dynamically combines the geographic context features of the specific location. This achieves a secondary fusion of global functions and local features at the decoding time, which can further improve the local accuracy and contextual rationality of the prediction. As a result, the final synthesized continuous soil organic matter 3D volume field maintains the overall data distribution and physical consistency learned from the diffusion model, while incorporating more refined geographic background information at each local location. This generates high-quality 3D prediction results that are rich in detail, spatially coordinated, and physically reliable.
[0067] S4. The model is trained and validated using a loss function that combines noise scheduling weighting and volume rendering consistency to obtain the optimized model parameters after training and validation.
[0068] S4.1. Combining the noise scheduling weighting and volume rendering consistency loss function, the noise prediction error in the diffusion process is coupled with the depth profile error generated based on the three-dimensional volume field rendering of continuous soil organic matter to construct a joint optimization loss for model optimization.
[0069] Furthermore, the loss function combining noise scheduling weighting and volumetric rendering consistency comprises two core computational branches. The first branch calculates the noise prediction error during the diffusion process, i.e., the standard diffusion model loss. It calculates the difference between the model's predicted noise and the actual added noise at different noise levels, and assigns different weights to the errors at different time steps through a noise scheduling strategy to emphasize learning at specific stages. The second branch calculates the volumetric rendering consistency error. This branch first utilizes a continuous 3D soil organic matter volume field generated by a probabilistic diffusion model with implicit neural representation as the decoder. Then, using volumetric rendering technology, it integrates the 3D volume field along virtual rays that perfectly correspond to the sampling positions and depths of the actual sparse profile, synthesizing the corresponding predicted depth profile curve. Subsequently, it compares the synthesized predicted depth profile with the actual depth profile measurements in the sparse profile labels, calculating the difference between them, such as the mean squared error. Finally, the weighted noise prediction error and the volumetric rendering consistency error are linearly added according to preset coefficients to construct a joint optimization loss for end-to-end optimization of the entire model generation quality.
[0070] S4.2 The joint optimization loss drives the parameter update in the probabilistic diffusion model with implicit neural representation as the decoder through the backpropagation process, and evaluates the accuracy and continuity of the model in generating a continuous three-dimensional soil organic matter volume field through the validation set.
[0071] Furthermore, during training, after the joint optimization loss is calculated, its gradient with respect to all trainable parameters in the probabilistic diffusion model using implicit neural representations as the decoder is calculated via backpropagation. These parameters include the weights of the conditional denoising network, the conditional encoding network, and the implicit neural representation decoder. An optimizer, such as a variant of stochastic gradient descent, is used to update these parameters based on the calculated gradients. Simultaneously, the training process is monitored and evaluated using a validation set. After each training cycle or a fixed number of iterations, a complete generation process is run on the validation set using the current model parameters to generate a continuous three-dimensional soil organic matter volume field, and accuracy and continuity metrics are calculated on the validation set. Accuracy metrics can include a measure of the error between predicted and true values at validation points, while continuity metrics assess the smoothness of the generated volume field in the vertical direction or its conformity to physical constraints. The changing trends of the validation set metrics are observed to determine whether the model is overfitting or converging.
[0072] S4.3 After iterative training and validation, the conditionalized denoising network and decoder with implicit neural representation as the decoder in the probability diffusion model with optimized parameters are saved, and the optimized model parameters after training and validation are obtained.
[0073] Furthermore, the training and validation process continues iteratively until a preset stopping criterion is met, such as the joint loss or key accuracy metrics on the validation set no longer significantly improving over multiple consecutive cycles, or the maximum number of training iterations being reached. Once the stopping criterion is met, the training process terminates. At this point, the optimal parameters of the core trainable components in the probabilistic diffusion model using implicit neural representations as decoders—namely, the conditionally optimized denoising network with optimized parameters and the decoder composed of a multilayer perceptron used to parse the parameters of the implicit neural representation field—are extracted from memory and serialized and saved to a storage medium. This saved set of parameters constitutes the optimized model parameters after training and validation. These optimized model parameters after training and validation are the necessary knowledge carriers of the learned data distribution, physical laws, and task constraints for reconstructing and running the entire probabilistic diffusion model using implicit neural representations as decoders when subsequently predicting new regions.
[0074] S5. Using the optimized model parameters after training and validation, predict the new region, and obtain the three-dimensional prediction field of soil organic matter at arbitrary resolution by decoding the implicit neural representation field.
[0075] S5.1 Load the optimized model parameters after training and validation into the probability diffusion model with implicit neural representation as the decoder, and perform feature extraction, self-supervised contrastive learning and regularized projection on the preprocessed multi-source aligned dataset of the new region to generate a modulated structured geographic-deep association representation of the new region.
[0076] Furthermore, the probabilistic diffusion model, using implicit neural representations as decoders, first loads previously saved, trained, and validated optimized model parameters from the storage medium and uses these parameters to initialize the network weights within the model. Then, for the new target region, its corresponding preprocessed multi-source alignment dataset is obtained. This preprocessed multi-source alignment dataset is input into a multi-scale spatial pyramid network with loaded parameters for feature extraction and fusion, obtaining an initial fused feature cube for the new region. Next, this initial fused feature cube is input into a self-supervised contrastive learning momentum encoder with loaded parameters. This encoder applies a spectral-spatial-depth triple enhancement strategy and uses forward propagation driven by an improved InfoNCE loss mechanism to extract the geographic-depth association representation of the new region. Finally, this geographic-depth association representation is input into a differentiable, physically constrained, regularized projection network with loaded parameters. After a physical constraint regularization transformation, a modulated, structured geographic-depth association representation of the new region is output.
[0077] S5.2 The probabilistic diffusion model with implicit neural representation as the decoder runs a reverse generation process with the modulated structured geographic-deep association representation of the new region as the condition, and outputs the implicit neural representation field parameters of the new region.
[0078] Furthermore, after obtaining the modulated structured geographic-deep association representation of the new region, the probabilistic diffusion model, using the implicit neural representation as a decoder, takes this representation as a conditional input. If needed, a small number of sparse profile labels that may exist in the new region can be input as local anchor constraints. The model runs a complete inverse generation process, starting with random noise conforming to a standard Gaussian distribution and iteratively calling a conditional denoising network with loaded parameters. In each denoising iteration, the conditional denoising network receives the current noise latent variable, the current denoising time step information, and the modulated structured geographic-deep association representation of the new region, and predicts the noise based on a physically constrained attention mechanism, gradually removing the noise. After completing all the preset number of inverse denoising iterations, the final output denoising result is the implicit neural representation field parameter of the new region, which encodes a continuous function of the three-dimensional spatial distribution of soil organic matter in the new region.
[0079] S5.3 The probability diffusion model using implicit neural representation as the decoder queries the implicit neural representation field parameters of the new region through the decoder to generate the three-dimensional prediction field of soil organic matter in the new region.
[0080] Furthermore, the probabilistic diffusion model, using implicit neural representations as the decoder, invokes a decoder composed of a multilayer perceptron with pre-loaded parameters. The user specifies the two-dimensional horizontal and vertical depth extent of the target prediction region, as well as the desired spatial query resolution. The decoder receives each coordinate point on a regular three-dimensional coordinate grid defined by these parameters and performs forward propagation computation in conjunction with the feature context of the corresponding location in the modulated structured geographic-depth association representation of the new region. For each input three-dimensional coordinate, the decoder outputs a predicted soil organic matter concentration value. By traversing or batch processing all target coordinate points and organizing the output predicted values into a three-dimensional array format according to coordinates, a three-dimensional predicted volume field of soil organic matter with the user-specified resolution covering the new target region is finally generated.
[0081] This embodiment also provides a system for constructing a three-dimensional prediction model of soil organic matter based on machine learning, including: a fusion module, which collects and preprocesses multi-source data to obtain a preprocessed multi-source aligned dataset, and performs feature extraction and fusion through a multi-scale spatial pyramid network to obtain an initial fused feature cube; The training module performs self-supervised pre-training using a momentum contrastive learning architecture that combines a spectral-spatial-depth triple enhancement strategy, and performs structured modulation using a regularized projection network based on differentiable physical constraints, resulting in a modulated structured geographic-deep relational representation. The reverse generation module, using the modulated structured geographic-deep association representation as a condition, utilizes a probability diffusion model with implicit neural representation as the decoder to reverse generate sparse profile labels, thereby obtaining a continuous three-dimensional soil organic matter volume field. The validation module uses a loss function that combines noise scheduling weighting and volume rendering consistency to train and validate the model, and obtains the optimized model parameters after training and validation. The prediction module uses the optimized model parameters after training and validation to predict new regions and obtains the three-dimensional prediction field of soil organic matter at arbitrary resolution through implicit neural representation field decoding.
[0082] This embodiment also provides a computer device applicable to the construction method of a three-dimensional prediction model of soil organic matter based on machine learning, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the construction method of a three-dimensional prediction model of soil organic matter based on machine learning as proposed in the above embodiment.
[0083] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0084] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0085] In summary, this invention employs a momentum contrastive learning architecture combining a spectral-spatial-depth triple enhancement strategy for self-supervised pre-training, and utilizes a regularized projection network based on differentiable physical constraints for structured modulation, resulting in a modulated structured geographic-depth correlation representation. Using this modulated representation as a condition, a probabilistic diffusion model with implicit neural representations as the decoder is used to inversely generate sparse profile labels, yielding a continuous 3D volumetric field of soil organic matter. The model is then trained and validated using a loss function combining noise scheduling weighting and volumetric rendering consistency, resulting in optimized model parameters. These optimized parameters are used to predict new regions, and the 3D predicted volumetric field of soil organic matter at arbitrary resolution is obtained through decoding the implicit neural representation field. From multi-source data acquisition, feature fusion, self-supervised and physically constrained representation learning, to diffusion model-driven 3D continuum generation, an end-to-end machine learning solution is formed.
[0086] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for constructing a three-dimensional prediction model for soil organic matter based on machine learning, characterized in that: include, Multi-source data is collected and preprocessed to obtain a preprocessed multi-source aligned dataset. Feature extraction and fusion are performed through a multi-scale spatial pyramid network to obtain an initial fused feature cube. By performing self-supervised pre-training using a momentum contrastive learning architecture that combines spectral-spatial-depth triple enhancement strategies, and then performing structured modulation using a regularized projection network based on differentiable physical constraints, a modulated structured geographic-deep relational representation is obtained. Using the modulated structured geographic-deep association representation as a condition, a probability diffusion model with implicit neural representation as the decoder is used to inversely generate sparse profile labels, resulting in a continuous three-dimensional soil organic matter volume field. The model was trained and validated using a loss function that combines noise scheduling weighting and volume rendering consistency, resulting in optimized model parameters after training and validation. The optimized model parameters after training and validation are used to predict new regions, and the three-dimensional prediction field of soil organic matter at arbitrary resolution is obtained by decoding the implicit neural representation field.
2. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 1, characterized in that: The preprocessed multi-source aligned dataset includes, Multispectral images, hyperspectral images, and lidar point clouds are acquired to obtain raw multi-source acquisition data, which are then corrected to obtain a preprocessed multi-source aligned dataset.
3. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 2, characterized in that: The initially fused feature cube includes, A multi-scale spatial pyramid network is constructed, and features are interpreted through heterogeneous receptive field paths and cross-path gating mechanisms. Then, hierarchical encoding and cross-scale interaction are performed through bidirectional paths within the feature pyramid to obtain the multi-scale feature set after interaction. Non-parametric feature recombination based on mutual information and conditional entropy among multiple sources is performed on the multi-scale feature set after interaction. By maximizing inter-group mutual information and minimizing intra-group conditional entropy, feature self-organization fusion is driven to obtain the recombined multi-scale features. The recombined multi-scale features and the preprocessed multi-source aligned dataset are fused through residual connections to obtain the initial fused feature cube.
4. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 3, characterized in that: The self-supervised pre-training includes, A spectral-spatial-depth triple enhancement strategy is applied to the initially fused feature cube. The spectral-spatial-depth triple enhancement strategy performs adaptive probabilistic masking on specific organic matter-sensitive bands in the spectral dimension, performs clipping within the boundaries of terrain units derived from the digital elevation model in the spatial dimension, and performs random masking and sequence rearrangement based on geostatistical laws in the depth dimension to generate feature view pairs after triple enhancement transformation. The feature view after triple enhancement transformation is input to the momentum contrast learning architecture. Features are extracted by an online encoder and a target encoder that updates momentum. A set of positive and negative sample pairs is constructed based on the proximity constraint of spatial location and the continuity constraint of depth sequence. The improved InfoNCE contrast loss, based on the set of positive and negative samples, adjusts the discriminative power of positive and negative samples through a temperature coefficient and drives the encoder to learn a representation that is invariant to spectral interference, spatial terrain transformation and depth sequence perturbation, thus obtaining a contrast loss value for parameter update. The online encoder parameters are optimized by backpropagation using contrastive loss values and the target encoder parameters are updated by momentum. The trained encoder processes the initial fused feature cube to obtain a geographic-deep association representation.
5. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 4, characterized in that: The modulated structured geographic-deep association representation includes, Differentiable physical constraints are used to process geographic-deep association representations. By using the differentiable form of energy constraints on soil aggregate formation dynamics and organic matter-mineral complex stabilization processes, specific soil structure formation and carbon stabilization processes are encoded as constraint operators, resulting in a representation of embedded process constraints. The representation of the embedded process constraint is subjected to process-guided feature transformation. The representation is mapped and regularized by differentiable aggregate dynamics and energy operators to obtain the transformed representation. The physical constraint loss function is calculated based on the transformed representation. The physical constraint loss function includes the aggregate formation constraint loss and the stabilization energy constraint loss, and the physical constraint loss value is obtained. A regularized projection network with differentiable physical constraints is combined with physical constraint loss and representation reconstruction loss to jointly optimize network parameters, resulting in a modulated structured geographic-depth association representation that satisfies physical constraints.
6. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 5, characterized in that: The continuous three-dimensional soil organic matter field includes... The probabilistic diffusion model, which uses implicit neural representation as the decoder, takes the modulated structured geographic-deep association representation as the global conditional vector. Using sparse profile labels as local anchor constraints, the modulated structured geographic-depth association representation is fused with sparse profile labels through a cross-scale conditional injection mechanism to initialize the weights of the conditional denoising network. In each iteration of the inverse denoising process, the conditional denoising network couples the current noise latent variable, the denoising time step embedding, and the modulated structured geographic-deep association representation. It dynamically corrects the noise prediction direction through a physically constrained attention mechanism to obtain the denoised implicit neural representation field parameters for the current step. After completing all inverse denoising iterations, the probability diffusion model using implicit neural representation as the decoder obtains the implicit neural representation field parameters and defines a continuous mapping function from three-dimensional spatial coordinates to soil organic matter concentration. The implicit neural representation field parameters are parsed by a decoder composed of a multilayer sensing mechanism. It receives arbitrary query coordinates and combines them with the modulated structured geographic-depth association representation. The soil organic matter value at the query coordinates is obtained through forward propagation, and a continuous three-dimensional soil organic matter volume field is synthesized.
7. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 6, characterized in that: The optimized model parameters after training and validation include: By combining the noise scheduling weighting and the loss function of volume rendering consistency, the noise prediction error in the diffusion process is coupled with the depth profile error generated by the rendering of the three-dimensional volume field based on continuous soil organic matter, and a joint optimization loss for model optimization is constructed. The joint optimization loss drives parameter updates in the probabilistic diffusion model, which uses implicit neural representation as a decoder, through a backpropagation process. The accuracy and continuity of the model's generation of a continuous three-dimensional soil organic matter volume field are evaluated through a validation set. After iterative training and validation, the conditionalized denoising network and decoder with implicit neural representation as the decoder in the probability diffusion model with parameter optimization are saved, and the optimized model parameters after training and validation are obtained.
8. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 7, characterized in that: The prediction of the new region includes, The probabilistic diffusion model with implicit neural representation as decoder is loaded with optimized model parameters after training and validation. Feature extraction, self-supervised contrastive learning and regularized projection are performed on the preprocessed multi-source aligned dataset of the new region to generate a modulated structured geographic-deep association representation of the new region.
9. The method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in claim 8, characterized in that: The three-dimensional predictive field of soil organic matter includes... The probabilistic diffusion model, which uses implicit neural representation as the decoder, runs an inverse generation process conditioned on the modulated structured geographic-deep association representation of the new region, and outputs the implicit neural representation field parameters of the new region. The probabilistic diffusion model, which uses implicit neural representation as a decoder, queries the implicit neural representation field parameters of the new region through the decoder to generate a three-dimensional predictive volume field of soil organic matter in the new region.
10. A system for constructing a three-dimensional prediction model of soil organic matter based on machine learning, based on the method for constructing a three-dimensional prediction model of soil organic matter based on machine learning as described in any one of claims 1 to 9, characterized in that: This includes a fusion module that collects and preprocesses multi-source data to obtain a preprocessed multi-source aligned dataset, and performs feature extraction and fusion through a multi-scale spatial pyramid network to obtain an initial fused feature cube. The training module performs self-supervised pre-training using a momentum contrastive learning architecture that combines a spectral-spatial-depth triple enhancement strategy, and performs structured modulation using a regularized projection network based on differentiable physical constraints, resulting in a modulated structured geographic-deep relational representation. The reverse generation module, using the modulated structured geographic-deep association representation as a condition, utilizes a probability diffusion model with implicit neural representation as the decoder to reverse generate sparse profile labels, thereby obtaining a continuous three-dimensional soil organic matter volume field. The validation module uses a loss function that combines noise scheduling weighting and volume rendering consistency to train and validate the model, and obtains the optimized model parameters after training and validation. The prediction module uses the optimized model parameters after training and validation to predict new regions and obtains the three-dimensional prediction field of soil organic matter at arbitrary resolution through implicit neural representation field decoding.