A multi-modal probabilistic modeling subsurface structure inference system and method

The underground structure inference system based on multimodal probabilistic modeling solves the problems of unstable inference results and poor multimodal data fusion caused by data sparsity or lack of modalities in existing technologies. It realizes intelligent exploration and decision support in underground space, quantifies exploration risks and improves information utilization efficiency.

CN122492982APending Publication Date: 2026-07-31TIANJIN UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-04-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing geological exploration and 3D modeling technologies are highly dependent on data completeness. When data is sparse or modalities are missing, the reliability of inference results decreases, uncertainty cannot be quantified, multimodal data fusion is ineffective, and it is difficult to form a unified and reliable understanding of underground space.

Method used

The underground structure inference system employing multimodal probabilistic modeling includes modules for multimodal feature extraction, fusion, spatial topology graph construction, spatial context aggregation, and probabilistic space inference. It achieves adaptive model updates through a diagnostic self-evolution module and combines a deep probabilistic model for probability distribution modeling and inference.

Benefits of technology

It achieves stable inference under conditions of sparse data or missing modalities, quantifies exploration risks, provides intelligent exploration and decision support for underground space, and improves information utilization efficiency and reliable understanding of underground structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492982A_ABST
    Figure CN122492982A_ABST
Patent Text Reader

Abstract

A multimodal probabilistic modeling system and method for inferring underground structures are disclosed. The method first constructs a systematic error calibration mechanism based on standard sample data with known geological attributes, correcting and optimizing the consistency of multi-source heterogeneous data. Then, data from different sources are mapped to a unified latent space, and spatial coordinate conditions are introduced to construct a cross-modal feature selection mechanism, achieving unified feature expression and enhanced spatial correlation. Based on this, a deep probabilistic model based on variational inference learns the mapping relationship between features and the probability distribution of underground structures, generating a three-dimensional probability distribution field with quantified uncertainty and its three-dimensional representation. Finally, a diagnostic mechanism is constructed based on prediction errors, and adaptive model updates are achieved through selective parameter updates. This invention realizes the quantitative characterization and dynamic updating of underground structure uncertainty and can be applied to underground space exploration, pollution assessment, and sampling fidelity verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of environmental engineering and geological information technology, and in particular to a multimodal probabilistic modeling system and method for inferring underground structures. Background Technology

[0002] In fields such as contaminated site investigation, engineering geological survey, and environmental sampling, accurately acquiring and characterizing the three-dimensional subsurface structure is a fundamental prerequisite for implementing precise assessments and scientific decision-making. Currently, mainstream methods for reconstructing the three-dimensional subsurface structure mainly follow the following two technical paths: 1) Deterministic modeling methods based on physical measurement data; These methods rely on physical measurement data such as borehole data, geophysical exploration curves, drilling process parameters, and geological profiles. They construct deterministic three-dimensional stratigraphic models using traditional spatial statistical methods such as kriging interpolation, co-kriging, or multi-source data fusion. However, this method has significant limitations: (1) Its modeling accuracy is heavily dependent on the density and uniformity of sampling points, and its reliability drops sharply in sparse data areas; (2) The model output is a deterministic result, which cannot quantify the uncertainty of stratigraphic interfaces and lithological spatial distribution, resulting in inaccurate risk assessment.

[0003] 2) Microstructure reconstruction method based on high-precision scanning; For example, patent application number CN202410828799.7 proposes reconstructing the pore and component structure inside soil and rock samples using CT scans. Although this method can finely characterize microscopic features, it is only applicable to the laboratory scale and cannot be extrapolated to the macroscopic scale of the site. Furthermore, it relies on a complete sample acquisition and scanning process, which is costly and unsuitable for large-scale field investigations.

[0004] However, the core challenges faced in practical engineering lie precisely in the incompleteness of data, the heterogeneity of modes, and the uncertainty of geological bodies. Specifically, this manifests as follows: (1) There are huge differences in scale and semantics among multimodal data such as drilling vibration, acoustic signals, geophysical fields, and visual images; (2) Due to cost and site conditions, the data often appears sparse, discontinuous, or even missing some modalities; (3) The presence of underground hidden structures, such as faults, lenses, and contamination plumes, makes structural changes highly nonlinear and abrupt. Existing technologies are unable to learn the inherent generation laws of underground structures under such complex data conditions and output a three-dimensional geological scene with probability confidence that can express multiple reasonable possibilities.

[0005] Meanwhile, deep probabilistic models have demonstrated a powerful ability to learn underlying distributions from complex data and perform probabilistic generation in fields such as computer vision and natural language processing. However, to date, no mature technology has been developed to systematically introduce such deep probabilistic models into the field of geological engineering to solve the problem of probabilistic modeling and inference of underground space under multimodal and sparse data.

[0006] Therefore, the applicant proposes a method and system for inferring underground structures using multimodal probabilistic modeling. Summary of the Invention

[0007] The purpose of this invention is to address the following technical problems: Existing geological exploration and 3D modeling technologies are highly dependent on data completeness. While they generally perform well when data is dense and modalities are complete, the reliability and stability of their inference results drop sharply in actual explorations where data is sparse, some modalities are missing, or the signal-to-noise ratio is low. Furthermore, mainstream technologies are mostly based on deterministic models, which can only output a single "most likely" structure and cannot quantify the uncertainty of prediction results. This makes it difficult for decision-makers to assess exploration risks and to scientifically formulate supplementary exploration plans in areas with low confidence. In addition, when faced with multimodal data from diverse sources and with varying structures, such as drilling, geophysical exploration, imagery, and sensing, existing methods lack deep semantic alignment and robust fusion mechanisms. Especially when there are contradictions between data or mismatches in spatiotemporal scales, the fusion effect is often poor, information utilization efficiency is low, and it is difficult to form a unified and reliable understanding of underground space. To address the aforementioned technical problems, this invention proposes a multimodal probabilistic modeling technique for inferring underground structures. This technique transforms deterministic interpolation into probabilistic inference and possesses adaptive model updating capabilities. It can be used for intelligent exploration, risk assessment, and decision support in underground space.

[0008] To solve the above-mentioned technical problems, the present invention proposes the following technical solution: A multimodal probabilistic modeling system for inferring underground structures includes a multimodal feature extraction module. The output of the multimodal feature extraction module is used to output several feature vectors and query vectors to a multimodal feature fusion module. The outputs of the multimodal feature fusion module and the spatial topology graph construction module are connected to the input of the spatial context aggregation module. The output of the spatial context aggregation module is connected to the input of the probabilistic space inference module.

[0009] The diagnostic self-evolution module triggers error assessment and parameter update processes based on newly added measured data. It determines the subset of parameters to be updated through error diagnosis and optimizes and updates the parameter subset under the condition of meeting preset performance constraints, so as to achieve adaptive control of the multimodal feature extraction module, multimodal feature fusion module and probability space inference module.

[0010] It also includes a multimodal data input and classification module. The output of the multimodal data input and classification module is connected to the input of the multimodal feature extraction module and the spatial topology map construction module, respectively. The multimodal data input and classification module is used to receive multimodal in-situ collaborative sensing data of the target area and classify it according to the data structure to provide standardized input for feature encoders of different modalities.

[0011] It also includes a data calibration module, whose input is used to receive multimodal in-situ collaborative sensing data and corresponding measured values ​​of known geological attributes obtained based on standard samples. Its output is connected to the input of the underground structure visualization module, and the output of the probability space inference module is connected to the input of the underground structure visualization module. It also includes an optimizer and a diagnostic self-evolution module. The output of the probability space inference module is connected to the input of the optimizer and the diagnostic self-evolution module, respectively. The output of the optimizer is connected to the input of the multimodal feature extraction module and the multimodal feature fusion module, respectively.

[0012] The data calibration module receives multimodal in-situ collaborative sensing data and corresponding measured values ​​of known geological attributes obtained based on standard samples, wherein the multimodal data and the measured values ​​are in one-to-one correspondence. By inputting the multimodal data into a model composed of a multimodal feature extraction module, a multimodal feature fusion module, a spatial context aggregation module, and a probability space inference module, a predicted probability distribution is obtained. The statistically optimal estimate of the predicted probability distribution is compared with the measured values, and a calibration function is constructed using regression analysis or error mapping methods to achieve the pre-calibration of system errors, providing a calibration benchmark for subsequent underground structure inference and three-dimensional representation.

[0013] The spatial topology graph construction module constructs spatial adjacency relationships based on the spatial coordinate data of sampling points and generates a graph structure describing the spatial connection relationships between sampling points, providing a topological foundation for spatial context modeling. The multimodal feature extraction module has a built-in modality-adaptive feature encoder group, which extracts features from different modal data to obtain feature vector representations corresponding to each modality. The multimodal feature fusion module constructs a conditional query mechanism based on spatial coordinate information, selectively fuses features of each modality through weighted aggregation, and maps them to a shared latent space to obtain a unified multimodal fused feature vector. The spatial context aggregation module receives the multimodal fused feature vector and the spatial relationship graph, and uses a graph neural network to propagate and aggregate neighborhood information of node features to obtain an enhanced feature vector that integrates spatial context information. The probability space extrapolation module incorporates a deep probability model, which models and extrapolates the probability distribution of multimodal fusion features through an encoder and decoder. Preferably, the deep probability model is a variational generation model based on variational inference. After model deployment, the diagnostic self-evolution module performs error assessment and diagnosis based on newly added measured data, determines the parameter subset to be updated based on the diagnostic results, and calls the optimizer to optimize and update the parameter subset, thereby realizing the online adaptive update of the feature extraction model and the feature fusion model. Based on the probability space extrapolation results, a three-dimensional visualization of the underground structure is generated.

[0014] When the system is in operation / use, the following steps are included: Step S1: Based on standard samples with measured values ​​of known geological attributes, acquire multimodal in-situ collaborative sensing data; Step S2: Based on the data obtained in step S1, classify the data according to the data structure type, and obtain the feature vector representation of each modality through feature extraction; Step S3: Construct a conditional query mechanism and perform weighted aggregation based on the feature fusion model, map the fusion result to a unified latent space representation, and obtain a multimodal fusion feature vector; Step S4: Construct spatial adjacency relationships and generate a spatial relationship diagram describing the spatial connection relationships between sampling points; Step S5: Combine the multimodal fusion feature vector with the spatial relationship graph to perform spatial propagation and context modeling of node features; Step S6: Train the variational generation model and use the variational generation model to infer the target area, generating a three-dimensional geological attribute probability distribution field covering the entire area.

[0015] It also includes step S7: based on the three-dimensional geological attribute probability distribution field generated in step S6, the statistical characteristics and uncertainties of each grid point are calculated and mapped, and multi-scenario inference is performed under different sampling conditions through the variational generation model to realize the visual expression of the underground structure. It also includes step S8: constructing an error assessment and parameter update mechanism based on the newly added sampled data, determining the parameter subset that needs to be updated by diagnosing and analyzing the prediction error of the variational generation model; and using an optimizer to iteratively update the model parameter subset under the condition of satisfying the preset performance constraints, thereby achieving adaptive updating.

[0016] In step S1: Based on standard samples with measured values ​​of known geological attributes, multimodal in-situ collaborative sensing data is acquired, and a mapping relationship between the prediction results and the measured values ​​is constructed to achieve system error calibration; In step S2: acquire multimodal in-situ collaborative sensing data of the target area, classify and process the data according to the data structure type, extract features from different types of data through the feature encoding model in the multimodal feature extraction module, and obtain feature vector representations of each modality; In step S3, a conditional query mechanism is constructed using the spatial coordinates of the sampling points as conditional variables. Through the feature fusion model in the multimodal feature extraction module, weighted aggregation is performed based on the coordinate-guided cross-modal attention mechanism, and the fusion result is mapped to a unified latent space representation to obtain a multimodal fusion feature vector with spatial correlation. In step S4, spatial adjacency relationships are constructed based on the spatial coordinate information of the sampling points, and a spatial relationship diagram describing the spatial connection relationships between the sampling points is generated to characterize the spatial topology of geological attributes. In step S5, the multimodal fusion feature vector is combined with the spatial relationship graph. The graph neural network model in the spatial context modeling module performs spatial propagation and context modeling on the node features, so that the features of each node are fused with its neighborhood structure information, thereby obtaining an enhanced feature representation with spatial relevance. In step S6, the enhanced feature vector is input into the variational generation model in the probability space inference module for training. The model parameters are optimized by constructing reconstruction loss and regularization loss, so that the variational generation model learns the mapping relationship between enhanced features and the probability distribution of geological attributes. After training, the variational generation model is used to infer the target area and generate a three-dimensional geological attribute probability distribution field covering the entire area.

[0017] Step 1 includes the following sub-steps: Step 1-1: Construct a standard sample set with known geological attribute measured values, and arrange it in the target area according to the preset three-dimensional spatial coordinates; Steps 1-2: Perform multimodal collaborative acquisition on the standard samples to obtain drilling signals, geophysical data, image data, in-situ sensing data and experimental analysis data at the corresponding locations; Steps 1-3: Obtain the geological attribute prediction probability distribution for each standard sample location based on the multimodal data; Steps 1-4: Based on the deviation relationship between the predicted probability distribution and the measured values ​​of known geological attributes, a systematic error calibration function is constructed through regression analysis or error mapping methods to correct the prediction results output by the subsequent probability space extrapolation module.

[0018] Step 2 includes the following sub-steps: Step 2-1: Collect and input multimodal in-situ collaborative sensing data of the target area to construct a sensing data set. Multimodal data includes time-series signals acquired during drilling operations, geophysical exploration data, geological image data, in-situ sensor monitoring data, and geological attribute data of sampling points; spatial topology data is also acquired. , representing the three-dimensional spatial coordinates of each sampling point, the spatial topology data C Used for subsequent spatial relationship modeling; in some embodiments, structured data describing the spatial topological relationship between sampling points may be further introduced to supplement the spatial topological data; Step 2-2: Classify the multimodal data according to its representation and structural characteristics, dividing it into numerical datasets. Time series datasets and spatial structure data sets ; Steps 2-3: Input each type of data into the corresponding feature encoding model in the multimodal feature extraction module for feature extraction. The feature encoding model, depending on the data type, can employ one or more of the following: convolutional neural network, recurrent neural network, or attention-based network structure. For any type... Its eigenvectors are represented as: ; ; in, For any type eigenvectors, For the first t The total number of data instances in a modal dataset. For modality t The Middle j Feature vectors of data instances For the first t Feature coding model corresponding to class data For the first t The first in the class modal data set j One data instance, Its learnable parameter set; Steps 2-4: Output the feature vector set for each modality , , , which serves as the input to the subsequent multimodal feature fusion model, where: , ; It is a set of numerical modal eigenvectors. It is a set of temporal modal feature vectors. It is the set of spatial structural modal feature vectors.

[0019] Step 3 includes the following sub-steps: Step 3-1: Obtain the spatial topology data stored in Step 2-1: ; in, For the first i The three-dimensional spatial coordinates of each sampling point These represent their coordinate components in space. N This represents the total number of sampling points; Simultaneously, obtain the set of modal feature vectors output from steps 2-4. , , ; Step 3-2: Set the spatial coordinates of each sampling point Input to coordinate encoding model Mapped to query vector : ; in, For coordinate encoding functions, a multilayer perceptron (MLP) or a position encoding network can be used to map spatial location information to the feature space. Its learnable parameters; Step 3-3: For each sampling point and each mode The query vector obtained in step 3-2 As a conditional query, for modality t Feature vector set Perform weighted aggregation to obtain the results with respect to the sampling points. i Spatial location-dependent modal conditional feature vectors: ; in, For the first i Each sampling point in the mode t The conditional feature vectors after spatial condition filtering. For modal t Weighted summation of all candidate feature vectors in the dataset. For modality t The total number of candidate feature vectors For the first i The query vector of each sampling point The assigned weights, For modality t The Middle j 1 eigenvector; Attention weight The similarity between the query vector and the feature vector is calculated and normalized to obtain the following: ; in, For the first i The query vector corresponding to each sampling point superscript T This is the transpose operation for a matrix or vector. For modality t The corresponding learnable projection matrix is ​​used to project the feature vectors. Mapping to query vector A consistent similarity computation space d Let the similarity calculation space be the representation dimension. This is a scaling factor used to stabilize the calculation process of attention weights. It is an exponential function. k The index variable for summing the denominators. For modality t The Middle k 1 candidate feature vector; Its denominator: ; Indicates mode t The exponential terms corresponding to all candidate feature vectors are normalized and summed so that the sum of each attention weight is 1; The above method can be used for the same sampling point i Obtain its spatial conditional eigenvectors under different modalities , , This provides input for the fusion of multimodal features at the same point; Steps 3-4: Combine the same sampling points i Conditional eigenvectors under each modality , , The input is fed into the feature fusion model for joint mapping to obtain the multimodal fusion feature vector of the sampling point: , ; in, This is a feature fusion model, which can employ one or more of the following methods: weighted summation, vector concatenation followed by mapping through a fully connected network, or attention-based fusion structures. Its set of learnable parameters; Step 3-5: Perform steps 3-3 to 3-4 on all sampling points to output a set of multimodal fusion feature vectors. ; in, For the first iThe multimodal fusion feature vector corresponding to each sampling point, the feature vector set As input to the graph neural network model in the spatial context aggregation module; Step 4 includes the following sub-steps: Step 4-1: Obtain the spatial topology data input in Step 2-1 ,in, N This represents the total number of sampling points; Step 4-2: Based on coordinate set C This establishes the spatial adjacency relationship between sampling points. For any two points... and Their adjacency relationship is determined by the decision function: ; in, This is the adjacency determination function. Parameters for its determination; The adjacency relationship is based on the Euclidean distance threshold. K One or more of the nearest neighbor or density criteria are constructed to determine the spatial proximity between sampling points; when When, it represents a point. i With point j Adjacent in space; Step 4-3: Construct a spatial relationship graph based on adjacency relationships: ; Among them, node set V Corresponding sampling point set, node With coordinates One-to-one correspondence, edge set Determined by adjacency: ; in, For the spatial relationship diagram corresponding to the first i The target node of each sampling point For the target node in the spatial relationship diagram Adjacent nodes that have a spatial adjacency relationship; Step 4-4: Output spatial relationship diagram G As the graph structure input for subsequent graph neural network models, it is used to characterize the spatial dependencies between sampling points; Step 5 includes the following sub-steps: Step 5-1: Convert the multimodal fusion feature vector output from Step 3-5 into a multimodal fusion feature vector. As the initial features of the nodes, and the spatial relationship graph output from step 4-4 Align the input to the graph neural network model, where nodes With coordinates One-to-one correspondence, the initial node features are represented as ; Step 5-2: Create a spatial relationship diagram G The node features are input into the graph neural network model, and the node features are updated through multi-layer neighborhood information propagation. For the ... l layer ,node The feature update process includes neighborhood feature aggregation and node state update: Neighborhood aggregation: ; Feature update: ; in, For the graph neural network model in the first... l Layer aggregation functions, For the graph neural network model in the first... l The layer update function; the aggregation function can employ summation, averaging, or attention-based weighted aggregation methods, the update function... This can be achieved through a multilayer sensor or a gating mechanism; For nodes In the l Layer feature representation, Adjacent nodes In the l Feature vector representation in layered graph neural networks; For nodes In the l The neighborhood message (or context feature) vector obtained by aggregation in a layered graph neural network; Step 5-3: After L After layer information transmission, the final node features are obtained as spatial context enhancement features: ; in, For nodes After all L The final output feature vector after information transmission and feature update in the layered graph neural network; the enhanced feature integrates the node's own features and the features of its neighboring nodes to characterize the spatial correlation between sampling points. Step 5-4: Aggregate the enhanced features of all nodes to form an enhanced feature vector set: ; in, For the first i Enhanced feature vectors of each node, NThe total number of nodes is used to enhance the feature vector set as the input to the variational generation model in the subsequent probability space inference module; Step 6 includes the following sub-steps: Step 6-1: Combine the enhanced feature vector set output from Step 5-4 As the encoder that inputs observed data into the variational generative model, the encoder maps the enhanced features of each node into latent variables. z Posterior distribution parameters: ; ; in, For encoder parameters; Let be the mean and variance vectors of the posterior Gaussian distribution, respectively. For variance vector The diagonal covariance matrix is ​​constructed from the elements of the main diagonal. Latent variables The approximate posterior probability distribution is generated by the encoder based on the enhanced feature vector mapping; The symbol for normal distribution; Step 6-2: Sample latent variables from the posterior distribution using a reparameterization method: ; in, For random noise variables sampled from a standard normal distribution, this is used to achieve differentiability and parameter optimization of the sampling process through reparameterization techniques; For the Hadama product operator; For the identity matrix; implicit variables Spatial coordinates of the corresponding sampling points As conditional variables, these are input to the decoder of the variational generation model to obtain the probability distribution parameters of the geological attributes at that point: ; in, For decoder parameters, These are the predicted geological attribute distribution parameters; Step 6-3: Construct the loss function for the variational generation model and optimize the model parameters through backpropagation: ; in, , indicating the first i The reconstruction loss term for each sample, where Given latent variables and spatial coordinates, this is the conditional likelihood distribution of geological attributes (i.e., the predicted distribution generated by the decoder). , indicating the first iThe KL divergence term between the posterior and prior distributions of each sample, where For relative entropy, prior distribution , To balance the hyperparameters of the two terms; Step 6-4: After the variational generation model is trained, save the parameters of the variational generation model obtained from the training; for the coordinates of any grid point in the prediction region... q From the prior distribution Mid-sampled latent variables z and will z With coordinates q The common inputs are fed into the decoder of the variational generation model to obtain the probability distribution of geological attributes at that point: ; By traversing all grid points, a three-dimensional geological attribute probability distribution field covering the entire region is obtained: ; in M The total number of grid points. The target variable is the geological attribute of the underground structure to be inferred (such as pollutant concentration, lithology, or soil density). Step 7 includes the following sub-steps: Step 7-1: For the three-dimensional probability distribution field generated in step 6-4 At each grid point At that point, the predicted distribution output from the variational generation model. Extracting the statistically optimal estimate This constitutes a three-dimensional scalar field: ; The optimal estimate of the statistic is either the expected value or the maximum a posteriori estimate. This represents the statistically optimal estimate of the geological properties at the first grid point in three-dimensional space. For the third in three-dimensional space M The statistically optimal estimate of the geological properties at the last (i.e., the last) grid point; Three-dimensional scalar field Post-processing correction is performed using the calibration function determined in steps 1-4 to eliminate systematic errors; Based on the corrected scalar field, a three-dimensional representation model of the underground structure is generated using a three-dimensional geometric modeling method. Step 7-2: Based on the three-dimensional probability distribution field generated in Step 6-4 Calculate the uncertainty measure for each grid point to construct a three-dimensional uncertainty field: ; in, This represents the uncertainty measure of the geological attribute prediction results at the first grid point in three-dimensional space. For the third in three-dimensional space M The uncertainty measure of the geological attribute prediction results at the last (i.e., the last) grid point, for a continuous distribution, is calculated by determining its standard deviation. For discrete distributions, calculate their entropy. ,in For belonging to the first k The probability of a class; the uncertainty field U The data is mapped to color or transparency information and superimposed onto the 3D model obtained in step 7-1 to characterize the prediction confidence of different regions; where, K This represents the total number of categories for discrete geological attributes (such as soil and rock categories). k For the corresponding category index variable; To find the variance operator; Step 7-3: From the prior distribution of latent variables In the process S Independent sampling yields a set of latent variable samples. For each sample Compare it with the coordinates of all grid points The common input is fed into the decoder of the variational generation model to generate the corresponding three-dimensional probability distribution field. And extract the corresponding optimal estimated scalar field from each distribution field. This enables multi-scene three-dimensional visualization of underground structures; Step 8 includes the following sub-steps: Step 8-1: When the system acquires new, field-verified sampling point data, the diagnostic self-evolution module receives the new data and inputs it into the multimodal feature extraction module, the multimodal feature fusion module, the spatial context aggregation module, and the probability space inference module for prediction calculation to obtain the corresponding prediction results; Step 8-2: Calculate the point-level loss between the predicted probability distribution of the sampling points and their measured geological attribute values. ; Step 8-3: Calculate the point-level loss. With preset threshold T The comparison is performed, where the threshold is set based on a statistic of the historical validation point loss distribution: when If the model's prediction performance at that point is deemed satisfactory, the parameter update process is terminated. when When it is determined that the model parameters need to be updated, proceed to step 8-4; Step 8-4: Construct an error diagnosis mechanism based on the point-level error, determine the source of the error, and determine the range of model parameters to be updated accordingly; the error diagnosis is implemented through preset discrimination rules or a learning-based diagnostic model; Step 8-4 includes the following sub-steps: Step 8-4-1: Calculate the prediction error corresponding to each modality based on the single modality input. According to the error distribution characteristics of different modalities, determine that the error mainly comes from the feature encoding model or feature fusion model in the multimodal feature extraction module, thereby determining the corresponding parameter subset. Step 8-4-2: Calculate the difference measure between the latent variable distribution corresponding to the new sampling point and the historical latent variable distribution. When the difference exceeds the preset range, it is determined that there is a local feature distribution shift, thereby determining the relevant parameters of the variational generation model in the probability space inference module that need to be updated. Step 8-5: Based on the parameter range determined in Step 8-4, select the subset of parameters to be updated. And construct the optimization objective function: ; in, The weight coefficients for the regularization term are... For regularization, the remaining unselected parameters are frozen during the parameter update process. For parameters in the parameter subset that belong to the multimodal feature extraction module and the multimodal feature fusion module, gradient optimization is performed using the optimizer. For the variational generation model parameters in the probability space inference module, they are updated through the backpropagation mechanism. Step 8-6: Evaluate the performance of the models determined and updated in Steps 8-4 and 8-5 on an independent validation set. The models include the feature encoding model in the multimodal feature extraction module, the graph neural network model in the spatial context aggregation module, and / or the variational generation model in the probability space inference module. The validation set includes historical validation points and preset validation points. When the following conditions are met: (1) loss of new sampling points (2) If the average reconstruction loss of historical verification points is not higher than the preset range, the updated model is accepted as the current model; otherwise, the original model is retained and the anomaly is recorded.

[0020] In addition, to address the following technical problems in existing underground structure inference technologies: 1) Difficulty in heterogeneous modal coordination: Existing methods struggle to deeply mine cross-modal collaborative correlations of multi-source sensing data (such as drilling and geophysical exploration) within a unified feature space, leading to limited inference accuracy in complex geological conditions; 2) Spatial topology fragmentation: Traditional point-by-point independent prediction methods ignore neighborhood dependencies between sampling points, resulting in spatially uneven three-dimensional attribute fields that are prone to "island effects" that violate the continuity of real geology; 3) Difficulty in quantifying inference risk: Existing deterministic models, when faced with extremely sparse measured samples, can only output a single predicted value and cannot quantitatively assess the uncertainty of inference results for unknown areas. This invention also provides a multimodal probabilistic modeling method for underground structure inference, comprising the following steps: Step S1: Based on standard samples with measured values ​​of known geological attributes, acquire multimodal in-situ collaborative sensing data; Step S2: Based on the data obtained in step S1, classify the data according to the data structure type, and obtain the feature vector representation of each modality through feature extraction; Step S3: Construct a conditional query mechanism and perform weighted aggregation based on the feature fusion model, map the fusion result to a unified latent space representation, and obtain a multimodal fusion feature vector; Step S4: Construct spatial adjacency relationships and generate a spatial relationship diagram describing the spatial connection relationships between sampling points; Step S5: Combine the multimodal fusion feature vector with the spatial relationship graph to perform spatial propagation and context modeling of node features; Step S6: Train the variational generation model and use the variational generation model to infer the target area, generating a three-dimensional geological attribute probability distribution field covering the entire area.

[0021] It also includes step S7: based on the three-dimensional geological attribute probability distribution field generated in step S6, the statistical characteristics and uncertainties of each grid point are calculated and mapped, and multi-scenario inference is performed under different sampling conditions through the variational generation model to realize the visual expression of the underground structure. It also includes step S8: constructing an error assessment and parameter update mechanism based on the newly added sampled data, determining the parameter subset that needs to be updated by diagnosing and analyzing the prediction error of the variational generation model, and iteratively updating the parameter subset under the condition of satisfying the preset performance constraints.

[0022] In step S1: Based on standard samples with measured values ​​of known geological attributes, multimodal in-situ collaborative sensing data is acquired, and a mapping relationship between the prediction results and the measured values ​​is constructed to achieve error calibration; In step S2: acquire multimodal in-situ collaborative sensing data of the target area, classify and process the data according to the data structure type, extract features from different types of data, and obtain feature vector representations of each modality; In step S3, a conditional query mechanism is constructed using the spatial coordinates of the sampling points as conditional variables. The feature fusion model is used to perform weighted aggregation based on a coordinate-guided cross-modal attention mechanism, and the fusion result is mapped to a unified latent space representation to obtain a multimodal fusion feature vector with spatial correlation. In step S4, spatial adjacency relationships are constructed based on the spatial coordinate information of the sampling points, and a spatial relationship diagram describing the spatial connection relationships between the sampling points is generated to characterize the spatial topology of geological attributes. In step S5, the multimodal fusion feature vector is combined with the spatial relationship graph to perform spatial propagation and context modeling of node features, so that the features of each node are fused with its neighborhood structure information, thereby obtaining an enhanced feature representation with spatial relevance. In step S6, the enhanced feature vector is input into the variational generation model for training. The model parameters are optimized by constructing reconstruction loss and regularization loss, so that the variational generation model learns the mapping relationship between the enhanced features and the probability distribution of geological attributes. After training, the variational generation model is used to infer the target area and generate a three-dimensional geological attribute probability distribution field covering the entire area.

[0023] Step 1 includes the following sub-steps: Step 1-1: Construct a standard sample set with known geological attribute measured values, and arrange it in the target area according to the preset three-dimensional spatial coordinates; Steps 1-2: Perform multimodal collaborative acquisition on standard samples to obtain drilling signals, geophysical data, image data, in-situ sensing data, and experimental analysis data at the corresponding locations; Steps 1-3: Obtain the probability distribution of geological attribute predictions for each standard sample location based on multimodal data; Steps 1-4: Based on the deviation relationship between the predicted probability distribution and the measured values ​​of known geological attributes, an error calibration function is constructed through regression analysis or error mapping methods to correct the prediction results output by the subsequent probability space extrapolation module.

[0024] Step 2 includes the following sub-steps: Step 2-1: Collect and input multimodal in-situ collaborative sensing data of the target area to construct a sensing data set. Multimodal data includes time-series signals acquired during drilling operations, geophysical exploration data, geological image data, in-situ sensor monitoring data, and geological attribute data of sampling points; spatial topology data is also acquired. This represents the three-dimensional spatial coordinates of each sampling point, and the spatial topology data. C Used for subsequent spatial relationship modeling; in some embodiments, structured data describing the spatial topological relationship between sampling points may be further introduced to supplement the spatial topological data; Step 2-2: Classify the multimodal data according to its representation and structural characteristics, dividing it into numerical datasets. Time series datasets and spatial structure data sets ; Steps 2-3: Extract features from various types of data. The feature encoding model can be one or more of the following, depending on the data type: convolutional neural network, recurrent neural network, or attention-based network structure. For any type... Its eigenvectors are represented as: ; ; in, For any type eigenvectors, For the first t The total number of data instances in a modal dataset. For the first j Feature vectors of data instances For the first t Feature coding model corresponding to class data For the first t The first in the class modal data set j One data instance, Its learnable parameter set; Steps 2-4: Output the feature vector set for each modality , , , which serves as the input to the subsequent multimodal feature fusion model, where: , ; It is a set of numerical modal eigenvectors. It is a set of temporal modal feature vectors. It is the set of spatial structural modal feature vectors.

[0025] Step 3 includes the following sub-steps: Step 3-1: Obtain the spatial topology data stored in Step 2-1: ; in, For the first i The three-dimensional spatial coordinates of each sampling point These represent their coordinate components in space. N This represents the total number of sampling points; Simultaneously, obtain the set of modal feature vectors output from steps 2-4. , , ; Step 3-2: Set the spatial coordinates of each sampling point Input to coordinate encoding model Mapped to query vector : ; in, For coordinate encoding functions, a multilayer perceptron (MLP) or a position encoding network can be used to map spatial location information to the feature space. Its learnable parameters; Step 3-3: For each sampling point and each mode The query vector obtained in step 3-2 As a conditional query, for modality t Feature vector set Perform weighted aggregation to obtain the results with respect to the sampling points. i Spatial location-dependent modal conditional feature vectors: ; in, For the first i Each sampling point in the mode t The conditional feature vectors after spatial condition filtering. For modal t Weighted summation of all candidate feature vectors in the dataset. For modality t The total number of candidate feature vectors For modality t The Middle j 1 eigenvector For the first i The query vector of each sampling point The assigned weights; Attention weight The similarity between the query vector and the feature vector is calculated and normalized to obtain the following: ; in, For the first i The query vector corresponding to each sampling point superscript T This is the transpose operation for a matrix or vector. For modalityt The corresponding learnable projection matrix is ​​used to project the feature vectors. Mapping to query vector A consistent similarity computation space d Let the similarity calculation space be the representation dimension. This is a scaling factor used to stabilize the calculation process of attention weights. It is an exponential function. k The index variable for summing the denominators. For modality t The Middle k 1 candidate feature vector; Its denominator: ; Indicates mode t The exponential terms corresponding to all candidate feature vectors are normalized and summed so that the sum of each attention weight is 1; The above method can be used for the same sampling point i Obtain its spatial conditional eigenvectors under different modalities , , This provides input for the fusion of multimodal features at the same point; Steps 3-4: Combine the same sampling points i Conditional eigenvectors under each modality , , The input is fed into the feature fusion model for joint mapping to obtain the multimodal fusion feature vector of the sampling point: , ; in, This is a feature fusion model, which can employ one or more of the following methods: weighted summation, vector concatenation followed by mapping through a fully connected network, or attention-based fusion structures. Its set of learnable parameters; Step 3-5: Perform steps 3-3 to 3-4 on all sampling points to output a set of multimodal fusion feature vectors. ; in, For the first i The multimodal fusion feature vector corresponding to each sampling point, the feature vector set As input to the graph neural network model in the spatial context aggregation module; Step 4 includes the following sub-steps: Step 4-1: Obtain the spatial topology data input in Step 2-1 ,in, N This represents the total number of sampling points; Step 4-2: Based on coordinate set C This establishes the spatial adjacency relationship between sampling points. For any two points... and Their adjacency relationship is determined by the decision function: ; in, This is the adjacency determination function. Parameters for its determination; The adjacency relationship is based on the Euclidean distance threshold. K One or more of the nearest neighbor or density criteria are constructed to determine the spatial proximity between sampling points; when When, it represents a point. i With point j Adjacent in space; Step 4-3: Construct a spatial relationship graph based on adjacency relationships: ; Among them, node set V Corresponding sampling point set, node With coordinates One-to-one correspondence, edge set Determined by adjacency: ; in, For the spatial relationship diagram corresponding to the first i The target node of each sampling point For the target node in the spatial relationship diagram Adjacent nodes that have a spatial adjacency relationship; Step 4-4: Output spatial relationship diagram G As the graph structure input for subsequent graph neural network models, it is used to characterize the spatial dependencies between sampling points; Step 5 includes the following sub-steps: Step 5-1: Convert the multimodal fusion feature vector output from Step 3-5 into a multimodal fusion feature vector. As the initial features of the nodes, and the spatial relationship graph output from step 4-4 Align the input to the graph neural network model, where nodes With coordinates One-to-one correspondence, the initial node features are represented as ; Step 5-2: Create a spatial relationship diagram G The node features are input into the graph neural network model, and the node features are updated through multi-layer neighborhood information propagation. For the ... l layer ,node The feature update process includes neighborhood feature aggregation and node state update: Neighborhood aggregation: ; Feature update: ; in, For the graph neural network model in the first... l Layer aggregation functions, For the graph neural network model in the first... l The layer update function; the aggregation function can employ summation, averaging, or attention-based weighted aggregation methods, the update function... This can be achieved through a multilayer sensor or a gating mechanism; For nodes In the l Layer feature representation, Adjacent nodes In the l Feature vector representation in layered graph neural networks; For nodes In the l The neighborhood message (or context feature) vector obtained by aggregation in a layered graph neural network; Step 5-3: After L After layer information transmission, the final node features are obtained as spatial context enhancement features: ; in, For nodes After all L The final output feature vector after information transmission and feature update in the layered graph neural network; the enhanced feature integrates the node's own features and the features of its neighboring nodes to characterize the spatial correlation between sampling points. Step 5-4: Aggregate the enhanced features of all nodes to form an enhanced feature vector set: ; in, For the first i Enhanced feature vectors of each node, N The total number of nodes is used to enhance the feature vector set as the input to the variational generation model in the subsequent probability space inference module; Step 6 includes the following sub-steps: Step 6-1: Combine the enhanced feature vector set output from Step 5-4 As the encoder that inputs observed data into the variational generative model, the encoder maps the enhanced features of each node into latent variables. zPosterior distribution parameters: ; ; in, For encoder parameters; Let be the mean and variance vectors of the posterior Gaussian distribution, respectively. For variance vector The diagonal covariance matrix is ​​constructed from the elements of the main diagonal. Latent variables The approximate posterior probability distribution is generated by the encoder based on the enhanced feature vector mapping; The symbol for normal distribution; Step 6-2: Sample latent variables from the posterior distribution using a reparameterization method: ; in, For random noise variables sampled from a standard normal distribution, this is used to achieve differentiability and parameter optimization of the sampling process through reparameterization techniques; For the Hadama product operator; For the identity matrix; implicit variables Spatial coordinates of the corresponding sampling points As conditional variables, these are input to the decoder of the variational generation model to obtain the probability distribution parameters of the geological attributes at that point: ; in, For decoder parameters, These are the predicted geological attribute distribution parameters; Step 6-3: Construct the loss function for the variational generation model and optimize the model parameters through backpropagation: ; in, , indicating the first i The reconstruction loss term for each sample, where Given latent variables and spatial coordinates, this is the conditional likelihood distribution of geological attributes (i.e., the predicted distribution generated by the decoder). , indicating the first i The KL divergence term between the posterior and prior distributions of each sample, where For relative entropy, prior distribution , To balance the hyperparameters of the two terms; Step 6-4: After the variational generation model is trained, save the parameters of the variational generation model obtained from the training; for the coordinates of any grid point in the prediction region... q From the prior distribution Mid-sampled latent variables z and will z With coordinates q The common inputs are fed into the decoder of the variational generation model to obtain the probability distribution of geological attributes at that point: ; By traversing all grid points, a three-dimensional geological attribute probability distribution field covering the entire region is obtained: ; in M The total number of grid points. The target variable is the geological attribute of the underground structure to be inferred (such as pollutant concentration, lithology, or soil density). Step 7 includes the following sub-steps: Step 7-1: For the three-dimensional probability distribution field generated in step 6-4 At each grid point At that point, the predicted distribution output from the variational generation model. Extracting the statistically optimal estimate This constitutes a three-dimensional scalar field: ; The optimal estimate of the statistic is either the expected value or the maximum a posteriori estimate. This represents the statistically optimal estimate of the geological properties at the first grid point in three-dimensional space. For the third in three-dimensional space M The statistically optimal estimate of the geological properties at the last (i.e., the last) grid point; Three-dimensional scalar field Post-processing correction is performed using the calibration function determined in steps 1-4 to eliminate systematic errors; Based on the corrected scalar field, a three-dimensional representation model of the underground structure is generated using a three-dimensional geometric modeling method. Step 7-2: Based on the three-dimensional probability distribution field generated in Step 6-4 Calculate the uncertainty measure for each grid point to construct a three-dimensional uncertainty field: ; in, This represents the uncertainty measure of the geological attribute prediction results at the first grid point in three-dimensional space. For the third in three-dimensional space M The uncertainty measure of the geological attribute prediction results at the last (i.e., the last) grid point, for a continuous distribution, is calculated by determining its standard deviation. For discrete distributions, calculate their entropy. ,in For belonging to the firstk The probability of a class; the uncertainty field U The data is mapped to color or transparency information and superimposed onto the 3D model obtained in step 7-1 to characterize the prediction confidence of different regions; where, K This represents the total number of categories for discrete geological attributes (such as soil and rock categories). k For the corresponding category index variable; To find the variance operator; Step 7-3: From the prior distribution of latent variables In the process S Independent sampling yields a set of latent variable samples. For each sample Compare it with the coordinates of all grid points The common input is fed into the decoder of the variational generation model to generate the corresponding three-dimensional probability distribution field. And extract the corresponding optimal estimated scalar field from each distribution field. This enables multi-scene three-dimensional visualization of underground structures; Step 8 includes the following sub-steps: Step 8-1: When the system acquires new, field-verified sampling point data, the diagnostic self-evolution module receives the new data and inputs it into the multimodal feature extraction module, the multimodal feature fusion module, the spatial context aggregation module, and the probability space inference module for prediction calculation to obtain the corresponding prediction results; Step 8-2: Calculate the point-level loss between the predicted probability distribution of the sampling points and their measured geological attribute values. ; Step 8-3: Calculate the point-level loss. With preset threshold T The comparison is performed, where the threshold is set based on a statistic of the historical validation point loss distribution: when If the model's prediction performance at that point is deemed satisfactory, the parameter update process is terminated. when When it is determined that the model parameters need to be updated, proceed to step 8-4; Step 8-4: Construct an error diagnosis mechanism based on the point-level error, determine the source of the error, and determine the range of model parameters to be updated accordingly; the error diagnosis is implemented through preset discrimination rules or a learning-based diagnostic model; Step 8-4 includes the following sub-steps: Step 8-4-1: Calculate the prediction error corresponding to each modality based on the single modality input. According to the error distribution characteristics of different modalities, determine that the error mainly comes from the feature encoding model or feature fusion model in the multimodal feature extraction module, thereby determining the corresponding parameter subset. Step 8-4-2: Calculate the difference measure between the latent variable distribution corresponding to the new sampling point and the historical latent variable distribution. When the difference exceeds the preset range, it is determined that there is a local feature distribution shift, thereby determining the relevant parameters of the variational generation model in the probability space inference module that need to be updated. Step 8-5: Based on the parameter range determined in Step 8-4, select the subset of parameters to be updated. And construct the optimization objective function: ; in, The weight coefficients for the regularization term are... For regularization, the remaining unselected parameters are frozen during the parameter update process. For parameters in the parameter subset that belong to the multimodal feature extraction module and the multimodal feature fusion module, gradient optimization is performed using the optimizer. For the variational generation model parameters in the probability space inference module, they are updated through the backpropagation mechanism. Step 8-6: Evaluate the performance of the models determined and updated in Steps 8-4 and 8-5 on an independent validation set. The models include the feature encoding model in the multimodal feature extraction module, the graph neural network model in the spatial context aggregation module, and / or the variational generation model in the probability space inference module. The validation set includes historical validation points and preset validation points. When the following conditions are met: (1) loss of new sampling points (2) If the average reconstruction loss of historical verification points is not higher than the preset range, the updated model is accepted as the current model; otherwise, the original model is retained and the anomaly is recorded.

[0026] Compared with the prior art, the present invention has the following technical effects: 1) This invention breaks through the strong dependence of traditional methods on data completeness. Through a deep probabilistic framework and a cross-modal robust fusion mechanism, it can still make reasonable inferences about underground structures even when key modes are missing or data is sparse, which significantly improves the applicability and reliability of the technology in real exploration scenarios. 2) It realizes a paradigm shift from "single deterministic output" to "multiple probabilistic scenarios". The system can not only generate multiple geologically reasonable three-dimensional structural models to express the non-uniqueness of interpretation, but also provide quantitative confidence intervals for geological attributes, thereby transforming uncertainty from qualitative description into quantifiable and mappable decision-making basis; 3) It reveals the underlying generative patterns of the data rather than surface interpolation. The generative model of this invention can learn the intrinsic patterns of the formation and distribution of underground structures from multi-source heterogeneous data. Its inference process is based on the deduction of patterns in the latent space, rather than simple mathematical interpolation of known points. Therefore, it has stronger geological rationality and predictive ability for unsampled areas. 4) A complete technical closed loop of deep integration and continuous evolution has been constructed. Through multimodal alignment and fusion, the system effectively improves the comprehensive utilization efficiency and inference accuracy of heterogeneous data; furthermore, its embedded diagnostic self-evolution module enables the model to continuously optimize itself using newly added validation data, thereby accumulating site-specific knowledge in long-term application and forming a virtuous cycle of becoming more and more refined with use; 5) Establish the core capability of an intelligent cognitive system for underground space, namely, to perform probabilistic three-dimensional inference with quantifiable risks, interpretable results, and adaptive evolution capabilities in an incomplete, multimodal real data environment, providing a more scientific and reliable decision-making basis for geological exploration, environmental assessment and engineering safety. Attached Figure Description

[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a block diagram of the overall system structure in this invention; Figure 2 This is a flowchart of step 1 of the present invention; Figure 3 This is a flowchart of steps 2 to 5 of the present invention; Figure 4 This is a flowchart of step 6 of the present invention; Figure 5 This is a flowchart of step 8 of the present invention. Detailed Implementation

[0028] like Figure 1 As shown, a multimodal probabilistic modeling underground structure inference system includes a multimodal feature extraction module 6. The output of the multimodal feature extraction module 6 is used to output several feature vectors and query vectors to the multimodal feature fusion module 7. The outputs of the multimodal feature fusion module 7 and the spatial topology graph construction module 9 are connected to the input of the spatial context aggregation module 10. The output of the spatial context aggregation module 10 is connected to the input of the probabilistic space inference module 3.

[0029] The diagnostic self-evolution module 5 triggers the error assessment and parameter update process based on newly added measured data. It determines the subset of parameters to be updated through error diagnosis and optimizes and updates the parameter subset under the condition of meeting preset performance constraints, so as to realize the adaptive control of the multimodal feature extraction module, the multimodal feature fusion module and the probability space inference module.

[0030] It also includes a multimodal data input and classification module 8. The output of the multimodal data input and classification module 8 is connected to the input of the multimodal feature extraction module 6 and the spatial topology map construction module 9, respectively. The multimodal data input and classification module 8 is used to receive multimodal in-situ collaborative sensing data of the target area and perform classification processing according to the data structure to provide standardized input for feature encoders of different modalities.

[0031] It also includes a data calibration module 1, whose input is used to receive multimodal in-situ collaborative sensing data and corresponding measured values ​​of known geological attributes obtained based on standard samples. Its output is connected to the input of the underground structure visualization module 2, and the output of the probability space inference module 3 is connected to the input of the underground structure visualization module 2. It also includes an optimizer 4 and a diagnostic self-evolution module 5. The output of the probability space inference module 3 is connected to the input of the optimizer 4 and the diagnostic self-evolution module 5, respectively. The output of the optimizer 4 is connected to the input of the multimodal feature extraction module 6 and the multimodal feature fusion module 7, respectively.

[0032] The data calibration module 1 is used to receive multimodal in-situ collaborative sensing data and corresponding measured values ​​of known geological attributes obtained based on standard samples, wherein the multimodal data and the measured values ​​are in one-to-one correspondence; by inputting the multimodal data into a model composed of multimodal feature extraction module 6, multimodal feature fusion module 7, spatial context aggregation module 10 and probability space inference module 3, the predicted probability distribution is obtained, and the statistical optimal estimate of the predicted probability distribution is compared with the measured value. A calibration function is constructed using regression analysis or error mapping method to realize the pre-calibration of system error, and to provide a calibration benchmark for subsequent underground structure inference and three-dimensional representation.

[0033] The spatial topology graph construction module 9 constructs spatial adjacency relationships based on the spatial coordinate data of sampling points and generates a graph structure describing the spatial connection relationships between sampling points, providing a topological foundation for spatial context modeling; the multimodal feature extraction module 6 has a built-in modality-adaptive feature encoder group, which extracts features from different modal data to obtain feature vector representations corresponding to each modality; the multimodal feature fusion module 7 constructs a conditional query mechanism based on spatial coordinate information, selectively fuses features of each modality through weighted aggregation, and maps them to a shared latent space to obtain a unified multimodal fusion feature vector; the spatial context aggregation module 10 receives the multimodal fusion feature vector and spatial relationship graph, and uses a graph neural network to propagate and aggregate neighborhood information of node features to obtain an enhanced feature vector that integrates spatial context information; The probability space extrapolation module 3 incorporates a deep probability model, which performs probability distribution modeling and spatial extrapolation of multimodal fusion features through encoders and decoders. Preferably, the deep probability model is a variational generation model based on variational inference. After model deployment, the diagnostic self-evolution module 5 performs error evaluation and diagnosis based on newly added measured data, determines the parameter subset to be updated based on the diagnostic results, and calls the optimizer 4 to optimize and update the parameter subset, thereby realizing the online adaptive update of the feature extraction model and the feature fusion model. Based on the probability space extrapolation results, a three-dimensional visualization representation of the underground structure is generated.

[0034] When the system is in operation / use, the following steps are included: Step S1: Based on standard samples with measured values ​​of known geological attributes, acquire multimodal in-situ collaborative sensing data; Step S2: Based on the data obtained in step S1, classify the data according to the data structure type, and obtain the feature vector representation of each modality through feature extraction; Step S3: Construct a conditional query mechanism and perform weighted aggregation based on the feature fusion model, map the fusion result to a unified latent space representation, and obtain a multimodal fusion feature vector; Step S4: Construct spatial adjacency relationships and generate a spatial relationship diagram describing the spatial connection relationships between sampling points; Step S5: Combine the multimodal fusion feature vector with the spatial relationship graph to perform spatial propagation and context modeling of node features; Step S6: Train the variational generation model and use the variational generation model to infer the target area, generating a three-dimensional geological attribute probability distribution field covering the entire area.

[0035] It also includes step S7: based on the three-dimensional geological attribute probability distribution field generated in step S6, the statistical characteristics and uncertainties of each grid point are calculated and mapped, and multi-scenario inference is performed under different sampling conditions through the variational generation model to realize the visual expression of the underground structure. It also includes step S8: constructing an error assessment and parameter update mechanism based on the newly added sampled data, determining the parameter subset that needs to be updated by diagnosing and analyzing the prediction error of the variational generation model; and using an optimizer to iteratively update the model parameter subset under the condition of satisfying the preset performance constraints, thereby achieving adaptive updating.

[0036] In step S1: Based on standard samples with measured values ​​of known geological attributes, multimodal in-situ collaborative sensing data is acquired, and a mapping relationship between the prediction results and the measured values ​​is constructed to achieve system error calibration; In step S2: acquire multimodal in-situ collaborative sensing data of the target area, classify and process the data according to the data structure type, extract features from different types of data through the feature encoding model in the multimodal feature extraction module, and obtain feature vector representations of each modality; In step S3, a conditional query mechanism is constructed using the spatial coordinates of the sampling points as conditional variables. Through the feature fusion model in the multimodal feature extraction module, weighted aggregation is performed based on the coordinate-guided cross-modal attention mechanism, and the fusion result is mapped to a unified latent space representation to obtain a multimodal fusion feature vector with spatial correlation. In step S4, spatial adjacency relationships are constructed based on the spatial coordinate information of the sampling points, and a spatial relationship diagram describing the spatial connection relationships between the sampling points is generated to characterize the spatial topology of geological attributes. In step S5, the multimodal fusion feature vector is combined with the spatial relationship graph. The graph neural network model in the spatial context modeling module performs spatial propagation and context modeling on the node features, so that the features of each node are fused with its neighborhood structure information, thereby obtaining an enhanced feature representation with spatial relevance. In step S6, the enhanced feature vector is input into the variational generation model in the probability space inference module for training. The model parameters are optimized by constructing reconstruction loss and regularization loss, so that the variational generation model learns the mapping relationship between enhanced features and the probability distribution of geological attributes. After training, the variational generation model is used to infer the target area and generate a three-dimensional geological attribute probability distribution field covering the entire area.

[0037] Step 1 includes the following sub-steps: Step 1-1: Construct a standard sample set with known geological attribute measured values, and arrange it in the target area according to the preset three-dimensional spatial coordinates; Steps 1-2: Perform multimodal collaborative acquisition on the standard samples to obtain drilling signals, geophysical data, image data, in-situ sensing data and experimental analysis data at the corresponding locations; Steps 1-3: Obtain the geological attribute prediction probability distribution for each standard sample location based on the multimodal data; Steps 1-4: Based on the deviation relationship between the predicted probability distribution and the measured values ​​of known geological attributes, a systematic error calibration function is constructed through regression analysis or error mapping methods to correct the prediction results output by the subsequent probability space extrapolation module.

[0038] Step 2 includes the following sub-steps: Step 2-1: Collect and input multimodal in-situ collaborative sensing data of the target area to construct a sensing data set. Multimodal data includes time-series signals acquired during drilling operations, geophysical exploration data, geological image data, in-situ sensor monitoring data, and geological attribute data of sampling points; spatial topology data is also acquired. , representing the three-dimensional spatial coordinates of each sampling point, the spatial topology data C Used for subsequent spatial relationship modeling; in some embodiments, structured data describing the spatial topological relationship between sampling points may be further introduced to supplement the spatial topological data; Step 2-2: Classify the multimodal data according to its representation and structural characteristics, dividing it into numerical datasets. Time series datasets and spatial structure data sets ; Steps 2-3: Input each type of data into the corresponding feature encoding model in the multimodal feature extraction module for feature extraction. The feature encoding model, depending on the data type, can employ one or more of the following: convolutional neural network, recurrent neural network, or attention-based network structure. For any type... Its eigenvectors are represented as: ; ; in, For any type eigenvectors, For the first t The total number of data instances in a modal dataset. For the first j Feature vectors of data instances For the first t Feature coding model corresponding to class data For the first t The first in the class modal data set j One data instance, Its learnable parameter set; Steps 2-4: Output the feature vector set for each modality , , , which serves as the input to the subsequent multimodal feature fusion model, where: , ; It is a set of numerical modal eigenvectors. It is a set of temporal modal feature vectors. It is the set of spatial structural modal feature vectors.

[0039] Step 3 includes the following sub-steps: Step 3-1: Obtain the spatial topology data stored in Step 2-1: ; in, For the first i The three-dimensional spatial coordinates of each sampling point These represent their coordinate components in space. N This represents the total number of sampling points; Simultaneously, obtain the set of modal feature vectors output from steps 2-4. , , ; Step 3-2: Set the spatial coordinates of each sampling point Input to coordinate encoding model Mapped to query vector : ; in, For coordinate encoding functions, multilayer perceptrons (MLPs) or position encoding networks can be used to map spatial location information to the feature space. Its learnable parameters; Step 3-3: For each sampling point and each mode The query vector obtained in step 3-2 As a conditional query, the cross-modal attention mechanism in the multimodal feature extraction module is used to analyze the modality. t Feature vector set Perform weighted aggregation to obtain the results with respect to the sampling points. i Spatial location-dependent modal conditional feature vectors: ; in, For the first i Each sampling point in the mode t The conditional feature vectors after spatial condition filtering. For modal t Weighted summation of all candidate feature vectors in the dataset. For modality t The total number of candidate feature vectors For modality t The Middle j 1 eigenvector For the first i The query vector of each sampling point The assigned weights; The attention weight The similarity between the query vector and the feature vector is calculated and normalized to obtain the following: ; in, For the first i The query vector corresponding to each sampling point superscript T This is the transpose operation for a matrix or vector. For modality t The corresponding learnable projection matrix is ​​used to project the feature vectors. Mapping to query vector A consistent similarity computation space d Let the similarity calculation space be the representation dimension. This is a scaling factor used to stabilize the calculation process of attention weights. It is an exponential function. k The index variable for summing the denominators. For modality t The Middle k 1 candidate feature vector; Its denominator: ; Indicates mode t The exponential terms corresponding to all candidate feature vectors are normalized and summed so that the sum of each attention weight is 1; The above method can be used for the same sampling point i Obtain its spatial conditional eigenvectors under different modalities , , This provides input for the fusion of multimodal features at the same point; Steps 3-4: Combine the same sampling points i Conditional eigenvectors under each modality , , The input is fed into the feature fusion model in the multimodal feature extraction module for joint mapping to obtain the multimodal fusion feature vector of the sampling point: , ; in, The feature fusion model can employ one or more of the following: weighted summation, vector concatenation followed by mapping via a fully connected network, or an attention-based fusion structure. Its set of learnable parameters; Step 3-5: Perform steps 3-3 to 3-4 on all sampling points to output a set of multimodal fusion feature vectors. ; in, For the first i The set of multimodal fusion feature vectors corresponding to each sampling point As input to the graph neural network model in the spatial context aggregation module; Step 4 includes the following sub-steps: Step 4-1: Obtain the spatial topology data input in Step 2-1 ,in, N This represents the total number of sampling points; Step 4-2: Based on the coordinate set C This establishes the spatial adjacency relationship between sampling points. For any two points... and Their adjacency relationship is determined by the decision function: ; in, This is the adjacency determination function. Parameters for its determination; The adjacency relationship is based on the Euclidean distance threshold. K One or more of the nearest neighbor or density criteria are constructed to determine the spatial proximity between sampling points; when When, it represents a point. i With point j Adjacent in space; Step 4-3: Construct a spatial relationship graph based on the adjacency relationships: ; Among them, node set V Corresponding sampling point set, node With coordinates One-to-one correspondence, edge set Determined by adjacency: ; in, For the spatial relationship diagram corresponding to the first i The target node of each sampling point For the target node in the spatial relationship diagram Adjacent nodes that have a spatial adjacency relationship; Step 4-4: Output spatial relationship diagram G As the graph structure input of the graph neural network model in the subsequent spatial context aggregation module, it is used to characterize the spatial dependencies between sampling points; Step 5 includes the following sub-steps: Step 5-1: Convert the multimodal fusion feature vector output from Step 3-5 into a multimodal fusion feature vector. As the initial features of the nodes, and the spatial relationship graph output from step 4-4 Align the input to the graph neural network model in the spatial context aggregation module, where nodes With coordinates One-to-one correspondence, the initial node features are represented as ; Step 5-2: Create a spatial relationship diagram G The node features are input into the graph neural network model, and the node features are updated through multi-layer neighborhood information propagation. For the first node... l layer ,node The feature update process includes neighborhood feature aggregation and node state update: Neighborhood aggregation: ; Feature update: ; in, For the graph neural network model in the first... Layer aggregation functions, For the graph neural network model in the first... The layer update function; the aggregation function can employ summation, averaging, or attention-based weighted aggregation methods, the update function This can be achieved through a multilayer sensor or a gating mechanism; For nodes In the l Layer feature representation, Adjacent nodes In the l Feature vector representation in layered graph neural networks; For nodes In the l The neighborhood message (or context feature) vector obtained by aggregation in a layered graph neural network; Step 5-3: After L After layer information transmission, the final node features are obtained as spatial context enhancement features: ; in, For nodes After all L The final output feature vector after information transmission and feature update in the layered graph neural network; the enhanced feature integrates the node's own features and the features of its neighboring nodes to characterize the spatial correlation between sampling points; Step 5-4: Aggregate the enhanced features of all nodes to form an enhanced feature vector set: ; in, For the firsti Enhanced feature vectors of each node, N The total number of nodes is represented by the enhanced feature vector set, which serves as the input to the variational generation model in the subsequent probability space inference module. Step 6 includes the following sub-steps: Step 6-1: Combine the enhanced feature vector set output from Step 5-4 The encoder, which inputs the observed data into the variational generative model in the probability space inference module, maps the enhanced features of each node into latent variables. z Posterior distribution parameters: ; ; in, For encoder parameters; Let be the mean and variance vectors of the posterior Gaussian distribution, respectively. For variance vector The diagonal covariance matrix is ​​constructed from the elements of the main diagonal. Latent variables The approximate posterior probability distribution is generated by the encoder based on the enhanced feature vector mapping; The symbol for normal distribution; Step 6-2: Sample latent variables from the posterior distribution using a reparameterization method: ; in, For random noise variables sampled from a standard normal distribution, this is used to achieve differentiability and parameter optimization of the sampling process through reparameterization techniques; For the Hadama product operator; For the identity matrix; implicit variables Spatial coordinates of the corresponding sampling points As conditional variables, these parameters are input together into the decoder of the variational generation model to obtain the probability distribution parameters of the geological attributes at that point. ; in, For decoder parameters, These are the predicted geological attribute distribution parameters; Step 6-3: Construct the loss function of the variational generation model and optimize the model parameters through backpropagation: ; in, , indicating the first i The reconstruction loss term for each sample, where Given latent variables and spatial coordinates, this is the conditional likelihood distribution of geological attributes (i.e., the predicted distribution generated by the decoder). , indicating the first i The KL divergence term between the posterior and prior distributions of each sample, where For relative entropy, prior distribution , To balance the hyperparameters of the two terms; Step 6-4: After the variational generation model is trained, save the parameters of the variational generation model obtained from the training; for the coordinates of any grid point in the prediction region... q From the prior distribution Mid-sampled latent variables z and will z With coordinates q The common inputs are fed into the decoder of the variational generation model to obtain the probability distribution of geological attributes at that point: ; By traversing all grid points, a three-dimensional geological attribute probability distribution field covering the entire region is obtained: ; in M The total number of grid points. The target variable is the geological attribute of the underground structure to be inferred (such as pollutant concentration, lithology, or soil density). Step 7 includes the following sub-steps: Step 7-1: For the three-dimensional probability distribution field generated in step 6-4 At each grid point At that point, the predicted distribution output from the variational generation model. Extracting the statistically optimal estimate This constitutes a three-dimensional scalar field: ; Wherein, the optimal estimate of the statistic is the expected value or the maximum a posteriori estimate. This represents the statistically optimal estimate of the geological properties at the first grid point in three-dimensional space. For the third in three-dimensional space M The statistically optimal estimate of the geological properties at the last (i.e., the last) grid point; The three-dimensional scalar field Post-processing correction is performed using the calibration function determined in steps 1-4 to eliminate systematic errors; Based on the corrected scalar field, a three-dimensional representation model of the underground structure is generated using a three-dimensional geometric modeling method. Step 7-2: Based on the three-dimensional probability distribution field generated in Step 6-4 Calculate the uncertainty measure for each grid point to construct a three-dimensional uncertainty field: ; in, This represents the uncertainty measure of the geological attribute prediction results at the first grid point in three-dimensional space. For the third in three-dimensional space M The uncertainty measure of the geological attribute prediction results at the last (i.e., the last) grid point, for a continuous distribution, is calculated by determining its standard deviation. For discrete distributions, calculate their entropy. ,in For belonging to the first k The probability of a class; the uncertainty field U The data is mapped to color or transparency information and superimposed onto the 3D model obtained in step 7-1 to characterize the prediction confidence of different regions; where, K This represents the total number of categories for discrete geological attributes (such as soil and rock categories). k For the corresponding category index variable; To find the variance operator; Step 7-3: From the prior distribution of latent variables In the process S Independent sampling yields a set of latent variable samples. For each sample Compare it with the coordinates of all grid points The common inputs are fed into the decoder of the variational generation model to generate the corresponding three-dimensional probability distribution field. And extract the corresponding optimal estimated scalar field from each distribution field. This enables multi-scene three-dimensional visualization of underground structures; Step 8 includes the following sub-steps: Step 8-1: When the system acquires new, field-verified sampling point data, the diagnostic self-evolution module receives the new data and inputs it into the multimodal feature extraction module, the multimodal feature fusion module, the spatial context aggregation module, and the probability space inference module for prediction calculation to obtain the corresponding prediction results; Step 8-2: Calculate the point-level loss between the predicted probability distribution of the sampling points and their measured geological attribute values. ; Step 8-3: Calculate the point-level loss. With preset threshold T The comparison is performed, where the threshold is set based on a statistic of the historical validation point loss distribution: when If the model's prediction performance at that point is deemed satisfactory, the parameter update process is terminated. when When it is determined that the model parameters need to be updated, proceed to step 8-4; Step 8-4: Construct an error diagnosis mechanism based on the point-level error, determine the source of the error, and determine the range of model parameters to be updated accordingly; the error diagnosis is implemented through preset discrimination rules or a learning-based diagnostic model; Step 8-4 includes the following sub-steps: Step 8-4-1: Calculate the prediction error corresponding to each modality based on the single modality input. According to the error distribution characteristics of different modalities, determine that the error mainly comes from the feature encoding model or feature fusion model in the multimodal feature extraction module, thereby determining the corresponding parameter subset. Step 8-4-2: Calculate the difference measure between the latent variable distribution corresponding to the new sampling point and the historical latent variable distribution. When the difference exceeds the preset range, it is determined that there is a local feature distribution shift, thereby determining the relevant parameters of the variational generation model in the probability space inference module that need to be updated. Step 8-5: Based on the parameter range determined in Step 8-4, select the subset of parameters to be updated. And construct the optimization objective function: ; in, The weight coefficients for the regularization term are... For regularization, the remaining unselected parameters are frozen during the parameter update process. For parameters in the parameter subset that belong to the multimodal feature extraction module and the multimodal feature fusion module, gradient optimization is performed using the optimizer. For the variational generation model parameters in the probability space inference module, they are updated through the backpropagation mechanism. Step 8-6: Evaluate the performance of the models determined and updated in Steps 8-4 and 8-5 on an independent validation set. The models include the feature encoding model in the multimodal feature extraction module, the graph neural network model in the spatial context aggregation module, and / or the variational generation model in the probability space inference module. The validation set includes historical validation points and preset validation points. When the following conditions are met: (1) loss of new sampling points (2) If the average reconstruction loss of historical verification points is not higher than the preset range, the updated model is accepted as the current model; otherwise, the original model is retained and the anomaly is recorded.

[0040] In addition, to address the following technical problems in existing underground structure inference technologies: 1) Difficulty in heterogeneous modal coordination: Existing methods struggle to deeply mine cross-modal collaborative correlations of multi-source sensing data (such as drilling and geophysical exploration) within a unified feature space, leading to limited inference accuracy in complex geological conditions; 2) Spatial topology fragmentation: Traditional point-by-point independent prediction methods ignore neighborhood dependencies between sampling points, resulting in spatially uneven three-dimensional attribute fields that are prone to "island effects" that violate the continuity of real geology; 3) Difficulty in quantifying inference risk: Existing deterministic models, when faced with extremely sparse measured samples, can only output a single predicted value and cannot quantitatively assess the uncertainty of inference results for unknown areas. This invention also provides a multimodal probabilistic modeling method for underground structure inference, comprising the following steps: Step S1: Based on standard samples with measured values ​​of known geological attributes, acquire multimodal in-situ collaborative sensing data; Step S2: Based on the data obtained in step S1, classify the data according to the data structure type, and obtain the feature vector representation of each modality through feature extraction; Step S3: Construct a conditional query mechanism and perform weighted aggregation based on the feature fusion model, map the fusion result to a unified latent space representation, and obtain a multimodal fusion feature vector; Step S4: Construct spatial adjacency relationships and generate a spatial relationship diagram describing the spatial connection relationships between sampling points; Step S5: Combine the multimodal fusion feature vector with the spatial relationship graph to perform spatial propagation and context modeling of node features; Step S6: Train the variational generation model and use the variational generation model to infer the target area, generating a three-dimensional geological attribute probability distribution field covering the entire area.

[0041] It also includes step S7: based on the three-dimensional geological attribute probability distribution field generated in step S6, the statistical characteristics and uncertainties of each grid point are calculated and mapped, and multi-scenario inference is performed under different sampling conditions through the variational generation model to realize the visual expression of the underground structure. It also includes step S8: constructing an error assessment and parameter update mechanism based on the newly added sampled data, determining the parameter subset that needs to be updated by diagnosing and analyzing the prediction error of the variational generation model, and iteratively updating the parameter subset under the condition of satisfying the preset performance constraints.

[0042] In step S1: Based on standard samples with measured values ​​of known geological attributes, multimodal in-situ collaborative sensing data is acquired, and a mapping relationship between the prediction results and the measured values ​​is constructed to achieve error calibration; In step S2: acquire multimodal in-situ collaborative sensing data of the target area, classify and process the data according to the data structure type, extract features from different types of data, and obtain feature vector representations of each modality; In step S3, a conditional query mechanism is constructed using the spatial coordinates of the sampling points as conditional variables. The feature fusion model is used to perform weighted aggregation based on a coordinate-guided cross-modal attention mechanism, and the fusion result is mapped to a unified latent space representation to obtain a multimodal fusion feature vector with spatial correlation. In step S4, spatial adjacency relationships are constructed based on the spatial coordinate information of the sampling points, and a spatial relationship diagram describing the spatial connection relationships between the sampling points is generated to characterize the spatial topology of geological attributes. In step S5, the multimodal fusion feature vector is combined with the spatial relationship graph to perform spatial propagation and context modeling of node features, so that the features of each node are fused with its neighborhood structure information, thereby obtaining an enhanced feature representation with spatial relevance. In step S6, the enhanced feature vector is input into the variational generation model for training. The model parameters are optimized by constructing reconstruction loss and regularization loss, so that the variational generation model learns the mapping relationship between the enhanced features and the probability distribution of geological attributes. After training, the variational generation model is used to infer the target area and generate a three-dimensional geological attribute probability distribution field covering the entire area.

[0043] Step 1 includes the following sub-steps: Step 1-1: Construct a standard sample set with known geological attribute measured values, and arrange it in the target area according to the preset three-dimensional spatial coordinates; Steps 1-2: Perform multimodal collaborative acquisition on standard samples to obtain drilling signals, geophysical data, image data, in-situ sensing data, and experimental analysis data at the corresponding locations; Steps 1-3: Obtain the probability distribution of geological attribute predictions for each standard sample location based on multimodal data; Steps 1-4: Based on the deviation relationship between the predicted probability distribution and the measured values ​​of known geological attributes, an error calibration function is constructed through regression analysis or error mapping methods to correct the prediction results output by the subsequent probability space extrapolation module.

[0044] Step 2 includes the following sub-steps: Step 2-1: Collect and input multimodal in-situ collaborative sensing data of the target area to construct a sensing data set. Multimodal data includes time-series signals acquired during drilling operations, geophysical exploration data, geological image data, in-situ sensor monitoring data, and geological attribute data of sampling points; spatial topology data is also acquired. This represents the three-dimensional spatial coordinates of each sampling point, and the spatial topology data. C Used for subsequent spatial relationship modeling; in some embodiments, structured data describing the spatial topological relationship between sampling points may be further introduced to supplement the spatial topological data; Step 2-2: Classify the multimodal data according to its representation and structural characteristics, dividing it into numerical datasets. Time series datasets and spatial structure data sets ; Steps 2-3: Extract features from various types of data. The feature encoding model can be one or more of the following, depending on the data type: convolutional neural network, recurrent neural network, or attention-based network structure. For any type... Its eigenvectors are represented as: ; ; in, For any type eigenvectors, For the first t The total number of data instances in a modal dataset. For the first j Feature vectors of data instances For the first t Feature coding model corresponding to class data For the first t The first in the class modal data set j One data instance, Its learnable parameter set; Steps 2-4: Output the feature vector set for each modality , , , which serves as the input to the subsequent multimodal feature fusion model, where: , ; It is a set of numerical modal eigenvectors. It is a set of temporal modal feature vectors. It is the set of spatial structural modal feature vectors.

[0045] Step 3 includes the following sub-steps: Step 3-1: Obtain the spatial topology data stored in Step 2-1: ; in, For the first i The three-dimensional spatial coordinates of each sampling point These represent their coordinate components in space. N This represents the total number of sampling points; Simultaneously, obtain the set of modal feature vectors output from steps 2-4. , , ; Step 3-2: Set the spatial coordinates of each sampling point Input to coordinate encoding model Mapped to query vector : ; in, For coordinate encoding functions, a multilayer perceptron (MLP) or a position encoding network can be used to map spatial location information to the feature space. Its learnable parameters; Step 3-3: For each sampling point and each mode The query vector obtained in step 3-2 As a conditional query, for modality t Feature vector set Perform weighted aggregation to obtain the results with respect to the sampling points. i Spatial location-dependent modal conditional feature vectors: ; in, For the first i Each sampling point in the mode t The conditional feature vectors after spatial condition filtering. For modal t Weighted summation of all candidate feature vectors in the dataset. For modality t The total number of candidate feature vectors For modality t The Middle j 1 eigenvector For the first i The query vector of each sampling point The assigned weights; Attention weight The similarity between the query vector and the feature vector is calculated and normalized to obtain the following: ; in, For the first i The query vector corresponding to each sampling point superscript T This is the transpose operation for a matrix or vector. For modalityt The corresponding learnable projection matrix is ​​used to project the feature vectors. Mapping to query vector A consistent similarity computation space d Let the similarity calculation space be the representation dimension. This is a scaling factor used to stabilize the calculation process of attention weights. It is an exponential function. k The index variable for summing the denominators. For modality t The Middle k 1 candidate feature vector; Its denominator: ; Indicates mode t The exponential terms corresponding to all candidate feature vectors are normalized and summed so that the sum of each attention weight is 1; The above method can be used for the same sampling point i Obtain its spatial conditional eigenvectors under different modalities , , This provides input for the fusion of multimodal features at the same point; Steps 3-4: Combine the same sampling points i Conditional eigenvectors under each modality , , The input is fed into the feature fusion model for joint mapping to obtain the multimodal fusion feature vector of the sampling point: , ; in, This is a feature fusion model, which can employ one or more of the following methods: weighted summation, vector concatenation followed by mapping through a fully connected network, or attention-based fusion structures. Its set of learnable parameters; Step 3-5: Perform steps 3-3 to 3-4 on all sampling points to output a set of multimodal fusion feature vectors. ; in, For the first i The multimodal fusion feature vector corresponding to each sampling point, the feature vector set As input to the graph neural network model in the spatial context aggregation module; Step 4 includes the following sub-steps: Step 4-1: Obtain the spatial topology data input in Step 2-1 ,in, N This represents the total number of sampling points; Step 4-2: Based on coordinate set C This establishes the spatial adjacency relationship between sampling points. For any two points... and Their adjacency relationship is determined by the decision function: ; in, This is the adjacency determination function. Parameters for its determination; The adjacency relationship is based on the Euclidean distance threshold. K One or more of the nearest neighbor or density criteria are constructed to determine the spatial proximity between sampling points; when When, it represents a point. i With point j Adjacent in space; Step 4-3: Construct a spatial relationship graph based on adjacency relationships: ; Among them, node set V Corresponding sampling point set, node With coordinates One-to-one correspondence, edge set Determined by adjacency: ; in, For the spatial relationship diagram corresponding to the first i The target node of each sampling point For the target node in the spatial relationship diagram Adjacent nodes that have a spatial adjacency relationship; Step 4-4: Output spatial relationship diagram G As the graph structure input for subsequent graph neural network models, it is used to characterize the spatial dependencies between sampling points; Step 5 includes the following sub-steps: Step 5-1: Convert the multimodal fusion feature vector output from Step 3-5 into a multimodal fusion feature vector. As the initial features of the nodes, and the spatial relationship graph output from step 4-4 Align the input to the graph neural network model, where nodes With coordinates One-to-one correspondence, the initial node features are represented as ; Step 5-2: Create a spatial relationship diagram G The node features are input into the graph neural network model, and the node features are updated through multi-layer neighborhood information propagation. For the ... l layer ,node The feature update process includes neighborhood feature aggregation and node state update: Neighborhood aggregation: ; Feature update: ; in, For the graph neural network model in the first... l Layer aggregation functions, For the graph neural network model in the first... l The layer update function; the aggregation function can employ summation, averaging, or attention-based weighted aggregation methods, the update function... This can be achieved through a multilayer sensor or a gating mechanism; For nodes In the l Layer feature representation, Adjacent nodes In the l Feature vector representation in layered graph neural networks; For nodes In the l The neighborhood message (or context feature) vector obtained by aggregation in a layered graph neural network; Step 5-3: After L After layer information transmission, the final node features are obtained as spatial context enhancement features: ; in, For nodes After all L The final output feature vector after information transmission and feature update in the layered graph neural network; the enhanced feature integrates the node's own features and the features of its neighboring nodes to characterize the spatial correlation between sampling points. Step 5-4: Aggregate the enhanced features of all nodes to form an enhanced feature vector set: ; in, For the first i Enhanced feature vectors of each node, N The total number of nodes is used to enhance the feature vector set as the input to the variational generation model in the subsequent probability space inference module; Step 6 includes the following sub-steps: Step 6-1: Combine the enhanced feature vector set output from Step 5-4 As the encoder that inputs observed data into the variational generative model, the encoder maps the enhanced features of each node into latent variables. zPosterior distribution parameters: ; ; in, For encoder parameters; Let be the mean and variance vectors of the posterior Gaussian distribution, respectively. For variance vector The diagonal covariance matrix is ​​constructed from the elements of the main diagonal. Latent variables The approximate posterior probability distribution is generated by the encoder based on the enhanced feature vector mapping; The symbol for normal distribution; Step 6-2: Sample latent variables from the posterior distribution using a reparameterization method: ; in, For random noise variables sampled from a standard normal distribution, this is used to achieve differentiability and parameter optimization of the sampling process through reparameterization techniques; For the Hadama product operator; For the identity matrix; implicit variables Spatial coordinates of the corresponding sampling points As conditional variables, these are input to the decoder of the variational generation model to obtain the probability distribution parameters of the geological attributes at that point: ; in, For decoder parameters, These are the predicted geological attribute distribution parameters; Step 6-3: Construct the loss function for the variational generation model and optimize the model parameters through backpropagation: ; in, , indicating the first i The reconstruction loss term for each sample, where Given latent variables and spatial coordinates, this is the conditional likelihood distribution of geological attributes (i.e., the predicted distribution generated by the decoder). , indicating the first i The KL divergence term between the posterior and prior distributions of each sample, where For relative entropy, prior distribution , To balance the hyperparameters of the two terms; Step 6-4: After the variational generation model is trained, save the parameters of the variational generation model obtained from the training; for the coordinates of any grid point in the prediction region... q From the prior distribution Mid-sampled latent variables z and will z With coordinates q The common inputs are fed into the decoder of the variational generation model to obtain the probability distribution of geological attributes at that point: ; By traversing all grid points, a three-dimensional geological attribute probability distribution field covering the entire region is obtained: ; in M The total number of grid points. The target variable is the geological attribute of the underground structure to be inferred (such as pollutant concentration, lithology, or soil density). Step 7 includes the following sub-steps: Step 7-1: For the three-dimensional probability distribution field generated in step 6-4 At each grid point At that point, the predicted distribution output from the variational generation model. Extracting the statistically optimal estimate This constitutes a three-dimensional scalar field: ; The optimal estimate of the statistic is either the expected value or the maximum a posteriori estimate. This represents the statistically optimal estimate of the geological properties at the first grid point in three-dimensional space. For the third in three-dimensional space M The statistically optimal estimate of the geological properties at the last (i.e., the last) grid point; Three-dimensional scalar field Post-processing correction is performed using the calibration function determined in steps 1-4 to eliminate systematic errors; Based on the corrected scalar field, a three-dimensional representation model of the underground structure is generated using a three-dimensional geometric modeling method. Step 7-2: Based on the three-dimensional probability distribution field generated in Step 6-4 Calculate the uncertainty measure for each grid point to construct a three-dimensional uncertainty field: ; in, This represents the uncertainty measure of the geological attribute prediction results at the first grid point in three-dimensional space. For the third in three-dimensional space M The uncertainty measure of the geological attribute prediction results at the last (i.e., the last) grid point, for a continuous distribution, is calculated by determining its standard deviation. For discrete distributions, calculate their entropy. ,in For belonging to the firstk The probability of a class; the uncertainty field U The data is mapped to color or transparency information and superimposed onto the 3D model obtained in step 7-1 to characterize the prediction confidence of different regions; where, K This represents the total number of categories for discrete geological attributes (such as soil and rock categories). k For the corresponding category index variable; To find the variance operator; Step 7-3: From the prior distribution of latent variables In the process S Independent sampling yields a set of latent variable samples. For each sample Compare it with the coordinates of all grid points The common input is fed into the decoder of the variational generation model to generate the corresponding three-dimensional probability distribution field. And extract the corresponding optimal estimated scalar field from each distribution field. This enables multi-scene three-dimensional visualization of underground structures; Step 8 includes the following sub-steps: Step 8-1: When the system acquires new, field-verified sampling point data, the diagnostic self-evolution module receives the new data and inputs it into the multimodal feature extraction module, the multimodal feature fusion module, the spatial context aggregation module, and the probability space inference module for prediction calculation to obtain the corresponding prediction results; Step 8-2: Calculate the point-level loss between the predicted probability distribution of the sampling points and their measured geological attribute values. ; Step 8-3: Calculate the point-level loss. With preset threshold T The comparison is performed, where the threshold is set based on a statistic of the historical validation point loss distribution: when If the model's prediction performance at that point is deemed satisfactory, the parameter update process is terminated. when When it is determined that the model parameters need to be updated, proceed to step 8-4; Step 8-4: Construct an error diagnosis mechanism based on the point-level error, determine the source of the error, and determine the range of model parameters to be updated accordingly; the error diagnosis is implemented through preset discrimination rules or a learning-based diagnostic model; Step 8-4 includes the following sub-steps: Step 8-4-1: Calculate the prediction error corresponding to each modality based on the single modality input. According to the error distribution characteristics of different modalities, determine that the error mainly comes from the feature encoding model or feature fusion model in the multimodal feature extraction module, thereby determining the corresponding parameter subset. Step 8-4-2: Calculate the difference measure between the latent variable distribution corresponding to the new sampling point and the historical latent variable distribution. When the difference exceeds the preset range, it is determined that there is a local feature distribution shift, thereby determining the relevant parameters of the variational generation model in the probability space inference module that need to be updated. Step 8-5: Based on the parameter range determined in Step 8-4, select the subset of parameters to be updated. And construct the optimization objective function: ; in, The weight coefficients for the regularization term are... For regularization, the remaining unselected parameters are frozen during the parameter update process. For parameters in the parameter subset that belong to the multimodal feature extraction module and the multimodal feature fusion module, gradient optimization is performed using the optimizer. For the variational generation model parameters in the probability space inference module, they are updated through the backpropagation mechanism. Step 8-6: Evaluate the performance of the models determined and updated in Steps 8-4 and 8-5 on an independent validation set. The models include the feature encoding model in the multimodal feature extraction module, the graph neural network model in the spatial context aggregation module, and / or the variational generation model in the probability space inference module. The validation set includes historical validation points and preset validation points. When the following conditions are met: (1) loss of new sampling points (2) If the average reconstruction loss of historical verification points is not higher than the preset range, the updated model is accepted as the current model; otherwise, the original model is retained and the anomaly is recorded.

[0046] The method provided by this invention has the following technical effects: 1) Spatial guidance mechanism breaks through the bottleneck of cross-modal collaboration: It breaks away from the conventional direct feature splicing and uses spatial coordinate information as a guiding condition to perform unified latent space weighted aggregation of multimodal features, which effectively eliminates the semantic gap of heterogeneous data and improves the accuracy of feature expression in complex environments.

[0047] 2) Graph network topology enhancement ensures geological continuity: A graph structure is constructed based on spatial adjacency relationships, and contextual modeling is performed, forcing each node's features to be integrated with its neighborhood structural information. This mechanism follows the "geographically similar" principle, fundamentally eliminating the "island effect" in 3D inference, and making the prediction results highly consistent with the actual physical and geological laws.

[0048] 3) Variational probabilistic extrapolation achieves global uncertainty quantification: A variational generative model replaces traditional deterministic interpolation, mapping enhanced features into a three-dimensional geological attribute probability distribution field. This not only provides high-precision optimal attribute estimation but also, for the first time, achieves uncertainty quantification in deep underground spatial extrapolation, providing reliable mathematical probability support for subsequent engineering drilling and risk decision-making.

[0049] Example 1: Soil samples were collected from a 50m×80m×20m decommissioning site of a chemical plant. The core pollutants were chlorinated hydrocarbon solvents, mainly tetrachloroethylene (PCE) and trichloroethylene (TCE). The upper layer consisted of miscellaneous fill and silty clay, while the lower layer was a sand and gravel aquifer. The pollutants exhibited a complex source-plume spatial distribution. Due to the uncertainty of the underground structure of the contaminated site, the systematic method of this invention was used to generate a three-dimensional probability concentration field of the main pollutants in the soil, clarifying the confidence intervals of the pollution plume boundaries, pollution hotspot locations, and vertical migration depths, providing a basis for accurate risk quantification. The table below shows the multimodal in-situ collaborative sensing data input to the system:

[0050] The primary programming language used is Python, the deep learning framework is PyTorch 1.13+Pyro, and the underground structure visualization module mainly uses PyVista. For hardware, a single server is used, configured with an NVIDIA RTX A6000 GPU (48GB of VRAM) for model training and large-scale 3D simulation.

[0051] In the data calibration module, five standard samples were selected, and the model predicted their concentration distribution. This prediction was then compared with laboratory test results to calibrate the systematic background signal shift, resulting in a linear correction function. .

[0052] In the multimodal data input and classification module, the multimodal in-situ collaborative sensing data from the table above is input and automatically classified to obtain a dataset. .

[0053] In the multimodal feature extraction module, for numerical data... Feature extraction can be performed using a multilayer perceptron. In this embodiment, the multilayer perceptron has a three-layer structure with an input dimension of 2, a hidden layer dimension of [16, 32], an output dimension of 32, and an activation function of ReLU. For time-series data... Feature extraction can be performed using a one-dimensional convolutional neural network. In this embodiment, the one-dimensional convolutional neural network includes two convolutional layers with a kernel size of 5 and channels of 32 and 64 respectively. After global max pooling, a 64-dimensional feature vector is output through a fully connected layer. For spatial structure data... Feature extraction is performed using a convolutional neural network (CNN). This CNN can be pre-trained on a large-scale image dataset and obtains high-dimensional feature representations by removing the classification layer. Preferably, the feature vector from the penultimate layer of the network is extracted as the image encoding representation, for example, resulting in a 512-dimensional feature vector. The multimodal feature extraction module outputs a set of feature vectors corresponding to each modality. , , The dimension of each modal feature vector is determined by the corresponding encoder structure.

[0054] In the multimodal feature fusion module, a simple MLP is used to fuse three-dimensional coordinates. The mapping is to a 32-dimensional query vector Q. For each modality, a learnable projection matrix is ​​used. Features Projection. For contaminated sites, especially for A larger attention head dimension, head_dim=64, is set to capture more refined vertical contamination variation signals. The features from the three modalities are concatenated and then followed by a fully connected layer with an output dimension of 128 to form a fused feature. .

[0055] In the spatial topology graph construction module, the input coordinate set employs a hybrid strategy of "distance threshold + K-nearest neighbors". First, the Euclidean distance between sampling points is calculated, with a horizontal threshold of 15m and a vertical threshold of 5m as preliminary filtering. For each sampling point, K=6 nearest points are selected from the initially filtered neighbors as the final neighbors. The output is a spatial relationship graph. It contains 150 nodes, and each node is connected to an average of about 6 edges.

[0056] In the spatial context aggregation module, a two-layer graph attention network is used. The first layer has 4 attention heads and an output dimension of 64, while the second layer has 2 attention heads for aggregation, resulting in an output dimension of 128. The ELU activation function is applied. Through message passing in the two-layer graph attention network, the features of each node are aggregated with information from approximately 12 of its neighbors, resulting in an enhanced feature vector with a dimension of 128. .

[0057] In the probability space inference module, the encoder is a 4-layer MLP with a 128-dimensional input, a hidden layer dimension of [256, 128], and an output of the 16-dimensional mean of the latent variable z. and 16-dimensional log-variance The decoder is a 4-layer MLP, with inputs consisting of a 16-dimensional latent variable z and 19-dimensional three-dimensional coordinates. The concatenation of the hidden layers [128, 256] outputs the log-mean and log-standard deviation of the predicted concentrations. , It is a 16-dimensional identity matrix. For the total loss... ,use -VAE framework, settings To balance reconstruction accuracy and latent space regularity, the optimizer Adam was chosen, with a learning rate of lr=1e-4 and 800 training epochs. After training, a 1m×1m×0.5m grid was used to reconstruct the entire site.

[0058] In the underground structure visualization module, the expected concentration value is calculated from the predicted distribution of each grid point. A linear correction function is applied to all expected values ​​to eliminate systematic errors. Based on the corrected concentration field, an isosurface with a concentration equal to the risk control standard of 1 mg / kg is extracted to generate a 3D model of the "most likely" pollution plume. Simultaneously, an uncertainty metric is calculated to obtain a 3D uncertainty field. ,Will The data is mapped to transparency and overlaid in a semi-transparent form on the 3D model of the "most likely" pollution plume. High transparency represents low uncertainty, and low transparency represents high uncertainty. Five possible pollution source migration path scenarios, generated from latent variables, are used to assess long-term risk.

[0059] In the diagnostic self-evolutionary module, a new sampling point V01 was added. Its measured PCE concentration was 15.2 mg / kg, while the model predicted the following distribution: The corresponding mean was approximately 8.2 mg / kg, with a 95% percentile range of [1.7, 39.5]. Although the measured values ​​were within this range, there was a point-level loss. The error was high, triggering an evolutionary process. During error diagnosis, it was found that the loss at this point was extremely high when using only MIP data, while the loss was normal when using other data. The problem was judged to stem from the fact that the specific response of the MIP sensor under this depth condition was not well learned. Simultaneously, the encoded data at this point... z With training set z The large KL divergence in the distribution confirms it as a "new mode." Based on the diagnostic results, only the parameters of the last two layers of the MIP feature extraction 1D-CNN and the projection matrix of the MIP modal in the fusion module were unfrozen and updated. To optimize the objective, 50 fine-tuning steps were performed. After fine-tuning, the average predicted concentration at this point increased to 12.8 mg / kg, the loss decreased by 60%, and the performance remained stable after validation at 5 other historical validation points, allowing for the deployment of the new model.

[0060] This embodiment demonstrates the ability of the present invention to infer and generate three-dimensional structures of underground structures in contaminated sites. It can deeply integrate sparse and heterogeneous laboratory data and geophysical data to generate a probability model with confidence that can directly support risk quantification decision-making. Furthermore, it has the ability to continuously self-optimize as the investigation progresses, achieving a leap from "data interpretation" to "intelligent inference".

[0061] Example 2: A 2km long exploration section of a subway extension project in a certain city has an exploration depth of 0-50m. Using the systematic method of this invention, and leveraging sparse borehole data, the three-dimensional lithological distribution and soil density along the route are probabilistically predicted, providing a quantitative assessment of tunnel boring machine (TBM) construction risks. A total of 32 boreholes were drilled, with an average spacing of 60m and a maximum depth of 50m. Other multimodal in-situ collaborative sensing data are shown in the table below.

[0062] The primary programming language used is Python, and the deep learning framework employed is PyTorch 1.12 + Pyro. Key module parameter settings are shown in the table below:

[0063] In the data calibration module, systematic error calibration was performed using boreholes from five known standard samples, and a linear correction function for density prediction was fitted. In the underground structure visualization module, the "most likely" lithological structure displays a three-dimensional spatial distribution of silty clay, sand, and weathered rock; the density uncertainty heatmap shows that in areas with borehole spacing greater than 80m, the prediction standard deviation increases significantly, exceeding 0.15g / L. It intuitively indicates sparse data areas; it generates 10 possible lithological structure scenarios from prior distribution sampling, demonstrating the different connectivity possibilities of local sand layer lenses.

[0064] In the diagnostic self-evolution module, the 95% confidence interval of the density prediction values ​​covered the measured values ​​of all five newly added boreholes. The system triggered self-evolution at one of the boreholes, and the diagnosis revealed that the resistivity image characteristics at that point differed significantly from historical patterns. Selective fine-tuning of relevant parameters of the image encoder resulted in a 65% reduction in loss at that point.

[0065] Example 3: Along a planned integrated utility tunnel in a major city, the tunnel is 5 meters long and 10-20 meters deep. Using the systematic method of this invention, along with survey data and ground-penetrating radar data, geological hazards such as hidden cavities, weak interlayers, and water-rich areas are probabilistically identified, providing a basis for design avoidance and reinforcement. The input multimodal in-situ collaborative sensing data is shown in the table below:

[0066] The primary programming language used is Python, and the deep learning framework employed is TensorFlow 2.8, accelerated by a GPU. Key module parameter settings are shown in the table below:

[0067] In the data calibration module, systematic error calibration was performed using boreholes from five known standard samples, and a linear correction function for density prediction was fitted. In the underground structure visualization module, the distribution map of the "most likely" hazard bodies is based on the pipeline axis, generating a longitudinal profile along the line. Different colors represent different hazard body types, and transparency indicates the probability. It shows three high-probability (P>0.85) continuous cavity areas, two soft soil interlayer zones, and one water-rich anomaly area. The uncertainty heatmap uses standard deviation as a measure. At the discontinuity of GPR survey lines and in sections with borehole spacing greater than 100m, the predicted standard deviation increases significantly (>0.25), indicating that these areas need further exploration. Eight possible geological profile scenarios are generated from prior distribution sampling. Two of these scenarios show different spatial configurations of water-rich areas and cavities, used to assess the coupling risk under extreme conditions. In the diagnostic self-evolution module, one newly added borehole was measured to be a small-scale cavity. The model predicts its "cavity" probability field to be 0.6, but the "scale intensity" prediction error is large, and the point-level loss exceeds the threshold. After diagnosis, it was found that the classification and regression errors for this point were high when using only GPR images, while the errors were normal when using other data. Furthermore, latent space analysis showed a significant difference between this point's z-shape and the training set, indicating that the reflection patterns of such small-scale holes in GPR images had not been sufficiently learned. Therefore, only the last two convolutional layers of the GPR encoder were unfrozen. To optimize the objective, a small learning rate of 1e-5 was used for fine-tuning over 80 steps. After fine-tuning, the probability of a "hole" at that point increased to 0.78, the scale prediction error decreased by approximately 70%, and the performance remained stable after validation at another 10 historical validation points, allowing for model deployment updates.

[0068] In summary, the method proposed in this invention first constructs a systematic error calibration mechanism based on standard sample data with known geological attribute measured values ​​to correct and optimize the consistency of multi-source heterogeneous data. Then, through multimodal feature encoding and fusion methods, data from different sources are mapped to a unified latent space, and a cross-modal feature selection mechanism is constructed by introducing spatial coordinate conditions to achieve unified feature expression and enhanced spatial correlation. Based on this, a topological structure is constructed based on the spatial relationship of sampling points, and spatial context information is propagated using a graph neural network. Furthermore, a deep probability model based on variational inference learns the mapping relationship between features and the probability distribution of underground structures, generating a three-dimensional probability distribution field and its three-dimensional representation of quantified uncertainty. Finally, a diagnostic mechanism is constructed based on prediction errors, and the model is adaptively updated through selective parameter updates. This invention addresses the problem that traditional methods struggle to effectively quantify uncertainty under sparse data conditions, achieving quantitative characterization and dynamic updating of underground structural uncertainty. It can be applied to underground space exploration, pollution assessment, and sampling fidelity verification.

Claims

1. A multi-modal probabilistic modeling subsurface structure inference system, characterized in that, It includes a multimodal feature extraction module (6), the output of which is used to output several feature vectors and query vectors to the multimodal feature fusion module (7); the output of the multimodal feature fusion module (7) and the spatial topology graph construction module (9) are connected to the input of the spatial context aggregation module (10), and the output of the spatial context aggregation module (10) is connected to the input of the probability space inference module (3).

2. The system according to claim 1, characterized in that, The diagnostic self-evolution module (5) triggers the error assessment and parameter update process based on the newly added measured data. It determines the parameter subset to be updated through error diagnosis and optimizes and updates the parameter subset under the condition of meeting the preset performance constraints, so as to realize the adaptive control of the multimodal feature extraction module, the multimodal feature fusion module and the probability space inference module.

3. The system according to claim 1, characterized in that, It also includes a multimodal data input and classification module (8), the output of which is connected to the input of the multimodal feature extraction module (6) and the spatial topology map construction module (9), respectively. The multimodal data input and classification module (8) is used to receive multimodal in-situ collaborative sensing data of the target area and classify it according to the data structure to provide standardized input for feature encoders of different modalities.

4. The system according to claim 3, characterized in that, It also includes a data calibration module (1), whose input is used to receive multimodal in-situ collaborative sensing data and its corresponding known geological property measured values ​​obtained based on standard samples. Its output is connected to the input of the underground structure visualization module (2), and the output of the probability space inference module (3) is connected to the input of the underground structure visualization module (2). It also includes an optimizer (4) and a diagnostic self-evolution module (5). The output of the probability space inference module (3) is connected to the input of the optimizer (4) and the diagnostic self-evolution module (5), respectively. The output of the optimizer (4) is connected to the input of the multimodal feature extraction module (6) and the multimodal feature fusion module (7), respectively.

5. The system according to claim 4, characterized in that, The data calibration module (1) is used to receive multimodal in-situ collaborative sensing data and corresponding known geological attribute measured values ​​obtained based on standard samples, wherein the multimodal data and the measured values ​​correspond one-to-one; by inputting the multimodal data into the model composed of the multimodal feature extraction module (6), the multimodal feature fusion module (7), the spatial context aggregation module (10) and the probability space inference module (3), the predicted probability distribution is obtained, and the statistical optimal estimate of the predicted probability distribution is compared with the measured value. A calibration function is constructed by using regression analysis or error mapping method to realize the pre-calibration of system error and provide a calibration benchmark for subsequent underground structure inference and three-dimensional expression.

6. The system according to claim 1, characterized in that, The spatial topology graph construction module (9) constructs spatial adjacency relationships based on the spatial coordinate data of sampling points and generates a graph structure describing the spatial connection relationship between sampling points, providing a topological basis for spatial context modeling; the multimodal feature extraction module (6) is equipped with a modality adaptation feature encoder group, which extracts features from different modal data to obtain the feature vector representation corresponding to each modality; The multimodal feature fusion module (7) constructs a conditional query mechanism based on spatial coordinate information, selectively fuses each modality feature through weighted aggregation, and maps it to a shared latent space to obtain a unified multimodal fusion feature vector; The context aggregation module (10) receives the multimodal fusion feature vector and spatial relationship graph, and performs neighborhood information propagation and aggregation on the node features through the graph neural network to obtain the enhanced feature vector with fused spatial context information.

7. The system according to any one of claims 1 to 6, characterized in that, The probability space inference module (3) has a built-in deep probability model, which performs probability distribution modeling and spatial inference of multimodal fusion features through encoder and decoder; preferably, the deep probability model is a variational generation model based on variational inference; the diagnostic self-evolution module (5) performs error evaluation and diagnosis based on newly added measured data after model deployment, and determines the parameter subset to be updated according to the diagnosis results, and calls the optimizer (4) to optimize and update the parameter subset, thereby realizing the online adaptive update of the feature extraction model and the feature fusion model; and generates a three-dimensional visualization expression of the underground structure based on the probability space inference results.

8. The system according to claim 7, characterized in that, When this system is in operation / use, the following steps are followed: Step S1: Based on standard samples with measured values ​​of known geological attributes, acquire multimodal in-situ collaborative sensing data; Step S2: Based on the data obtained in step S1, classify the data according to the data structure type, and obtain the feature vector representation of each modality through feature extraction; Step S3: Construct a conditional query mechanism and perform weighted aggregation based on the feature fusion model, map the fusion result to a unified latent space representation, and obtain a multimodal fusion feature vector; Step S4: Construct spatial adjacency relationships and generate a spatial relationship diagram describing the spatial connection relationships between sampling points; Step S5: Combine the multimodal fusion feature vector with the spatial relationship graph to perform spatial propagation and context modeling of node features; Step S6: Train the variational generation model and use the variational generation model to infer the target area, generating a three-dimensional geological attribute probability distribution field covering the entire area; It also includes step S7: based on the three-dimensional geological attribute probability distribution field generated in step S6, the statistical characteristics and uncertainties of each grid point are calculated and mapped, and multi-scenario inference is performed under different sampling conditions through the variational generation model to realize the visual expression of the underground structure. It also includes step S8: constructing an error assessment and parameter update mechanism based on the newly added sampled data, determining the parameter subset that needs to be updated by diagnosing and analyzing the prediction error of the variational generation model; and using an optimizer to iteratively update the model parameter subset under the condition of satisfying the preset performance constraints, thereby achieving adaptive updating. In step S1: Based on standard samples with measured values ​​of known geological attributes, multimodal in-situ collaborative sensing data is acquired, and a mapping relationship between the prediction results and the measured values ​​is constructed to achieve system error calibration; In step S2: acquire multimodal in-situ collaborative sensing data of the target area, classify and process the data according to the data structure type, extract features from different types of data through the feature encoding model in the multimodal feature extraction module, and obtain feature vector representations of each modality; In step S3, a conditional query mechanism is constructed using the spatial coordinates of the sampling points as conditional variables. Through the feature fusion model in the multimodal feature extraction module, weighted aggregation is performed based on the coordinate-guided cross-modal attention mechanism, and the fusion result is mapped to a unified latent space representation to obtain a multimodal fusion feature vector with spatial correlation. In step S4, spatial adjacency relationships are constructed based on the spatial coordinate information of the sampling points, and a spatial relationship diagram describing the spatial connection relationships between the sampling points is generated to characterize the spatial topology of geological attributes. In step S5, the multimodal fusion feature vector is combined with the spatial relationship graph. The graph neural network model in the spatial context modeling module performs spatial propagation and context modeling on the node features, so that the features of each node are fused with its neighborhood structure information, thereby obtaining an enhanced feature representation with spatial relevance. In step S6, the enhanced feature vector is input into the variational generation model in the probability space inference module for training. The model parameters are optimized by constructing reconstruction loss and regularization loss, so that the variational generation model learns the mapping relationship between enhanced features and the probability distribution of geological attributes. After training, the variational generation model is used to infer the target area and generate a three-dimensional geological attribute probability distribution field covering the entire area.

9. The system according to claim 8, characterized in that, Step 1 includes the following sub-steps: Step 1-1: Construct a standard sample set with known geological attribute measured values, and arrange it in the target area according to the preset three-dimensional spatial coordinates; Steps 1-2: Perform multimodal collaborative acquisition on the standard samples to obtain drilling signals, geophysical data, image data, in-situ sensing data and experimental analysis data at the corresponding locations; Steps 1-3: Obtain the geological attribute prediction probability distribution for each standard sample location based on the multimodal data; Steps 1-4: Based on the deviation relationship between the predicted probability distribution and the measured values ​​of known geological attributes, a systematic error calibration function is constructed through regression analysis or error mapping methods to correct the prediction results output by the subsequent probability space extrapolation module.

10. The system according to claim 8, characterized in that, Step 2 includes the following sub-steps: Step 2-1: Collect and input multimodal in-situ collaborative sensing data of the target area to construct a sensing data set. Multimodal data includes time-series signals acquired during drilling operations, geophysical exploration data, geological image data, in-situ sensor monitoring data, and geological attribute data of sampling points; spatial topology data is also acquired. , representing the three-dimensional spatial coordinates of each sampling point, the spatial topology data C Used for subsequent spatial relationship modeling; in some embodiments, structured data describing the spatial topological relationship between sampling points may be further introduced to supplement the spatial topological data; Step 2-2: Classify the multimodal data according to its representation and structural characteristics, dividing it into numerical datasets. Time series data sets and spatial structure data sets ; Steps 2-3: Input each type of data into the corresponding feature encoding model in the multimodal feature extraction module for feature extraction. The feature encoding model, depending on the data type, can employ one or more of the following: convolutional neural network, recurrent neural network, or attention-based network structure. For any type... , Its eigenvectors are represented as follows: ; ; in, For any type eigenvectors, For the first t The total number of data instances in a modal dataset. For the first j Feature vectors of data instances For the first t Feature coding model corresponding to class data For the first t The first in the class modal data set j One data instance, Its learnable parameter set; Steps 2-4: Output the feature vector set for each modality , , , which serves as the input to the subsequent multimodal feature fusion model, where: , ; It is a set of numerical modal eigenvectors. It is a set of temporal modal feature vectors. It is the set of spatial structural modal feature vectors.