A Smart Prediction Method for Peak Elastic Performance of Rock Integrating Physical Constraints and Multimodal Deep Learning
By integrating physical constraints and multimodal deep learning, scanning electron microscope images, energy-dispersive X-ray spectroscopy data, and X-ray diffraction data are combined to solve the problem of heterogeneous data integration for predicting the peak elastic properties of rocks, and achieve high-precision and interpretable prediction of the peak elastic properties of rocks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANHUA UNIV
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to effectively integrate scanning electron microscope images, energy-dispersive X-ray spectroscopy data, and X-ray diffraction data to achieve high-precision, interpretable predictions of rock peak elastic properties. Furthermore, the lack of physical constraints in traditional deep learning models leads to inconsistent prediction results.
A visual transformer network and a fully connected neural network are used to extract multimodal data features. The elastic strain energy density theory and mineral elastic modulus are used as physical constraints. Feature alignment is achieved through optimal transport theory and entropy regularization. A composite loss function is constructed to drive model training, and parameters are optimized by combining mini-batch stochastic gradient descent algorithm.
It achieves high-precision and interpretable prediction of peak elastic properties of rocks, is applicable to different lithologies and confining pressure conditions, shortens the evaluation cycle, reduces costs, and improves the reliability and stability of prediction results.
Smart Images

Figure CN122494020A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rock mechanics testing and intelligent evaluation technology for geological engineering, specifically to an intelligent prediction method for peak elastic properties of rocks that integrates physical constraints and multimodal deep learning. Background Technology
[0002] With the implementation of the deep-earth strategy, deep-earth engineering projects such as deep mineral resource development, underground energy storage, nuclear waste disposal, and deep geothermal development are advancing in depth. Peak elastic performance evaluation of rocks has become a key parameter for disaster prevention and control, rock mass stability analysis, and engineering design in deep engineering. Deep rocks are affected by complex geological environments such as high ground stress, high ground temperature, and high osmotic pressure. Peak energy characteristics directly affect rockburst prediction, surrounding rock fracturing, support structure design, and the effectiveness of deep reservoir stimulation. However, my country's deep geological conditions are extremely complex, with diverse rock types and the coupling effects of multi-scale factors such as mineral composition, microstructure, and chemical composition. This results in significant nonlinearity and multi-source heterogeneity in peak energy characterization, making it difficult for traditional single-method testing to fully capture the physical essence of rock peak elastic performance. Therefore, accurate and rapid prediction of rock peak elastic performance has always been a challenge in deep-earth engineering.
[0003] Peak elastic properties of rocks are controlled by multiple factors, including microstructure, elemental composition, and mineral phases, thus requiring comprehensive multimodal data for evaluation. While existing methods, such as stress-strain curve analysis based on mechanical tests, can directly obtain the elastic strain energy of rocks, these methods are time-consuming, labor-intensive, and require sophisticated sample preparation, making them unsuitable for large-scale engineering applications. Empirical formula methods based on well logging data can achieve rapid prediction, but they rely on experimental calibration in specific blocks, limiting their generalization ability and making them unsuitable for evaluating peak elastic properties of rocks under different lithologies and geological conditions.
[0004] At the data acquisition level, scanning electron microscopy can acquire images of the microscopic morphology of rocks, reflecting key information such as pore structure, grain boundaries, and microcrack distribution; energy-dispersive X-ray spectroscopy can provide elemental composition data of the target area, revealing the distribution characteristics of elastic reinforcing elements such as Si and Al and plastic reinforcing elements such as Ca and Mg; X-ray diffraction can obtain mineral phase composition and relative content, clarifying the ratio of high-energy-storage minerals such as quartz and feldspar to low-energy-storage minerals such as chlorite. These three characterization techniques reveal the physical essence of rocks from three dimensions: microscopic morphology, chemical composition, and mineralogical properties, forming a complementary multi-scale analysis system. However, how to effectively integrate these three types of heterogeneous multimodal data and establish a quantitative mapping relationship between microstructure and macroscopic energy storage capacity is a key technical bottleneck restricting the intelligent prediction of peak elastic performance of rocks.
[0005] In recent years, deep learning methods have been widely applied in the field of rock property prediction due to their powerful feature extraction and nonlinear mapping capabilities. Visual transformer networks excel in image feature extraction, capturing global dependencies in images through self-attention mechanisms; fully connected neural networks can effectively process numerical data, extracting deep features of elemental composition and mineral components. However, existing technologies still face the following challenges: First, SEM images, EDS data, and XRD data are located in different feature spaces, and their dimensions, units, and physical meanings are significantly different. The inconsistency of feature spaces of heterogeneous data restricts the effective fusion of cross-modal information. Second, a single prediction model is insufficient to simultaneously characterize the synergistic regulatory effects of microstructure, elemental composition, and relative peak elastic properties of minerals, and the prediction accuracy and interpretability need to be improved. Third, deep learning models are usually black-box models, and their internal decision-making processes lack the constraints of physical mechanisms. The prediction results may contradict the theory of rock elasticity and mechanics, making it difficult to meet the reliability requirements of engineering applications.
[0006] Therefore, there is an urgent need for an intelligent evaluation method for the peak elastic properties of rocks that can effectively integrate multimodal heterogeneous data, establish physical mechanism constraints, and achieve high-precision interpretable prediction.
[0007] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention provides an intelligent prediction method for peak elastic properties of rocks that integrates physical constraints and multimodal deep learning, thereby solving the problems mentioned in the background section.
[0009] This invention provides the following technical solution: an intelligent prediction method for peak elastic properties of rocks that integrates physical constraints and multimodal deep learning, comprising the following steps: S1. Obtain scanning electron microscope images, energy dispersive X-ray spectroscopy elemental composition data, and X-ray diffraction mineral phase composition data of the rock sample to be tested. S2. The scanning electron microscope images are cropped and augmented, and the elemental composition data (EDS) and mineral phase composition data (XRD) from the energy dispersive X-ray spectrometer are standardized to construct a multimodal dataset. S3. A visual transformer network is used to extract the microscopic morphology features of scanning electron microscope images. A fully connected neural network is used to extract the elemental composition data of energy dispersive X-ray spectrometer and the mineral phase composition data of X-ray diffractometer, respectively. The elastic strain energy density theory and the mineral elastic modulus mixing law are used as physical constraints and embedded into the loss function. S4. Calculate the cost matrix between different modal features, adopt the optimal transfer method with entropy regularization, and solve the transfer matrix iteratively through the Sinkhorn algorithm to achieve soft alignment of the feature spaces of three modal features: scanning electron microscope, energy dispersive X-ray spectrometer, and X-ray diffractometer. S5. A multi-head attention mechanism is used to integrate the aligned cross-modal features, construct the microstructure energy storage index, elemental energy storage index and mineral phase energy storage index, and dynamically adjust the contribution of different modes through an adaptive weight calculation method to establish a multi-scale physical mechanism mapping relationship between micromorphology, chemical composition and mineralogical properties. S6. The model training is driven by a composite loss function that integrates physical mechanisms, and the model parameters are updated using a mini-batch stochastic gradient descent algorithm, combined with regularization techniques to prevent overfitting. S7. Input the multimodal data of the rock sample to be predicted into the trained model, and output the predicted value of the peak elastic performance of the rock.
[0010] Preferably, step S1 further includes: obtaining a standard cylindrical rock sample, conducting a uniaxial compression test using a rock mechanics testing system, simultaneously acquiring a complete stress-strain curve, and calculating the elastic strain energy density U as the target value for model training based on the peak stress σp and elastic modulus E according to the formula U=σp² / 2E.
[0011] Preferably, in step S1, the scanning electron microscope imaging scale is 5 μm, the elements obtained by the energy dispersive X-ray spectrometer include at least O, Na, Mg, Al, Si, K, Ca, and Fe, the X-ray diffractometer scanning angle range is 5-90°, and the particle size of the rock sample is less than 75 μm.
[0012] Preferably, in step S3, the visual transformer network extracts microscopic morphological features including: The input image X is divided into fixed-size patches, and each patch is mapped to a D-dimensional feature space through linear projection, resulting in... ; Add position encoding to each patch: A learnable positional encoding matrix; Feature extraction is performed using an L-layer Transformer encoder, with each layer containing a multi-head self-attention mechanism and a feedforward neural network. The computation process is as follows: z' l = MHA(LN(z){l-1})) + z_{l-1}; z_l = MLP(LN(z'_l)) + z'_l; Where z is the feature vector of each layer of the Transformer encoder; MHA is multi-head self-attention; LN is layer normalization; MLP is multilayer perceptron; and l is the layer index.
[0013] Preferably, the extraction of deep features by the fully connected neural network in step S3 includes: For EDS and XRD data, construct multi-layer fully connected neural networks respectively; the forward propagation process can be represented as: ,and ; In the formula, l The first term of the neural network l layer, Indicates the first l The output of layer neurons, To connect the weight matrix, For bias vectors, For activation function, This represents the weighted input received by the neuron; This represents the number of nodes in the current layer. This represents the number of nodes in the previous layer. The parameters are updated using the backpropagation algorithm, and the gradient of the loss function L with respect to the weights and biases is calculated: ; ; ; In the formula, For the weight gradient, For the bias gradient, Here, T represents the error term, and T is the transpose operation. It is the reciprocal of the activation function.
[0014] Preferably, step S4 further includes: Calculate the Euclidean distance between different modal features as the transmission cost: ; ; ; in, ,and ,and These represent the feature vectors of each mode; is the cross-modal distance between the $i$-th SEM sample and the $j$-th EDS sample; is the cross-modal distance between the $i$-th SEM sample and the $k$-th XRD sample; is the cross-modal distance between the $j$-th EDS sample and the $k$-th XRD sample; Solve the optimal transport problem in the form of entropy regularization: $\min_{T\in\Pi(\mu,\gamma)}\langle T,C\rangle-\lambda H(T)$; where $T$ is the transport matrix, $\mu$ is the marginal probability distribution of the source modality, $\gamma$ is the marginal probability distribution of the target modality, and $H(T)$ is the entropy regularization term, where $H(T)=-\sum_{i,j}T_{ij}(\log T_{ij}-1)$, $T$ , , ,
[0015] , , j , , i , , β , C , , - , , + , , , C , , , is the transport volume from the $i$-th source sample to the $j$-th target sample, and $\lambda$ is the regularization parameter; Achieve feature alignment through the transport matrix: ; ; ; where represents transforming the features of the EDS modality through the optimal transport matrix to the SEM modality space, is the optimal transport matrix between the ESD and SEM modalities, which needs to be recalculated through the reverse OT problem rather than directly transposed.
[0015] Preferably, in step S5, the microstructure energy storage index is calculated by the following formula: ; where is the microstructure energy storage index, is the porosity, is the fracture density parameter, is the weight coefficient. The lower the porosity and fracture density, the stronger the energy storage capacity of the rock; The element energy storage index is calculated by the following formula: ; where C i + are the contents of the elastic strengthening elements Si and Al, C j - are the contents of the plastic strengthening elements Ca and Mg, β >i + and β j - These are the corresponding weighting coefficients; Mineral phase energy storage index The equivalent elastic modulus of mineral phases, calculated based on the Voigt-Reuss-Hill mixing law, is obtained using the following formula: ; in, The mineral phase energy storage index, Let be the relative content of the i-th mineral. Let n be the elastic modulus of the mineral, and n be the total number of mineral types.
[0016] Preferably, the adaptive weight calculation method in step S5 is as follows: ; in, For the adaptive weights of mode i, The independent peak energy prediction score for this mode. This is a sensitivity parameter that controls the sensitivity of weight allocation.
[0017] In step S5, feature fusion is performed end-to-end through a deep neural network, as shown below: ; ; in, X The input vector is a weighted concatenation of the three modal energy storage indices. Let L be the activation function of the Lth layer. For the first L The weight matrix of the layer, , , For each modality, adaptive weights, This is the predicted value of the peak elastic energy of the rock.
[0018] Preferably, the composite loss function in step S6 is: ; For mean square error loss, ; Physical consistency penalty item , ; This is a soft constraint penalty term for monotonicity. ; in, mFor the sample size, To predict the peak energy value, These are the weighting coefficients for each constraint term.
[0019] In step S6, parameter optimization uses the mini-batch stochastic gradient descent algorithm, and the parameter update formula is as follows: ; ; in, θ For model parameters, Let be the parameter value at step t. ηt Let be the learning rate at step t. The gradient of the loss function with respect to the parameters. The initial learning rate is 0.01; The attenuation coefficient is... T This is the decay period.
[0020] Preferably, step S7 also includes feature importance analysis: using three measurement methods—Pearson correlation coefficient, Spearman rank correlation coefficient, and mutual information value—to evaluate the contribution of each feature to the prediction of peak elastic performance, and to identify key control factors, including quartz content, feldspar content, estimated elastic modulus, and porosity; for different lithologies (granite, sandstone, etc.), the model can automatically adjust the relative importance of each physical factor to achieve generalized prediction of peak energy storage capacity of rocks under different lithologies and confining pressures.
[0021] The present invention provides an intelligent prediction method for peak elastic properties of rocks that integrates physical constraints and multimodal deep learning, which has the following beneficial effects: First, this invention establishes a multi-scale analysis system encompassing microstructure, chemical composition, and mineralogical properties by acquiring rock microstructure images using scanning electron microscopy, elemental composition data using energy-dispersive X-ray spectroscopy, and mineral phase composition data using X-ray diffraction. Based on this, it constructs a cost matrix between different modal features using optimal transport theory, and iteratively solves the transport matrix using entropy regularization and the Sinkhorn algorithm, achieving soft alignment of the three modal feature spaces. This effectively solves the problem of inconsistencies in dimension, scale, and physical meaning among heterogeneous data, laying a technical foundation for the deep integration of cross-modal information.
[0022] Second, this invention embeds the elastic strain energy density theory and the mineral elastic modulus mixing law as physical constraints into the loss function, while simultaneously applying soft constraints of peak energy non-negativity and monotonicity, thus constructing a composite loss function that integrates physical mechanisms. This design ensures that the model's prediction results are consistent with the theory of rock elasticity, overcoming the shortcomings of traditional deep learning models where the opaque internal decision-making process leads to prediction results that may violate physical laws, significantly improving the reliability and credibility of the prediction results.
[0023] Third, this invention extracts microscopic morphology features through a visual transformer network and extracts elemental composition and mineral component features through a fully connected neural network. Based on this, it constructs microstructure energy storage indices, elemental energy storage indices, and mineral phase energy storage indices. By integrating aligned cross-modal features through a multi-head attention mechanism and dynamically adjusting the contribution of different modes using an adaptive weight calculation method, a multi-scale physical mechanism mapping relationship of rock peak elastic energy under the synergistic effect of microscopic morphology, chemical composition, and mineralogical properties is established. This significantly enhances the interpretability of the model and provides a quantitative basis for understanding the physical nature of rock peak energy storage capacity.
[0024] Fourth, this invention uses the elastic strain energy density obtained from uniaxial compression tests as a supervision label, employs a composite loss function that integrates physical mechanisms to drive model training, and combines mini-batch stochastic gradient descent algorithm with regularization techniques to prevent overfitting. Experimental results show that the model's predictions under different rock types are in high agreement with measured and theoretical values. Compared with single-modal or non-physical constraint prediction methods, its prediction accuracy and stability are significantly improved, and it has good cross-lithological generalization ability, making it applicable to the prediction of peak elastic energy in various rock types such as granite, sandstone, and shale under different confining pressures.
[0025] Fifth, compared to traditional stress-strain curve analysis methods based on uniaxial compression tests, this invention only requires the acquisition of scanning electron microscope images, energy-dispersive X-ray spectroscopy elemental composition data, and X-ray diffraction mineral phase composition data of rock samples to quickly output predicted peak elastic properties of the rock, without the need for destructive mechanical tests. This method significantly shortens the evaluation cycle, reduces sample preparation and testing costs, and provides an efficient and accurate peak elastic property evaluation tool for deep-earth engineering problems such as rockburst prediction, surrounding rock fracturing, support structure design, and deep reservoir stimulation. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating an intelligent prediction method for peak elastic properties of rocks that integrates physical constraints and multimodal deep learning. Figure 2 This is a schematic diagram of a multimodal data peak elasticity prediction framework; where, Figure 2 a is a schematic diagram of the overall structure for multimodal data peak elastic performance prediction; Figure 2 b is a schematic diagram of the feature extraction structure; Figure 2 c is a schematic diagram of the feature alignment structure; Figure 2 d is a schematic diagram of the feature reconstruction structure; Figure 2 e is a schematic diagram of the physical mechanism mapping structure; Figure 3 This is a diagram illustrating a framework for applying the Transformer architecture to image feature extraction. Figure 4 This is the FCNN network structure for numerical feature extraction. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] To address the problems mentioned in the background section, this invention provides an intelligent prediction method for the peak elastic properties of rocks that integrates physical constraints and multimodal deep learning, thereby solving the aforementioned technical problems. The technical solution is as follows: The following is in conjunction with the appendix Figure 1-4 The present invention will be further described in detail below, along with specific embodiments.
[0029] I. An intelligent prediction method for peak elastic properties of rocks, which integrates physical constraints and multimodal deep learning, is provided in this embodiment of the invention, as follows: S1. Multimodal data acquisition: Scanning electron microscopy was used to acquire images of the rock microstructure at an imaging scale of 5 μm; energy dispersive X-ray spectroscopy was used to acquire volume fraction data of elements such as O, Na, Mg, Al, Si, K, Ca, and Fe; and X-ray diffractometer was used to acquire mineral phase composition and relative content data within a scanning angle range of 5-90°. S2. Data Preprocessing: The acquired SEM images are cropped to 224×224 pixels; data augmentation is performed using methods such as contrast enhancement, vertical flipping, horizontal flipping, and center flipping; a multimodal dataset is constructed and divided into a 70% training set, a 15% validation set, and a 15% test set. S3. Feature Extraction: The Vision Transformer (ViT) network is used to extract microscopic morphological features through image segmentation, position encoding, and multi-layer self-attention mechanism; the fully connected neural network (FCNN) is used to process EDS and XRD data respectively, and deep features of elements and mineral components are extracted through multi-layer nonlinear transformation. S4. Feature Alignment: Calculate the cost matrix between different modal features and use the optimal transfer method with entropy regularization; solve the transfer matrix iteratively through the Sinkhorn algorithm to achieve soft alignment of the feature spaces of SEM, EDS, and XRD. S5. Feature Fusion and Physical Mechanism Mapping: A multi-head attention mechanism is used to integrate aligned cross-modal features; peak elastic energy indices of microstructure, elements, and mineral phases are constructed; and the contribution of different modes is dynamically adjusted through an adaptive weight calculation method. S6. Model Training: The mean squared error is used as the loss function, and the model parameters are updated using the mini-batch stochastic gradient descent algorithm; the initial learning rate is set to 0.01, and an exponential decay strategy is adopted; L2 regularization and Dropout techniques are combined to prevent overfitting; the model performance is evaluated on the validation set. S7. Peak Elastic Performance Prediction: Input the multimodal data of the rock sample to be predicted into the trained model, and output the predicted value of the peak elastic performance index of the rock.
[0030] In this embodiment, step S1 further includes: S11. Rock Sample Preparation and Combined SEM-EDS Characterization: Select the rock sample to be tested, preferably typical deep-earth engineering rocks such as granite, sandstone, or shale. Cut the sample into pieces of approximately 1 cm³, with the specific size adjustable within the range of 0.8-1.5 cm³. Fix the sample onto the stage using epoxy resin or conductive adhesive. Use a scanning electron microscope, preferably a Hitachi SU1510 or equivalent, to image the sample surface. The imaging parameters are set as follows: imaging scale is 5 μm, which can be adjusted within the range of 1 to 10 μm as needed. Smaller scales allow for the observation of finer microstructural features. Accelerating voltage is 15 kV, ranging from 10 to 20 kV. Working distance is 10 mm, ranging from 8 to 15 mm. When imaging, select a representative area of the sample. It is recommended to acquire at least 3-5 images from different locations for each sample to avoid the influence of edge effects and local anomalies. The acquired SEM images should clearly show the microscopic morphological features of the rock, including key information such as grain boundaries, pore structure, and microcrack distribution.
[0031] Simultaneously, elemental composition analysis of the same target area was performed using an energy-dispersive X-ray spectrometer integrated with SEM. EDS analysis parameters were set as follows: acquisition time no less than 60 seconds to ensure data quality, adjustable to 60-120 seconds depending on elemental content; stable electron beam current; and spot size matching the SEM imaging area. EDS can acquire volume fraction data for eight major elements: O, Na, Mg, Al, Si, K, Ca, and Fe, with an accuracy of ±0.5% atomic percentage. Accurate electron beam focusing was crucial during analysis to avoid data distortion caused by sample drift. For each SEM imaging area, at least one full-spectrum EDS scan and analysis of 2 to 3 points were performed to verify data reliability.
[0032] S12. XRD Mineral Phase Analysis: Grind the rock sample into a fine powder with a particle size less than 75μm, pass it through a 200-mesh standard sieve, and grind for 15-30 minutes to ensure uniform particle size. Use an X-ray diffractometer, preferably a TD-3500 or Bruker D8Advance or equivalent, for crystal phase analysis. Scanning parameters: Scanning angle range 5-90°, covering the characteristic diffraction peaks of major minerals, step size 0.02°, which can be adjusted to 0.01-0.05° according to resolution requirements, scanning speed 2° / min, range 1-5° / min, slower speeds result in higher accuracy but longer scan time. After scanning, use Jade software for phase analysis. By analyzing the position and intensity of diffraction peaks, the composition and relative content data of major mineral phases such as quartz, feldspar, calcite, chlorite, and dolomite can be obtained with a quantitative accuracy of ±2%.
[0033] The three characterization techniques mentioned above form a complementary multi-scale analysis system: SEM provides spatial morphology information of microstructure at the μm scale; EDS provides elemental distribution information of chemical composition at the atomic scale; and XRD provides crystal structure information of mineralogical properties at the lattice scale. The combination of the three can comprehensively characterize the physical, chemical, and mineralogical properties of rocks, providing a sufficient data foundation for subsequent peak elastic performance prediction.
[0034] S13. Standard Mechanical Test Data Acquisition: To obtain the target value of the peak elastic energy of multimodal rocks and establish a reliable supervised learning model, the rock specimens were processed into standard cylindrical specimens. The specimen size was strictly controlled as follows: diameter 50 mm, height-to-diameter ratio of 2:1, in accordance with ISRM recommendations. After the specimens were prepared, they were dried at a constant temperature of 110±5°C for 24 hours to eliminate the influence of pore water on mechanical properties. The dried specimens should be cooled to room temperature in a desiccator before testing. A rock mechanics testing system was used to conduct uniaxial compression tests under strain-controlled loading conditions.
[0035] Throughout the experiment, high-precision sensors were used to synchronously acquire complete stress-strain curves at a sampling frequency of no less than 10 Hz; key parameters recorded included: peak stress. The maximum stress at rock failure; the elastic modulus E; the slope of the elastic segment of the stress-strain curve; based on experimentally obtained stress-strain data, the elastic strain energy density U is calculated as the target value for model training: ; In this embodiment, step S2 further includes: S21. Image Preprocessing and Standardization: The original SEM images are typically 1980×890 pixels, 2048×1536 pixels, or other non-standard sizes. Different devices and imaging parameters will produce different image sizes. To meet the input requirements of deep learning models, these images need to be cropped to a standard size. Cropping strategy: A sliding window method is used to crop each original image into multiple 224×224 pixel sub-images. There are three reasons for choosing the 224×224 size: This is the standard input size of mainstream pre-trained vision models such as Vision Transformer (ViT-B / 16), which can make full use of pre-trained weights to accelerate model convergence; This size can retain sufficient microstructural details such as individual particles, pores, cracks, etc., without causing excessive consumption of computational resources; It facilitates batch processing and GPU parallel computing, improving training efficiency.
[0036] Specific cropping method: For the original image of 1980×890 pixels, a sliding window with a step size of 224 pixels (no overlap) or a step size of 180 pixels (20% overlap) is used for cropping; no-overlap cropping can obtain approximately (1980 / 224)×(890 / 224)≈8×3=24 sub-images; overlapping cropping can obtain approximately (1980-224) / 180+1×(890-224) / 180+1≈10×4=40 sub-images; overlapping cropping can increase the number of samples, but the overlap rate should be avoided to prevent it from being too high, and it is recommended to be ≤30% to avoid data leakage.
[0037] S22. Data Augmentation Processing: Data augmentation is an important means to improve the generalization ability of deep learning models; considering the characteristics of rock microscopic images, this invention adopts the following augmentation methods: a. Contrast Enhancement: The contrast-limited adaptive histogram equalization (CLAHE) algorithm is adopted; parameter settings: clipLimit=2.0, which limits the contrast enhancement amplitude to avoid noise amplification, tileGridSize=(8, parallel 8), which divides the image into 8×8 sub-regions and equalizes them separately; CLAHE can significantly enhance the local contrast of the image, making microstructural features, especially low-contrast features such as pore boundaries and cracks, clearer, while avoiding the over-enhancement caused by traditional histogram equalization.
[0038] b. Geometric transformations: including vertical flip (50% probability), horizontal flip (50% probability), and center flip (rotation of 180°, 25% probability). These transformations can increase the diversity of samples and enable the model to learn rotation-invariant features. Since the microstructure of rocks is usually isotropic, these geometric transformations will not change the physical nature of the image except for layered rocks.
[0039] c. Optional enhancement operations: Slight rotation, ±5°, can be added as needed to avoid loss of edge information due to excessive angle; scaling, 0.95 to 1.05 times; brightness adjustment, ±10%; Gaussian noise, σ=0.01 to 0.02; however, it should be noted that these operations should be kept moderate to avoid over-enhancement that may cause image distortion or introduce unrealistic features.
[0040] d. Enhancement strategy: Apply the above enhancement operations to the training set images, and apply only necessary normalization processing such as CLAHE contrast enhancement to the validation and test sets, without applying random geometric transformations; this strategy can increase sample diversity during the training phase while maintaining the authenticity of the data during the evaluation phase.
[0041] Through the above enhancement operations, assuming the original acquisition of 180 SEM images, after cropping, approximately 1080 sub-images (180×6) are obtained. After data augmentation considering combinations such as flipping and rotation, the number can be expanded to approximately 5400 images (1080×5), significantly increasing the quantity and diversity of training samples. The core purpose of these image enhancement techniques is to make the model pay more attention to the intrinsic morphological features of the rock sample, such as pore shape, grain size distribution, and crack network topology, rather than secondary image features such as brightness, shooting angle, and position.
[0042] S23. Numerical Data Standardization: For EDS elemental composition data and XRD mineral phase data, due to significant differences in the dimensions and numerical ranges of different characteristics (e.g., Si content may be 30% to 50%, while the content of some trace elements may be <1%), standardization is required to eliminate the influence of dimensions. The Z-score standardization method is used. ; In the formula: x is the original feature value; μ is the mean of the feature in the training set; σ is the standard deviation of the feature in the training set; These are the standardized feature values, typically distributed in the range of -3 to 3, with a mean of 0 and a standard deviation of 1.
[0043] The specific steps for standardization are as follows: a. Calculate the mean μ and standard deviation σ for each feature on the training set; b. Use these statistics to standardize the corresponding features on the training, validation, and test sets; c. Save the μ and σ values so that the same standardization process can be applied to new samples in practical applications. Note that the standardized statistics μ and σ must be calculated only from the training set, and statistics from the entire dataset cannot be used, otherwise it will lead to data leakage and affect the objectivity of model evaluation; for the validation and test sets, μ and σ from the training set should be used for standardization.
[0044] S24. Dataset Partitioning: The dataset is randomly partitioned using a common deep learning approach: 70% training set, 15% validation set, and 15% test set. The training set (70%) is used for learning and optimizing model parameters (weights and biases). By repeatedly traversing the data and using the backpropagation algorithm to adjust the parameters, the prediction results gradually approach the true values. The validation set (15%) is used for hyperparameter tuning (such as learning rate, regularization coefficient, number of network layers, etc.) and preventing overfitting. Performance metrics (RMSE, R², etc.) are calculated after each training cycle, and an early stopping strategy is implemented by monitoring their changing trends. The validation set does not participate in the direct training of model parameters, but it indirectly affects the selection of hyperparameters and the control of the training process. The test set (15%) remains "hidden" throughout the entire model development process and does not participate in any training or hyperparameter tuning decisions. It is only evaluated once after the model is fully trained and the hyperparameters are fixed. Its performance metrics represent the expected performance of the model in real-world application scenarios. This partitioning ratio ensures sufficient training samples and provides enough independent data for model evaluation and hyperparameter tuning, thus ensuring the objectivity of model performance evaluation and the effectiveness of generalization ability testing.
[0045] Points to note when partitioning: a. Use random partitioning to ensure similar data distribution across the training, validation, and test sets, avoiding data bias in any subset; b. For multiple images of the same rock sample, all should be assigned to the same subset (training, validation, or test) to avoid data leakage; for example, if an SEM image of a granite sample is cropped into 10 sub-images, all 10 sub-images should be assigned to the same subset; c. Maintain approximately equal proportions of the three rock types (granite, sandstone, shale, etc.) in each subset to ensure the representativeness of stratified sampling; d. Record the random seed used for dataset partitioning to ensure reproducible results.
[0046] In this embodiment, step S3 further includes: S31. Feature Extraction of the Vision Transformer (ViT) Network: The Vision Transformer is an innovative model that applies the Transformer architecture to computer vision tasks. It abandons the local receptive field limitations of traditional convolutional neural networks and captures the global dependencies of images through a self-attention mechanism, achieving excellent performance in tasks such as image classification and object detection; Image Patch Embedding: ViT first transforms the input image X∈R... (H×W×C) Where H=224, W=224, C=1 is a grayscale image, divided into N fixed-size non-overlapping patches; for the ViT-B / 16 model, the patch size is 16×16 pixels, so N=(224 / 16)×(224 / 16)=14×14=196 patches; each patch is flattened into a vector with a length of 16×16×1=256.
[0047] The advantages of this block-based strategy are: a. It decomposes the image into a serialized input, allowing the Transformer architecture to directly process image data; b. Each patch contains local microstructural information such as individual mineral grains and pores, and the relationships between patches are learned through an attention mechanism; c. Compared to convolution operations, patch embedding is more computationally efficient.
[0048] Positional Encoding: Since the Transformer architecture itself does not contain positional information, meaning it is independent of the processing order, explicit positional encoding is needed to preserve the spatial relationship of patches within the image; ViT employs learnable positional encoding. ; In the formula, The initial input features represent the model. Positional encoding enables the model to understand the relative positional relationships between different patches. The functions of positional encoding are: a. to enable the model to distinguish patches in different positions, such as the pores in the upper left corner and the pores in the lower right corner, which have similar features but different positions; b. to encode the relative positional relationships between patches, such as adjacent patches possibly belonging to the same mineral grain; c. to preserve the spatial structure information of the image, which is crucial for understanding the microstructure of rocks.
[0049] Multilayer Transformer Encoder: Input Features z 0After L layers (L=12) for ViT-B / 16, the Transformer encoder performs feature extraction. Each Transformer encoder layer contains two main modules: Multi-Head Attention (MHA) and Feed-Forward Network (FFN), both employing Residual Connections and Layer Normalization. The computation process of the Multi-Head Attention mechanism is as follows: ; In the formula, This represents the output features of the l-th layer Transformer encoder. MLP stands for Multilayer Perceptron. This multilayer structure can progressively extract more complex feature representations, with each layer maintaining the integrity of information through residual connections.
[0050] The specific calculation of multi-head self-attention is as follows: ; In the formula, Q, and K, and V These are vectors representing queries, keys, and values. To control the scaling factor and prevent numerical instability, T It is the matrix transpose symbol, used for calculation Q and K The inner product correlation.
[0051] Feature aggregation: To obtain a global feature representation of the entire image, this invention employs average pooling. ; In the formula: The final extracted global feature vector of the SEM image has a dimension of 768. This vector integrates information from all 196 patches, represents the microscopic morphological features of the entire image, and can be used for subsequent cross-modal fusion and peak elastic performance prediction.
[0052] S32. Feature Extraction using Fully Connected Neural Network (FCNN): For both EDS elemental composition data and XRD mineral phase data, two types of numerical data, a fully connected neural network (FCNN) is used for feature extraction. FCNN is a basic yet efficient deep learning architecture that can learn higher-order nonlinear relationships and complex patterns in the input data through multiple layers of nonlinear transformations. The forward propagation process of FCNN can be represented as: ,and
[0053] In the formula, This represents the output of the neurons in layer 1. To connect the weight matrix, For bias vectors, For activation function, This represents the weighted input received by the neuron.
[0054] For 8-dimensional input EDS data, the FCNN structure is designed as follows: 8 neurons in the input layer → 64 neurons in the first hidden layer → 128 neurons in the second hidden layer → 64 neurons in the third hidden layer → 64 neurons in the output layer, i.e., the feature vector; the activation function is ReLU: f(x) = max(0, x). For 12-dimensional input XRD data, the FCNN structure is designed as follows: 12 neurons in the input layer → 64 neurons in the first hidden layer → 128 neurons in the second hidden layer → 64 neurons in the third hidden layer → 64 neurons in the output layer, i.e., the feature vector; the activation function is also ReLU.
[0055] The design considerations for the number of network layers and neurons are as follows: a. The input dimension is relatively low, with 8 or 12 dimensions, so the network does not need to be too deep, with 3 to 4 hidden layers being sufficient; b. The number of neurons in the hidden layers follows an "expansion-contraction" pattern: 64→128→64. First, expansion can increase expressive power, and then contraction can reduce dimensionality and extract the most critical features; c. The output layer is uniformly 64-dimensional, which facilitates fusion with the 768-dimensional features extracted by ViT, and the dimensions are aligned through subsequent linear transformations or attention mechanisms.
[0056] Backpropagation Algorithm: The parameters of FCNN, namely the weights W and biases b, are learned through the backpropagation algorithm. The core of backpropagation is to use the chain rule to calculate the gradient of the loss function L with respect to the parameters of each layer, and then use gradient descent to update the parameters. The loss function is calculated... L Weights W(l) and bias b(l) gradient: ; (9) ; (10) ; (11) In the formula, the error term As a central component, the calculation of weights and bias gradients is linked. Equation (11) defines how the error term is recursively calculated through network layers, and Equations (9) and (10) use this term to determine the parameter gradients. Specifically, Equation (11) establishes the backflow of the error signal, where the error of the first layer is calculated by the Hadamard product (⊙) of the backpropagation error of the next layer and the local derivative of the activation function. The calculated error term is then used.
[0057] In this embodiment, step S4 further includes: S41. Introduction to Optimal Transport Theory: Optimal Transport (OT) theory studies how to transform one probability distribution into another with minimal cost. It originated from Monge's "earthwork transportation problem" in the 19th century. In machine learning, OT theory provides a rigorous mathematical framework for measuring distances between different feature spaces, making it particularly suitable for multimodal data fusion scenarios. In the multimodal feature alignment problem of this invention, the feature extraction networks ViT and FCNN for three modalities (SEM, EDS, and XRD) are trained in different data distributions and feature spaces, leading to variations in the extracted feature vector F. SEM ∈R 768 F EDS ∈R 64 F XRD ∈R 64 They exist in different feature spaces; this inconsistency, or heterogeneity, of feature spaces can hinder the effective fusion of cross-modal information; OT theory can achieve soft alignment of feature spaces by learning the optimal transport mapping between modalities, so that the features of different modalities are semantically aligned.
[0058] S42. Constructing the Cost Matrix: The first step in OT theory is to define the transport cost, which is the cost of moving a feature point from the source space to the target space; this invention uses Euclidean distance as the cost metric. ; ; ; in, ,and ,and These represent the feature vectors of each mode. Since... f SEM The dimension is 768 and f SEM With a dimension of 64, it is necessary to first project the two to the same dimension, such as 128, through linear projection before calculating the distance, or use a kernel function for implicit alignment; i, j, and k are the sample indices of the three modalities, respectively.
[0059] S43. Entropy Regularized Optimal Transport: The classic Monge-Kantorovich OT problem has a high computational complexity of O(n³log n), making it impractical for large-scale data. This invention adopts an entropy regularization approach, adding an entropy term to the optimization objective, which allows for efficient solution using the Sinkhorn algorithm, reducing the complexity to O(n²). Entropy regularized OT optimization problem: ; Where: T is the transport matrix, with dimension R, where T(i, j) represents the "mass" or probability of transporting the i-th sample from the source distribution to the j-th sample of the target distribution; Π(μ, γ) is the set of transport plans that satisfy the marginal constraints, i.e., the row sums of T must equal the source distribution μ, Σ_j T(i, j) = μ_i, and the column sums must equal the target distribution γ; <T, C> represents the Frobenius inner product of the matrices, and this term represents the total transport cost; H(T) is the entropy of the transport matrix: H(T) = -ΣiΣjT(i, j)·[logT(i, j) - 1], and the entropy term encourages the transport matrix to be more dispersed or smooth, avoiding overly sparse solutions; λ is the regularization coefficient, with typical values ranging from 0.01 to 0.1, which controls the weight of the entropy term, The larger λ is, the smoother the solution but it may deviate from the optimal transport. The physical meaning of this optimization problem is: under the premise of satisfying the mass conservation constraint, i.e., the marginal constraint, find a transport plan T that minimizes the total transport cost <T, C>, while avoiding overly concentrated transport, which is achieved through the entropy regularization term.
[0060] S44. Sinkhorn algorithm for iterative solution: The entropy-regularized OT problem can be efficiently solved by the Sinkhorn-Knopp algorithm, which is based on the idea of matrix scaling. It approximates the optimal transport matrix by alternately updating two scaling vectors; first, define the Gibbs kernel matrix: ; Where K is the Gibbs kernel matrix, with dimension R NXN ; is the element-wise exponential operation; is the negative normalization of the cost matrix; then, by iteratively updating the two scaling vectors and : ; Where: and are the scaling vectors at the l-th iteration; and are the probability vectors of the source distribution and the target distribution, usually uniform distributions; the final transport matrix is: ; In the formula: diag(u) is the operation to convert vector u into a diagonal matrix; T is an n×n transfer matrix; the advantages of the Sinkhorn algorithm are: a. The computational complexity is O(n²), which is significantly reduced compared to the original OT's O(n³log n); b. The algorithm is stable and converges quickly, usually within 10 to 50 iterations; c. It is suitable for large-scale data; d. It is simple to implement, requiring only matrix multiplication and element-wise operations.
[0061] S45. Feature Space Alignment: After obtaining the transfer matrix T, the alignment of features from different modes can be achieved. Alignment means mapping the features of one mode to the feature space of another mode through a weighted combination of the transfer matrix. ; ; ; in, This represents the new feature matrix after aligning EDS features to the SEM feature space; note that: Since it is not a symmetric matrix, reverse alignment of SEM→EDS requires recalculating the transfer matrix. ≠T^ The reverse transfer matrix is obtained by rerunning the Sinkhorn algorithm by exchanging the source and target distributions.
[0062] In this embodiment, step S5 further includes: S51. The Necessity of Physical Mechanism Mapping: Although traditional deep learning models have excellent predictive performance, their internal decision-making process is a "black box," and the physical meaning of the model parameters, i.e., the weight matrix, is unclear, which limits their application in basic science. The peak elastic energy of rocks is jointly regulated by multiple physical mechanisms, including microstructure porosity, cracks, particle cementation state, elemental composition (the ratio of elastic reinforcing elements such as Si and Al to plastic reinforcing elements such as Ca and Mg), and mineral phases (the content of high-energy-storage minerals such as quartz and low-energy-storage minerals such as chlorite). If the model can explicitly quantify the contribution of these physical factors to the peak energy, it can not only improve the interpretability of the prediction, but also provide a deeper understanding of the mechanism for rock engineering.
[0063] Based on the microstructure features extracted from SEM images, the energy storage index of the microstructure is calculated: ; in: The microstructure energy storage index, Porosity For the fracture density parameter, As a weighting coefficient, the lower the porosity and fracture density, the stronger the rock's energy storage capacity.
[0064] Calculate the elemental energy storage index based on EDS elemental composition data: ; in, The content of elastic reinforcing elements (Si, Al) The content of plasticity-enhancing elements (Ca, Mg), and These are the corresponding weighting coefficients.
[0065] Based on XRD mineral phase data, the equivalent elastic modulus of the mineral phases was calculated according to the Voigt-Reuss-Hill mixing law, and then the mineral phase energy storage index was constructed. ; in, The mineral phase energy storage index, Let be the relative content of the i-th mineral. Let n be the elastic modulus of the mineral, and n be the total number of mineral types.
[0066] S52, Adaptive Weight Calculation: To accommodate the differences in characteristics among different rock types, an adaptive weighting calculation method is adopted: ; in, For the adaptive weights of mode i, The independent peak energy prediction score for this mode. This is a sensitivity parameter that controls the sensitivity of weight allocation.
[0067] S53. Feature Fusion and Peak Elastic Performance Prediction: The peak elastic performance exponents of the three modalities and adaptive weights are integrated and mapped end-to-end using a deep neural network: ; ; in, X The input vector is a weighted concatenation of the three modal energy storage indices. Let l be the activation function of the l-th layer. Let be the weight matrix of the l-th layer. , , For each modality, adaptive weights, This is the predicted peak energy value for rocks. The mapping function is a shallow neural network used to learn the nonlinear relationship and interaction between the three energy storage indices, capture the physical characteristics of the elastic energy storage stage in the stress-strain curve, and distinguish the essential differences between high-energy-storage and low-energy-storage rocks.
[0068] In this embodiment, step S6 further includes: a. Parameter Initialization: The parameter initialization of deep neural networks has a significant impact on training results; improper initialization may lead to vanishing gradients, exploding gradients, or convergence to local optima. For the weight parameters of the ViT network, using the weights of the pre-trained model as initial values is called transfer learning. The ViT-B / 16 model was pre-trained on the ImageNet dataset and learned rich general visual features such as edges, textures, and shapes, which are also effective for microscopic rock images. Using pre-trained weights can: speed up convergence, usually requiring only 1 / 5 to 1 / 10 of the original training time; improve final accuracy, especially when the amount of data is limited; and reduce the risk of overfitting.
[0069] b. Parameter optimization strategy: The model parameters are updated using the mini-batch stochastic gradient descent (SGD) algorithm; the parameter update formula is as follows: ; Where θ represents the model parameters. Let η be the parameter value at step t, ηt be the learning rate at step t, and η0 be the initial learning rate of 0.01. c. Learning rate decay strategy: A fixed learning rate may cause oscillations in the early stages of training and may not be able to accurately adjust parameters in the later stages; therefore, a learning rate decay strategy is used; exponential decay: ; In the formula: ηt is the learning rate of the t-th epoch; η0 is the initial learning rate, such as 0.01; β is the decay factor, typically 0.95 to 0.99, which is set to 0.95 in this invention; t is the current epoch number; this strategy causes the learning rate to decrease exponentially, with a larger learning rate in the early stage of training for rapid convergence, and a smaller learning rate in the later stage of training for fine adjustment.
[0070] e. Overfitting Prevention: To prevent the model from overfitting, which performs well on the training set but poorly on the test set, various regularization techniques are employed: L2 Regularization (Weight Decay): This involves adding an L2 norm penalty term for the weights to the loss function. ; In the formula: λ This is the regularization coefficient, typically ranging from 0.0001 to 0.01. This invention sets its value range to 0.01 to 0.1 to control the regularization strength; W ||² represents the sum of squares of all weight parameters; this term penalizes models with excessively large weights, encourages the model to learn smaller weights, and reduces model complexity; excessively large weights usually correspond to overfitting, where the model excessively memorizes noise and details from the training data.
[0071] f. Dropout technique: During training, the outputs of some hidden layer neurons that have been set to zero are randomly dropped with a probability of 0.5. This is equivalent to training multiple subnetworks with different structures and then integrating these subnetworks during testing, which can effectively prevent overfitting. Dropout is usually applied after fully connected layers. During testing, Dropout is turned off, and all neurons are used.
[0072] g. Early stopping strategy: After completing each epoch and traversing the training set once, calculate the root mean squared error (RMSE), coefficient of determination (COP), and R² evaluation metric on the validation set; RMSE: ; In the formula: n is the number of validation set samples; RMSE has the same dimensions as the peak elasticity exponent and can be intuitively understood as the magnitude of the average prediction error. For example, RMSE = 0.3 means that the average prediction error is approximately 0.3 peak elasticity exponent units; coefficient of determination R²: ; In the formula: SSres The sum of squared residuals represents the unexplained variation in the model; SStot The total sum of squares represents the total variation in the data; R² The range is [0, 1], and the closer to 1, the better the model fit. R² =1 indicates a perfect fit. R² =0 indicates that the model's predictive ability is no better than the mean prediction.
[0073] Closely monitor the changes in these metrics. If the RMSE on the validation set no longer shows a downward trend for 5 consecutive epochs, terminate training early to prevent overfitting and select the best-performing model. The benefits of stopping early include: avoiding overfitting and maintaining the model's generalization ability; saving training time by not waiting for the preset maximum number of epochs; automatically selecting the optimal model and saving the model weights when the validation set performance is best.
[0074] Repeat steps b to f until the training stopping condition is met, reaching the preset maximum number of training epochs of 100, or until the performance on the validation set no longer improves. Save the best model obtained from training, including the model structure and parameters, for subsequent prediction applications.
[0075] In this embodiment, step S7 further includes: S71. Model Application Process: For a new rock sample to be predicted, collect SEM, EDS, and XRD data according to method S1; perform preprocessing according to method S2, including image cropping, contrast enhancement, and numerical standardization; input the preprocessed data into the trained model, and perform forward propagation to obtain the predicted peak elastic energy index. .
[0076] S72. Feature Importance Analysis: After model training, the contribution of each feature to the prediction of peak elastic strain energy can be further analyzed to identify key control factors. This invention employs three feature importance assessment methods: a. Pearson Correlation Coefficient: measures the linear correlation between a feature and peak elastic strain energy; b. Spearman Rank Correlation Coefficient: measures the monotonic relationship between a feature and peak elastic strain energy, including nonlinear monotonic relationships. The calculation method involves first sorting the data into ranks, then calculating the Pearson correlation coefficient between the ranks. It is suitable for nonlinear but monotonic relationships; c. Mutual Information: measures any dependency between a feature and peak elastic strain energy, including nonlinear and non-monotonic relationships. Mutual information is a concept in information theory, requiring no assumptions about the form of the relationship, and has the widest applicability.
[0077] S73. Key Control Factor Identification: The contribution of each feature to the prediction of peak elastic strain energy is comprehensively evaluated by three measurement methods: Pearson correlation coefficient, Spearman rank correlation coefficient, and mutual information value. The key control factors affecting the peak elastic strain energy of rocks and their effects are identified. The model can automatically adjust the contribution of each physical factor through an adaptive weighting mechanism to achieve generalized prediction of peak elastic strain energy across lithologies.
[0078] S74. Cross-lithological generalization ability: The multimodal deep learning model of this invention has good cross-lithological generalization ability; by training on three representative rocks, granite, sandstone and shale, the model can automatically learn the characteristic patterns and peak elastic strain energy laws of different rocks; for new rock types such as limestone and basalt, the model can dynamically adjust the relative importance of each physical factor through an adaptive weighting mechanism to achieve accurate prediction of peak elastic strain energy under different rocks and confining pressures.
[0079] II. Experimental Results
[0080] The results are shown in Table 1. Table 1 provides a comparison of the actual peak elastic energy values and model predictions for different rock types. The following conclusions can be drawn from the data in Table 1: The intelligent prediction method for peak elastic properties of rocks proposed in this invention, which integrates physical constraints and multimodal deep learning, demonstrates excellent prediction accuracy across various rock types. For nine test samples, including blue sandstone, red sandstone, and granite, the absolute value of the relative error between the model's predicted values and the actual values is controlled within 5%, with the minimum relative error being only 0.7% and the maximum relative error being 4.8%.
[0081] Specifically, the relative error range for blue sandstone samples was -2.6% to 4.3%, for red sandstone samples it was -3.6% to 4.8%, and for granite samples it was -4.2% to 3.5%. The predicted values for all samples showed a high degree of agreement with the actual values, verifying the accuracy and generalization ability of the method under cross-lithological conditions. This method can effectively capture the nonlinear mapping relationship between the microstructure and macroscopic energy storage performance of different rock types, providing reliable technical support for the non-destructive and rapid prediction of peak elastic properties of rocks.
[0082] Table 1. Comparison and analysis of actual and model-predicted peak elastic energy values for different rock types.
[0083] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for intelligent prediction of peak elastic properties of rocks that integrates physical constraints and multimodal deep learning, characterized in that, Includes the following steps: S1. Obtain scanning electron microscope images, energy dispersive X-ray spectroscopy elemental composition data, and X-ray diffraction mineral phase composition data of the rock sample to be tested. S2. The scanning electron microscope images are cropped and augmented, and the elemental composition data (EDS) and mineral phase composition data (XRD) from the energy dispersive X-ray spectrometer are standardized to construct a multimodal dataset. S3. A visual transformer network is used to extract the microscopic morphology features of scanning electron microscope images. A fully connected neural network is used to extract the elemental composition data of energy dispersive X-ray spectrometer and the mineral phase composition data of X-ray diffractometer, respectively. The elastic strain energy density theory and the mineral elastic modulus mixing law are used as physical constraints and embedded into the loss function. S4. Calculate the cost matrix between different modal features, adopt the optimal transfer method with entropy regularization, and solve the transfer matrix iteratively through the Sinkhorn algorithm to achieve soft alignment of the feature spaces of three modal features: scanning electron microscope, energy dispersive X-ray spectrometer, and X-ray diffractometer. S5. A multi-head attention mechanism is used to integrate the aligned cross-modal features, construct the microstructure energy storage index, elemental energy storage index and mineral phase energy storage index, and dynamically adjust the contribution of different modes through an adaptive weight calculation method to establish a multi-scale physical mechanism mapping relationship between micromorphology, chemical composition and mineralogical properties. S6. The model training is driven by a composite loss function that integrates physical mechanisms, and the model parameters are updated using a mini-batch stochastic gradient descent algorithm, combined with regularization techniques to prevent overfitting. S7. Input the multimodal data of the rock sample to be predicted into the trained model, and output the predicted value of the peak elastic performance of the rock.
2. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, Step S1 also includes: obtaining a standard cylindrical rock sample, conducting a uniaxial compression test using a rock mechanics testing system, simultaneously acquiring a complete stress-strain curve, and calculating the elastic strain energy density U as the target value for model training based on the peak stress σp and elastic modulus E according to the formula U=σp² / 2E.
3. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, In step S1, the scanning electron microscope imaging scale is 5 μm, the elements obtained by the energy dispersive X-ray spectrometer include at least O, Na, Mg, Al, Si, K, Ca, and Fe, the X-ray diffractometer scanning angle range is 5-90°, and the particle size of the ground rock sample is less than 75 μm.
4. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, In step S3, the visual transformer network extracts microscopic morphological features, including: The input image X is divided into fixed-size patches, and each patch is mapped to a D-dimensional feature space through linear projection, resulting in... ; Add position encoding to each patch: A learnable positional encoding matrix; Feature extraction is performed using an L-layer Transformer encoder, with each layer containing a multi-head self-attention mechanism and a feedforward neural network. The computation process is as follows: z' l = MHA(LN(z) {l-1})) + z_{l-1}; z_l = MLP(LN(z'_l)) + z'_l; Where z is the feature vector of each layer of the Transformer encoder; MHA is multi-head self-attention; LN is layer normalization; MLP is multilayer perceptron; and l is the layer index.
5. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, Step S3, in which the fully connected neural network extracts deep features, includes: For EDS and XRD data, construct multi-layer fully connected neural networks respectively; the forward propagation process can be represented as: ,and ; In the formula, l The first term of the neural network l layer, Indicates the first l The output of layer neurons, To connect the weight matrix, For bias vectors, For activation function, This represents the weighted input received by the neuron; This represents the number of nodes in the current layer. This represents the number of nodes in the previous layer. Update the parameters using the backpropagation algorithm, and calculate the gradients of the loss function L with respect to the weights and biases: ; ; ; In the formula, For the weight gradient, For the bias gradient, Here, T represents the error term, and T is the transpose operation. It is the reciprocal of the activation function.
6. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, Step S4 also includes: Calculate the Euclidean distance between different modal features as the transportation cost: ; ; ; in, ,and ,and These represent the feature vectors of each mode; Let be the cross-modal distance between the i-th SEM sample and the j-th EDS sample; Let be the cross-modal distance between the i-th SEM sample and the k-th XRD sample; Let be the cross-modal distance between the j-th EDS sample and the k-th XRD sample; Solve the optimal transportation problem in the form of entropy regularization: min_{T∈Π(μ, γ)} <T, C> - λH(T); Where T is the transfer matrix, μ is the marginal probability distribution of the source mode, γ is the marginal probability distribution of the target mode, and H(T) is the entropy regularization term, where H(T) = -Σ_{i,j} T_{ij} (log T_{ij} - 1), T ij Let λ be the transmission amount from the i-th source sample to the j-th target sample, and λ be the regularization parameter. Achieve feature alignment through the transportation matrix: ; ; ; in, This indicates that the characteristics of the EDS mode are transmitted through the optimal transfer matrix. Transform to SEM modal space, The optimal transfer matrix for both ESD and SEM modes needs to be recalculated through the reverse OT problem, rather than being directly transposed.
7. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, Microstructure energy storage index in step S5 Calculated using the following formula: ; in, The microstructure energy storage index, Porosity For the fracture density parameter, As a weighting coefficient, the lower the porosity and fracture density, the stronger the rock's energy storage capacity. Element energy storage index Calculated using the following formula: ; in, C i + The content of elastic reinforcing elements Si and Al, C j - The content of plasticity-enhancing elements Ca and Mg, β i + and β j - These are the corresponding weighting coefficients; Mineral phase energy storage index The equivalent elastic modulus of mineral phases, calculated according to the Voigt-Reuss-Hill mixing law, is obtained through the following formula: ; in, The mineral phase energy storage index, Let be the relative content of the i-th mineral. Let n be the elastic modulus of the mineral, and n be the total number of mineral types.
8. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, The method for calculating the adaptive weights in Step S5 is: ; in, For the adaptive weights of mode i, The independent peak energy prediction score for this mode. This is a sensitivity parameter that controls the sensitivity of weight allocation. In Step S5, feature fusion is performed through an end-to-end mapping by a deep neural network, expressed as: ; ; in, X The input vector is a weighted concatenation of the three modal energy storage indices. Let L be the activation function of the Lth layer. For the first L The weight matrix of the layer, , , For each modality, adaptive weights, This is the predicted value of the peak elastic energy of the rock.
9. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, The composite loss function in Step S6 is: ; For mean square error loss, ; Physical consistency penalty item , ; This is a soft constraint penalty term for monotonicity. ; in, m For the sample size, To predict the peak energy value, These are the weighting coefficients for each constraint term. In Step S6, the small-batch stochastic gradient descent algorithm is used for parameter optimization, and the parameter update formula is: ; ; in, θ For model parameters, Let be the parameter value at step t. ηt Let be the learning rate at step t. The gradient of the loss function with respect to the parameters. The initial learning rate is 0.01; The attenuation coefficient is... T This is the decay period.
10. The intelligent prediction method for peak elastic properties of rocks integrating physical constraints and multimodal deep learning according to claim 1, characterized in that, Step S7 also includes feature importance analysis: Use three measurement methods, namely the Pearson correlation coefficient, the Spearman rank correlation coefficient, and the mutual information value, to evaluate the contribution of each feature to the peak elastic energy prediction and identify the key control factors.