Microstructure reconstruction method and system based on rock image deep learning

By extracting and fusing the global and local features of rock images, generating multi-scale joint embedding vectors, and reconstructing grain boundary profiles using generator and discriminator frameworks, the problem of loss of rock mesostructure reconstruction accuracy in the prior art is solved, and the coordinated evolution of high-accuracy reconstruction and physical attribute optimization is achieved.

CN120235984AActive Publication Date: 2025-07-01NORTHEASTERN UNIV CHINA +1

Patent Information

Application Number
CN202510726700.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The prior art has the problem of lacking the global correlation characteristics that take into account the mineral distribution and the local detailed characterization of grain boundary microstructures in the feature extraction process in the rock mesoscopic reconstruction, resulting in the accuracy loss of the reconstruction results on the macro-microscopic scale. At the same time, the existing methods have not yet established an end-to-end mapping framework between rock mesostructural image data and key physical parameters, resulting in the inability to evolve in concert with the optimization of physical attributes.

Method used

By acquiring rock images, chunking and mapping them to the space of the specified dimensions, global and local features are extracted, fused into multi-scale joint embedding vectors, and adjusting the embedding vectors to generate reconstructed images. The cross-modal attention mechanism and generator discriminator framework are adopted to achieve accurate reconstruction of grain boundary contours.

Benefits of technology

The accuracy of rock mesostructure reconstruction is improved, the completeness of information of the reconstruction results on multiple scales is ensured, and the coordinated evolution of mesostructure reconstruction and physical attribute optimization is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235984A_ABST
    Figure CN120235984A_ABST
Patent Text Reader

Abstract

The invention provides a microstructure reconstruction method and system based on rock image deep learning, and relates to the technical field of image processing. The method comprises the following steps: acquiring an image of a rock to be reconstructed; the method comprises the following steps of: partitioning an image and mapping the image to a space with a specified dimension to obtain a plurality of subspaces, and obtaining a global feature according to a long-range dependency relationship among the plurality of subspaces; crystal boundaries and microfractures in the image are extracted to obtain local features, the crystal boundaries are boundaries of crystals in the image, the crystals are constituent parts of the rock to be reconstructed, and the microfractures are gaps between every two adjacent crystals; fusing the global features and the local features to obtain a multi-scale joint embedding vector; and reconstructing the contour of the crystal boundary according to the multi-scale joint embedded vector to obtain a reconstructed image. When the image of the rock to be reconstructed is reconstructed, information of the image on multiple scales is concerned, and the accuracy of the obtained reconstructed image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method and system for reconstructing a microstructure based on deep learning of rock images. Background Art

[0002] The mesostructure of rocks has a significant impact on their macroscopic mechanical behavior. At present, the study of rock mesostructure mainly relies on physical experiments and empirical models based on statistical characteristics. Physical experiments, such as the use of electronic equipment such as computed tomography (CT) and scanning electron microscope (SEM) to scan the rock mesostructure, and perform corresponding experiments based on the scanning results. The empirical model mainly uses deep learning technology, which has been increasingly applied in the field of rock mesostructure reconstruction with the development of deep learning technology.

[0003] However, the above research methods still have significant technical bottlenecks: first, deep learning technology usually uses convolutional neural networks (CNN). In the feature extraction process, the empirical model based on CNN lacks the technical means to take into account both the global correlation characteristics of mineral distribution and the local detail characterization of grain boundary microstructure, resulting in accuracy loss in the reconstruction results at both macro and micro scales; second, the existing methods have not yet established an end-to-end mapping framework between rock mesostructure image data and key physical parameters (such as mineral content, grain boundary roughness, etc.), resulting in the inability to achieve coordinated evolution of the mesostructure reconstruction process and physical property optimization, and the accuracy of mesostructure reconstruction cannot be guaranteed.

[0004] In view of this, the present invention is proposed. Summary of the invention

[0005] In order to solve the above problems, this application provides a method and system for reconstructing microstructures based on deep learning of rock images, which ensures the accuracy of the reconstruction results by fusing information at multiple scales of the image. The technical solution of this application is as follows: In the first aspect, a method for reconstructing a microstructure based on deep learning of rock images is provided, comprising: Acquire an image of the rock to be reconstructed; After dividing the image into blocks and mapping them to a space of a specified dimension, a plurality of subspaces are obtained, and a global feature is obtained according to a long-range dependency relationship between the plurality of subspaces; Extracting grain boundaries and microcracks in the image to obtain local features, wherein the grain boundaries are boundaries of crystals in the image, the crystals are components of the rock to be reconstructed, and the microcracks are gaps between two adjacent crystals; Fuse the global feature and the local feature to obtain a multi-scale joint embedding vector; Reconstruct the contour of the grain boundary according to the multi-scale joint embedding vector to obtain a reconstructed image.

[0006] In a possible implementation manner, the reconstructing the contour of the grain boundary according to the multi-scale joint embedding vector to obtain a reconstructed image includes: Obtain the contour offset of the grain boundary; Adjust the multi-scale joint embedding vector according to the contour offset to obtain an input image; Obtain the reconstructed image according to the input image and a standard image, where the standard image refers to an image showing the mineral composition, crystal arrangement, and pore distribution of the rock to be reconstructed.

[0007] In a possible implementation manner, the adjusting the multi-scale joint embedding vector according to the contour offset to obtain an input image includes: , where, is the input image, is the multi-scale joint embedding vector, is the image of the rock to be reconstructed, is the preset weight corresponding to the multi-scale joint embedding vector, F( ) is a deformable convolution function, and the deformable convolution function is used to adjust the multi-scale joint embedding vector according to the contour offset to obtain an input image, is the contour offset.

[0008] In a possible implementation manner, the obtaining the reconstructed image according to the input image and a standard image includes: Use a generator to obtain at least one candidate image according to the input image; Use a discriminator to obtain the similarity between the candidate image and the standard image; Select a candidate image with the highest similarity as the reconstructed image.

[0009] In a possible implementation manner, after obtaining the image of the rock to be reconstructed, the method further includes: Use a Gaussian filtering algorithm to filter out interference signals in the image; Use a histogram equalization method to enhance the contrast of the image.

[0010] In a second aspect, a mesoscopic structure reconstruction system based on deep learning of rock images is provided. The system is used to execute the above-mentioned mesoscopic structure reconstruction method based on deep learning of rock images, and includes: The data acquisition layer is used to acquire the image of the rock to be reconstructed; The data extraction layer includes a first extraction sub-layer and a second extraction sub-layer arranged in parallel. The first extraction sub-layer is used to extract the global features of the image, and the second extraction sub-layer is used to extract the local features of the image; The data processing layer is respectively connected to the first extraction sub-layer and the second extraction sub-layer, and is used to fuse the global features and the local features to obtain the multi-scale joint embedding vector; The data generation layer is used to obtain the reconstructed image according to the multi-scale joint embedding vector.

[0011] In a possible implementation manner, the first extraction sub-layer uses a ViT model to extract the global features of the image, and the ViT model is used to perform the following operations: Linear embedding and tokenization: After dividing the image into blocks, map it to a space of a specified dimension to obtain multiple sub-spaces; Position encoding: After capturing the spatial order of multiple sub-spaces, generate a spatial sequence; Weight adjustment: According to the long-range dependence relationship between multiple sub-spaces in the spatial sequence, adjust the weight of each sub-space; Extract global features.

[0012] In a possible implementation manner, the second extraction sub-layer uses a MobileNetV3 model to extract the local features of the image; The MobileNetV3 model at least includes a separable convolutional network, a non-linear activation network, and a residual network.

[0013] In a possible implementation manner, the data processing layer uses a cross-modal attention mechanism to fuse the global features and the local features; The calculation formula of the cross-modal attention mechanism is: , where G represents the global feature, L represents the local feature, both G and L are dimensional vectors, A is the multi-scale joint embedding vector, is the function of the cross-modal attention mechanism.

[0014] In a possible implementation manner, the system further includes a model training layer, and the model training layer is used to train the reconstruction model, and the reconstruction model is composed of the data acquisition layer, the data extraction layer, the data processing layer, and the data generation layer; The model training layer includes: The first training sub-layer is used to construct a loss function, and the loss function is used to constrain the reconstruction model; The second training sub-layer evaluates the accuracy of the reconstructed image output by the reconstruction model using the five-fold cross-validation method; The third training sub-layer: adjusts the parameters of the reconstruction model and selects the best set of parameters as the parameters of the reconstruction model.

[0015] The technical solution provided by the embodiment of the present application can achieve the following technical effects.

[0016] (1) The present application extracts the global features and local features of the image, and then fuses the global features and local features to obtain data involving multiple-scale information, that is, obtains a multi-scale joint embedding vector. Then, the obtained multi-scale joint embedding vector is adjusted to generate a reconstructed image. It can be seen that when reconstructing the image, the present application pays attention to the information on multiple scales of the image, thereby improving the accuracy of the obtained reconstructed image.

[0017] (2) The present application is also provided with a reconstruction model, which is composed of multiple layers, and each layer contains multiple network models. Then, through the mutual cooperation between the multiple network models, technical support is provided for paying attention to and fusing the information on multiple scales of the image, thereby ensuring the accuracy of the obtained reconstructed image.

[0018] (3) The present application is also provided with a model training layer, which trains and optimizes the reconstruction model through the model training layer to ensure the accuracy of the reconstructed image obtained when the reconstruction model reconstructs the image. Description of the Drawings

[0019] In order to more clearly illustrate the technical solution of the embodiment of the present application, the drawings required to be used in the description of the embodiment of the present application will be briefly introduced below.

[0020] Figure 1 is a flowchart of a mesoscopic structure reconstruction method based on deep learning of rock images provided by the embodiment of the present application.

[0021] Figure 2 is a structural diagram of a mesoscopic structure reconstruction system based on deep learning of rock images provided by the embodiment of the present application.

[0022] Figure 3 is a structural diagram of the data extraction layer and the data processing layer in the system embodiment of the present application.

[0023] Figure 4 is a structural diagram of the first extraction sub-layer in the system embodiment of the present application.

[0024] Figure 5 is a flowchart of establishing a training set in the system embodiment of the present application.

[0025] Figure 6 is a structural diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners

[0026] Exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0027] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such use can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "including" and its variants should be interpreted as open-ended terms meaning "including but not limited to".

[0028] As Figure 1 shown, the present application provides a mesoscopic structure reconstruction method based on deep learning of rock images, and the method mainly includes the following steps S101 to step S105.

[0029] Step S101, obtaining an image of the rock to be reconstructed.

[0030] The rock to be reconstructed refers to the rock to be subjected to image reconstruction, and the image of the rock to be reconstructed refers to: an image obtained by scanning the rock to be reconstructed using an electronic device such as CT or SEM. This image not only shows the components of the rock to be reconstructed - crystals, but also shows the boundaries of the crystals - grain boundaries, and even includes the gaps between the crystals - microcracks.

[0031] After obtaining the image, preprocess the image, including: (1) Denoising processing: Use the Gaussian filtering algorithm to filter out the interference signals in the image. The specific calculation formula is: , where is the pixel coordinate of the image, is the Gaussian kernel function, , is the Gaussian kernel standard deviation, which is used to control the filtering intensity, determines the Gaussian kernel size. By adjusting the values of and , the noise interference in the image can be removed while retaining the mesoscopic structure of the image. The mesoscopic structure mainly refers to the above-mentioned crystals, grain boundaries, microcracks, and the shapes and sizes of each structure, etc.

[0032] (2)Enhancement processing: The histogram equalization method is adopted. By adjusting the grayscale histogram of the image, the dynamic range of the image grayscale values is expanded, the contrast of the image is enhanced, and the structures such as crystals and grain boundaries in the image become clearer. In a specific example, the initial grayscale value range of the image is [a, b], denoted by After adopting histogram equalization, the grayscale value range of the image is expanded to [0, 255], and we get: , where is the image after expanding the image grayscale values, is the number of pixels with the initial grayscale value of the image being , is the cumulative grayscale histogram.

[0033] After preprocessing the image, it enters steps S102 and S103 respectively, and further processing is continued on the image.

[0034] Step S102: After dividing the image into blocks and mapping them to a space of a specified dimension, multiple subspaces are obtained. Global features are obtained based on the long-range dependence relationship between the multiple subspaces.

[0035] In this embodiment, the image can be divided into multiple blocks of the same size, or the image can be divided into multiple blocks of different sizes, as long as the obtained blocks do not overlap. When dividing, the tiling division method can be adopted, that is, each area on the image is divided in a certain order; of course, the classification division method can also be adopted for division, that is, different types of areas on the image are distinguished according to their types or attributes. For example, the area where a crystal is located is divided into one block, the area where the grain boundary between two adjacent crystals is located is divided into another block, and other blank areas are used as another separate block.

[0036] Based on the obtained multiple blocks, first, each block is converted into a vector, and then the vector is linearly projected into a space of a specified dimension to obtain a subspace corresponding to each block. Secondly, analyze the spatial position of each subspace in the actual scene, determine the spatial order of each subspace based on this position information, and then output the multiple subspaces as a spatial sequence according to this spatial order. For example, output the multiple subspaces as a spatial sequence in the spatial order from inside to outside, from top to bottom.

[0037] Based on the generated spatial sequence, using this spatial sequence as the current spatial sequence, retrieve multiple other spatial sequences before and after the current spatial sequence. The retrieved other spatial sequences and the current spatial sequence belong to the same rock to be reconstructed. Then, compare the dependency relationships between the current spatial sequence and the multiple other spatial sequences before and after. For example, if the crystals of the rock to be reconstructed change after the current spatial sequence, then in the other spatial sequences after the current spatial sequence, the crystals of the rock to be reconstructed maintain the same change. Therefore, by comparing the current spatial sequence with the multiple other spatial sequences before and after it, the dependency relationships between multiple sub-spaces within a certain time range can be obtained. This dependency relationship is also called a long-range dependency relationship. Finally, use the self-attention mechanism to adjust the weights of each sub-space in the spatial sequence. For example, assign higher weights to the sub-spaces with significant long-range dependency relationships, and then adjust the weights of each sub-space in the current spatial sequence to ensure that each sub-space more accurately reflects the actual changes of the rock to be reconstructed.

[0038] Through the adjustment of weights, strengthen the sub-spaces that have a significant impact on global features, and at the same time reduce the interference of other sub-spaces. Therefore, after adjusting the weights of the sub-spaces, extract the global features of the image to obtain accurate global features, which contain the overall semantic information of the image, such as semantic information about the shape and size of the rock to be reconstructed.

[0039] Step S103, extract the grain boundaries and microcracks in the image to obtain local features.

[0040] In this embodiment, use a deep learning edge detection algorithm to perform operations such as convolution, weight adjustment, and feature fusion on the image. After a series of operations, the feature expression ability of high-frequency details such as grain boundaries and microcracks in the image is enhanced. Then, determine the positions of high-frequency details such as grain boundaries and microcracks, and extract local features containing information such as grain boundaries and microcracks.

[0041] The deep learning edge detection algorithm can be any one of the rich feature hierarchies (RCF) detection algorithm, the Holistically-Nested Edge Detection (HED) algorithm, and the lightweight MobileNetV3 model algorithm. In this embodiment, taking the lightweight MobileNetV3 model algorithm as an example, the MobileNetV3 model algorithm is based on depthwise separable convolution and inverted residual structures. The model first uses 3×3 convolution in the shallow layer to capture low-order edge features, such as grain boundaries and microcracks, and then refines the edge features layer by layer through multiple inverted residual blocks. At the same time, introduce the self-attention mechanism to enhance the response of important channels, and also use activation functions to retain high-frequency details, finally achieving accurate detection of high-frequency details, so as to obtain accurate local features.

[0042] Step S104: Fuse the global features and local features to obtain a multi-scale joint embedding vector.

[0043] By combining features at different levels and scales such as global features and local features, a more expressive joint embedding vector is generated, which is also called a multi-scale joint embedding vector.

[0044] In this embodiment, the specific fusion process mainly includes: First, adjust the size of the local features to match the global features, then unify the number of channels of the global features and local features, then flatten the global features and local features, splice the flattened global features and local features to form a long vector, and then adjust the weights of each vector through the attention mechanism to obtain the multi-scale joint embedding vector. The calculation formula is: , where G represents the global features, L represents the local features, means that both G and L are dimensional vectors, A is the data obtained after fusing the global features and local features. Since this data contains features in multiple scales such as global and local, and both global features and local features are represented by vectors, this data is also called a multi-scale joint embedding vector.

[0045] It can be seen from this that the multi-scale joint embedding vector can not only describe the overall semantic information of the image, but also capture the high-frequency details in the image.

[0046] Step S105: Reconstruct the contour of the grain boundary according to the multi-scale joint embedding vector to obtain a reconstructed image.

[0047] First, obtain the contour offset of the grain boundary. This contour offset can be obtained by pre-training a deformable convolution function, which is the error existing in the deformable convolution function due to its own parameters. The deformable convolution function is the key function for reconstructing the grain boundary contour. Therefore, it is necessary to adjust the contour of the grain boundary based on this contour offset. The specific calculation formula is: , where, is the multi-scale joint embedding vector, is the image of the rock to be reconstructed, is the preset weight corresponding to the multi-scale joint embedding vector, F( ) is the deformable convolution function, and the deformable convolution function is used to adjust the multi-scale joint embedding vector according to the contour offset to obtain the input image , is the contour offset.

[0048] Then, use the input image as the input of the generator, and the generator receives the input image At least one candidate image is output, and then the discriminator retrieves the standard image. The standard image refers to an image obtained by scanning the rock to be reconstructed using electronic devices such as CT and SEM, and after reconstruction or other adjustment means, clearly showing information such as the mineral composition, crystal arrangement, and pore distribution inside the rock to be reconstructed. Based on the obtained candidate image and the standard image, the discriminator judges the similarity between the candidate image and the standard image. Finally, the candidate image with the largest similarity value is used as the reconstructed image. By this way, the obtained reconstructed image is constrained to make the mesoscopic structure of the rock to be reconstructed in the reconstructed image more in line with the actual situation. In addition, when there is only one candidate image, the discriminator uses the candidate image as the reconstructed image and outputs the similarity difference between the reconstructed image and the standard image, so as to remind the user to retrain and optimize the reconstruction model of the execution entity when the similarity difference exceeds the preset difference.

[0049] It should be noted that the execution entity can be any one of a processor, a server, a terminal, and a component, and the reconstruction model is an algorithm model used to implement the reconstruction method, which is composed of multiple network models.

[0050] Based on the mesoscopic structure reconstruction method based on deep learning of rock images provided by the above method embodiments, based on the same inventive concept, an embodiment of the present application also provides a mesoscopic structure reconstruction system based on deep learning of rock images, and this system is used to execute the above reconstruction method.

[0051] As Figure 2 shown, the present application provides a mesoscopic structure reconstruction system based on deep learning of rock images. From top to bottom, this system includes a data acquisition layer, a data extraction layer, a data processing layer, and a data generation layer.

[0052] Data acquisition layer: It is mainly used for information interaction with external terminal devices to obtain images of the rock to be reconstructed. Since the image of the rock to be reconstructed in this embodiment refers to an image obtained by scanning the rock to be reconstructed using electronic devices such as CT or SEM, the above terminal device can be an electronic device with a scanning function such as CT and SEM, or an electronic device used to store images obtained by scanning such as CT and SEM.

[0053] Data extraction layer: The data extraction layer includes a first extraction sub-layer and a second extraction sub-layer arranged in parallel.

[0054] Among them, the first extraction sub-layer is mainly composed of a ViT (Vision Transformer) model. The ViT model is a deep learning model for image processing. As Figure 3 and Figure 4 shown, the ViT model is mainly used to implement the following operations: Linear embedding marker: Divide the image into N non-overlapping blocks P i , where i = 1, 2, …, N. Assume that the size of each block is , then the block P i can be vectorized into x i . Specifically: Project x i onto the -dimensional space to obtain the subspace , where is the projection weight matrix, and the dimension of is . Its function is to convert the two-dimensional vector of the block P i to the specified -dimensional space. is the bias vector, which also represents the -dimensional vector, and it is used to adjust the offset of the vector so that the ViT model can better learn the feature representation of the block.

[0055] Position encoding: Capture the spatial order of each subspace through learnable position embeddings . The calculation formula is: . Before capturing the spatial order of the subspace, a set of learnable classification markers needs to be added. Based on this set of classification markers , determine the basic spatial order z0, and then locate the spatial order of each subspace on the basis of this basic spatial order z0, so as to obtain the spatial order corresponding to each subspace . And so on, so as to obtain the spatial orders corresponding to all subspaces, and then output multiple subspaces as a spatial sequence according to this spatial order.

[0056] Weight adjustment: Use the self-attention mechanism to analyze the long-range dependence relationship between multiple subspaces in the spatial sequence, and adjust the weights of each subspace according to this long-range dependence relationship. In a specific example, the query vector, key vector, and value vector of the self-attention mechanism are represented by Q, K, and V respectively, and all three are set to be -dimensional vectors. Then the calculation formula of the self-attention mechanism is: , where K T represents the transpose of the key vector.

[0057] If the self-attention mechanism is executed in parallel h times, then we get: , where is the weight matrix, and the dimension supported by the self-attention mechanism is , the self-attention mechanism is used to project the query vector Q, key vector K, and value vector V into different subspaces respectively to learn the feature relationships from different angles, such as the above The dimension of , then the outputs of multiple dimensions can be fused. Therefore, the self-attention mechanism also becomes the multi-head self-attention mechanism, which can fuse multiple outputs. In this embodiment, the self-attention mechanism is executed 8 times in parallel, that is, h = 8. Then the self-attention mechanism can simultaneously capture the long-range dependence relationships between subspaces from 8 different subspaces, so as to more comprehensively obtain the spatial distribution and correlation information of the crystal in the image.

[0058] In this embodiment, normalization layers are arranged at both the front end and the back end of the self-attention mechanism. The idle sequence is converted into the data format supported by the self-attention mechanism through the normalization layer at the front end, and the data output by the self-attention mechanism is converted into the data format supported by the MLP Head through the normalization layer at the back end.

[0059] Extracting global features: The MLP Head is used to extract the reconstruction features of the image. The MLP Head consists of a series of fully connected layers (FCLs), and each layer is connected by an activation function to form a multi-layer perceptron (MLP). The spatial distribution and correlation information of the crystal in the image are enhanced through the multi-layer perceptron, and then the global features of the image are extracted on this basis. Specifically, the structure of the MLP Head is: , Among them, , are both preset weights of the fully connected layer, and , , refers to the dataset from the outside world, such as the datasets from corpora such as ImageNet and Wikipedia. b1 and b2 are both preset biases, is the activation function.

[0060] It can be seen from this that the work implemented by the first extraction sub-layer through the ViT model is as follows: First, an image is obtained. After being processed by linear embedding and tokenization, the image is divided into multiple blocks, and each block is mapped into a space of a specified dimension through linear projection to obtain the corresponding subspace. Second, positional encoding adds learnable positional information to each subspace to capture the spatial order relationship of the subspaces and generate a spatial sequence containing positional information. Thirdly, the self-attention mechanism is used to process the spatial sequence, including: first normalizing the input spatial sequence, and then calculating the correlation weights of multiple subspaces in the spatial sequence through the self-attention mechanism to mine the long-range dependence relationship of the subspaces, so as to adjust the weights of the subspaces according to the long-range dependence relationship. Finally, the image with adjusted weights is normalized again and sent to the MLP Head for global feature extraction. The output of the MLP Head is fused with the front-end input through residual connection to systematically extract the global features of the image.

[0061] As Figure 2 and Figure 3 shown, the second extraction sub-layer is mainly composed of a lightweight MobileNetV3 model. The MobileNetV3 model achieves a balance between lightweight and high performance by integrating a depthwise separable convolutional network, a non-linear activation network, a squeeze-and-excitation (SE) network, and a residual network connection. In this embodiment, the working process executed by the MobileNetV3 model based on the above-mentioned various components is as follows: First, the image passes through a 1×1 convolution to scale the number of channels by adjusting the number of convolution kernels to reduce the computational amount, and then an activation function is applied to optimize the computational efficiency while retaining the non-linear expression. The specific calculation formula is: .

[0062] Secondly, for each channel in the image, a 3×3 convolution kernel is used to independently extract features, and the number of output image channels is the same as the input. The output end is connected to a BN network and a Hardswish activation function; the image after convolution is downsampled and then passes through two fully connected layers, denoted as FC1 and FC2 respectively. The first fully connected layer FC1 uses a ReLU activation function to increase non-linearity, and the second fully connected layer FC2 uses a hard-σ activation function to generate attention weight and other information. The features obtained from the original input image and the processed image are added and fused through skip connection, and finally the number of channels is adjusted through a 1×1 convolution to output the final feature representation. That is to say, after the image passes through multiple networks in the MobileNetV3 model, local features of high-frequency details such as grain boundaries and microcracks can be extracted.

[0063] It should be noted that for the images input by the data acquisition layer, the first extraction sub-layer and the second extraction sub-layer process the input images in parallel. After processing, the first extraction sub-layer outputs global features, while the second extraction sub-layer outputs local features. The obtained global features and local features are uniformly input into the data processing layer for processing.

[0064] Data processing layer: As Figure 2 shown, the data processing layer is respectively connected to the first extraction sub-layer and the second extraction sub-layer. The data processing layer adopts a cross-modal attention mechanism to dynamically fuse global features and local features. The specific calculation formula is: , where G represents global features, L represents local features, both G and L are dimensional vectors, and A is the data obtained after fusing global features and local features. Since this data contains features in multiple scales such as global and local, and both global features and local features are represented by vectors, this data is also called a multi-scale joint embedding vector.

[0065] It can be seen from this that the cross-modal attention mechanism calculates the similarity score between global features and local features, and through weighted adjustment of local features, the fused features can not only contain the spatial information of the crystal but also highlight high-frequency details such as grain boundaries and microcracks, providing rich feature information for the subsequent data generation layer.

[0066] Data generation layer: As Figure 3 shown, the data generation layer is established based on an improved U-Net++ model. Specifically, the improved U-Net++ model means that dense skip connections are established between the encoder and the decoder to enhance multi-scale feature transfer, that is, to enhance the transfer of multi-scale joint embedding vectors. In the Figure 3 example, assuming that the output feature of the i-th layer of the encoder is E i , the left dotted box part represents the encoder, the input feature of the j-th layer of the decoder is D j , and the part symmetric to the encoder represents the decoder. By setting dense skip connections between the encoder and the decoder, D j not only receives the output of the previous layer of the decoder but also receives the features of the corresponding encoder layer and intermediate layers, that is, , is the fusion function. Specifically, can be a splicing operation, splicing multiple feature maps in the channel dimension, and then adjusting the number of channels and fusing features through 1×1 convolution. In the Figure 3In the example, when fusing the output features E3 of the third layer of the encoder and D2 after two layers of decoding, first concatenate E3 and D2 along the channel dimension to obtain a feature map with a dimension of and then adjust the number of channels to an appropriate dimension through a 1×1 convolution to achieve the effective transmission and fusion of the multi-scale joint embedding vectors, enhance the ability of the U-Net++ model to utilize different-scale structural information, and enable better recovery of the details and overall morphology of the mesoscopic structure of the rock to be reconstructed during the reconstruction process.

[0067] Based on the improved U-Net++ model, the model also includes the following core components: Deformable convolution: Deformable convolution is used to achieve dynamic upsampling and improve the accuracy of grain boundary contour reconstruction. The calculation formula of the deformable convolution has been described in step S105 of the method embodiment, so it will not be elaborated here. It should be noted that during the process of reconstructing the contour of the grain boundary, the fixed sampling positions of the traditional convolution kernels are difficult to accurately capture the irregular shape of the grain boundary, while the deformable convolution obtains the contour offset through training and learning so that when sampling, the convolution kernel can dynamically adjust the upsampling position, making the contour of the grain boundary in the obtained input image more fitting to the actual grain boundary contour.

[0068] Conditional generative adversarial network: It mainly includes a generator and a discriminator. The generator receives the input image to obtain a candidate image, and the discriminator judges the similarity between the candidate image and the standard image. Finally, the candidate image with the largest similarity value is used as the reconstructed image.

[0069] It should be noted that when the generator generates a candidate image, in addition to relying on the input image, it will also retrieve the noise vector and the mesoscopic structure feature parameters of the rock to be reconstructed. The noise vector is obtained through prior training, and the feature parameters include but are not limited to data such as mineral content and crystal shape. Therefore, the generator generates a candidate image based on multiple factors to ensure the accuracy of the obtained candidate image. When the discriminator judges the similarity between the candidate image and the standard image, it also considers the mesoscopic structure feature parameters of the rock to be reconstructed. If the feature parameters reflected by the candidate image and the standard image do not match, the similarity value between the candidate image and the standard image is lower. By this way, the obtained reconstructed image is constrained to make the mesoscopic structure of the rock to be reconstructed in the reconstructed image more in line with the actual situation.

[0070] In a possible implementation manner, the reconstruction system further includes a model training layer, and the model training layer is used to train the reconstruction model composed of the above-mentioned data acquisition layer, data extraction layer, data processing layer and data generation layer.

[0071] Before training the reconstruction model using the model training layer, it is necessary to first construct a multi-modal training data set, such as Figure 5As shown, the construction process of the data set is as shown in steps S201 - S206.

[0072] Step S201, obtain SEM images.

[0073] First, use electronic devices such as CT or SEM to scan the rock specimen slices to obtain clear SEM images. During the scanning process, by controlling parameters such as the acceleration voltage and magnification of the electronic device, ensure that the image can clearly show the mesoscopic structure of the rock, such as features like mineral crystals, grain boundaries, and microcracks.

[0074] Then, after denoising and feature enhancement of the SEM images respectively using the preprocessing method in step S101 of the method embodiment, save the SEM images.

[0075] Step S202, count the mineral content in the SEM images. Identify different mineral phases in the SEM images, segment the regions belonging to different mineral phases, and calculate the proportion of the pixel numbers occupied by each mineral phase in the SEM image, so as to obtain the corresponding mineral content. The specific calculation formula is: , Among them, is the total number of pixels in the SEM image, is the number of pixels of one of the mineral phases, is the mineral content of this mineral phase.

[0076] Step S203, identify the crystal shape to obtain characteristic parameters. Based on the digital image processing recognition algorithm, analyze the polygonal structure of the crystals in the SEM image. For each crystal polygonal structure, extract the major axis and the minor axis lengths, then obtain; equiaxed ratio , λ is used to measure the degree of closeness to an equiaxed shape of the crystal; roundness , where s is the area of the crystal polygonal structure, is the perimeter of the crystal polygonal structure. The closer the roundness is to 1, the closer the shape is to a circle; sphericity , is the volume of the crystal, is the equivalent diameter of the crystal. The sphericity is used to describe the degree of closeness of the crystal to a sphere; In addition, the roughness rate and characteristic parameters such as vertices and extreme points of the crystals are also calculated. Among them, the roughness rate is obtained by analyzing the change degree of the pixel points on the crystal boundary.

[0077] After obtaining various characteristic parameters of the crystal as described above, by statistically analyzing the characteristic parameters of a large number of crystal polygonal structures, a probability distribution function for each characteristic parameter is established.

[0078] Step S204, statistically analyze the grain boundary distribution: Identify and label the grain boundaries in the SEM image, and statistically analyze information such as the grain boundary length and area of different grain boundary types (such as coherent grain boundaries, semi-coherent grain boundaries, incoherent grain boundaries, etc.). Calculate the ratio of this type of grain boundary to the SEM image based on the area of the grain boundary, and characterize the proportion of this type of grain boundary in the overall grain boundary distribution through the size of the ratio.

[0079] Step S205, match topological labels for the crystal. Based on the gradient boosting machines (GBM) and random joint modeling method of digital image recognition, match topological labels for the crystal. The topological labels cover information such as the position, shape, contact relationship, and grain boundary type of the crystal. Among them, the GBM model is an ensemble learning algorithm that improves the performance of the model by combining multiple weak prediction models. In this application, the GBM model is used to achieve directional modeling.

[0080] During directional modeling, determine the grain boundaries based on the characteristic parameters of the crystal. For example, determine the grain boundaries through the constraints of mineral content, crystal shape characteristics, and grain boundary distribution. Specifically: First, based on the pixel ratio of the mineral phase in the SEM image, map each ratio to the modeling area, and pre-allocate the area of each phase region in the modeling area. Second, based on the probability distribution function of each characteristic parameter, perform directional adjustment on each phase in the modeling area. Then, based on the ratio of different grain boundary types, pre-define the spatial distribution probability of each grain boundary type, and at the same time assign a corresponding contact type label to each grain boundary, and adjust the grain boundary length and area distribution to match the statistical results. Finally, based on the Delaunay triangulation, generate a background grid, superimpose the constraints of mineral content, crystal shape, and grain boundary type in the grid, and establish a multi-dimensional mapping relationship through quantitative analysis of the statistical parameters of mineral content, crystal shape, and grain boundary distribution. This mapping process uses Kriging interpolation technology to extend the two-dimensional characteristic parameters to three-dimensional space to ensure the spatial continuity of the characteristic parameter distribution, thereby completing the construction of the boundary of the mesoscopic structure, that is, obtaining the grain boundaries. Based on the obtained grain boundaries, then adjust the volume ratio of the crystal blocks according to the shape constraint polygon of the crystal and the mineral content, so as to obtain a refined digital model. Each digital model is pre-set with a corresponding topological label. Therefore, after obtaining the digital model of the crystal, match the topological label corresponding to this type of digital model, and use the matched topological label as the topological label of the crystal.

[0081] The random joint modeling method is used for random modeling, which means that the shape of the polyhedron is determined by sampling the characteristic parameters of the crystal. For example, the volume ratio of the crystal is randomly determined according to the mineral content distribution, and the geometric structure of the crystal is constructed. At the same time, considering the spatial relationship of the crystal, the polygon boundary is adjusted by combining dynamic collision detection to avoid crystal overlap. Then, the Monte Carlo method is used to randomly assign corresponding label types to each grain boundary, and the label ratio corresponding to each type of grain boundary is statistically calculated. The topological label of this type of grain boundary is the label with the highest ratio.

[0082] Based on the topological labels obtained from orientation modeling and random modeling respectively, a spatial partitioning strategy is adopted to divide the modeling area into an orientation area and a random area. The overlapping area of the topological labels in the orientation modeling and random modeling is eliminated through an edge matching algorithm, and finally merged into a unified topological label set. The topological labels in the topological label set are stored in a hierarchical structure, including an orientation label layer, a random label layer, and an integrated label layer.

[0083] Step S206, construct a multi-modal training data set. By integrating SEM images, characteristic parameters, and topological labels, a multi-modal training data set is obtained, such as , where are characteristic parameters, is the SEM image, is the topological label.

[0084] Based on the obtained multi-modal training data set , randomly split the multi-modal training data set into three groups, one of which is the training group, and the other two are the control group and the verification group respectively. The training group and the control group are input into the model training layer, and the model training layer trains the reconstruction model according to the data of the training group and the control group.

[0085] Such as Figure 2 shown, the model training layer mainly includes a first training sub-layer, a second training sub-layer, and a third training sub-layer connected in series in sequence.

[0086] Among them, the first training sub-layer is used to construct a loss function, specifically: a multi-task loss function that combines the pixel-level reconstruction loss of the image, the grain boundary contrast loss (Dice coefficient), and the regression loss of the characteristic parameters, etc., to optimize the training process of the reconstruction model.

[0087] The pixel-level reconstruction loss combines the structural similarity index (SSIM index) and the absolute error loss (L1). Among them, the SSIM index is an index used to measure the similarity degree between two digital images. When one of the two images is a distortion-free image, also known as the standard image, and the other is a distorted image, also known as the reconstructed image, then the structural similarity between the two can be regarded as a measure of the image quality of the distorted image. In this embodiment, the calculation formula of the loss function of the SSIM index is: , where is the standard image, is the reconstructed image, , and are the means of the standard image and the reconstructed image respectively, which are used to measure the average brightness of the image. and are the standard deviations of the standard image and the reconstructed image respectively. The standard deviation is used to reflect the degree of dispersion of the image pixel values, that is, the larger the standard deviation, the richer the details of the image; is the covariance of the standard image and the reconstructed image, and this covariance is used to reflect the correlation of the change of their pixel values; and are constants used to maintain stability. L is the dynamic range of pixel values (for 8-bit images, L = 255). and are both preset values. In this embodiment, , . During training, it is necessary to continuously adjust the loss function of the SSIM index to ensure that the calculated loss value approaches 0, so that the reconstructed image is closer to the standard image.

[0088] L1 is used to calculate the absolute difference between the reconstructed image and the standard image. The calculation formula of the L1 loss function is: , where and are the height and width of the image respectively, and are the pixel values of the standard image and the reconstructed image at the position respectively.

[0089] It can be seen from this that through the L1 loss function, the pixel values of the reconstructed image can be made as close as possible to the pixel values of the standard image, which helps to improve the accuracy of the reconstructed image.

[0090] The Dice coefficient loss function is mainly used to measure the degree of overlap between the grain boundaries in the reconstructed image and those in the standard image. The calculation formula of the Dice coefficient loss function is as follows: , where A is the set of pixels of the grain boundaries in the standard image, and B is the set of pixels of the grain boundaries in the reconstructed image. represents the number of pixels in the overlapping part of the grain boundaries in the standard image and those in the reconstructed image. |A| and |B| are the total numbers of pixels of the grain boundaries in the standard image and the reconstructed image respectively. In this embodiment, by controlling the value of the Dice coefficient loss function to continuously approach 0, it shows that the grain boundaries in the reconstructed image are more coincident with those in the standard image, which can effectively improve the accuracy of grain boundary contour reconstruction.

[0091] The regression loss of the characteristic parameters mainly uses the mean absolute error (MAE) loss function. The MAE loss function is used to evaluate the accuracy of the reconstructed model in identifying the characteristic parameters. Specifically, the calculation formula of the MAE loss function is as follows: , where is the characteristic parameter in the standard image, is the characteristic parameter in the reconstructed image, is the number of samples. In this embodiment, by controlling the value of the MAE loss function to continuously approach 0, the average deviation between the reconstructed image output by the reconstructed model and the standard image can be adjusted to be smaller, which helps to ensure that the reconstructed model can accurately reflect the characteristic parameters of the mesoscopic structure of the rock.

[0092] Based on the above multiple loss functions, in order to overall control the optimization direction of the reconstructed model, this embodiment also sets a total loss function: , where , , , are all weight coefficients, which can be adaptively adjusted according to the importance of each loss function. This embodiment does not make any restrictions.

[0093] The second training sublayer mainly uses the five-fold cross-validation method to evaluate the accuracy of the reconstructed image obtained by the reconstructed model. In this embodiment, the training group is called to train the reconstructed model, and then the control group is used to evaluate the performance of the reconstructed model, and each evaluation index is recorded. By looping at least five times, the average value of the five evaluation results is selected as the final performance index of the reconstructed model to more accurately evaluate the generalization ability and reconstruction accuracy of the reconstructed model.

[0094] The third training sub-layer mainly optimizes the parameters of the reconstruction model. During the above training process, by adjusting the network structure parameters and training parameters of the reconstruction model, the network structure parameters include but are not limited to the size of the subspace in the ViT model, the specified dimension of the projection, the number of attention dimensions of the multi-head self-attention mechanism; the number of layers and channels of the MobileNetV3 network; the number of layers of the encoder and decoder in the U-Net++ architecture, the size of the convolutional kernel, the way of skip connection, etc. The training parameters include but are not limited to the size of the loss value, the number of iterations, the iteration duration, etc. After repeatedly adjusting the parameters of the reconstruction model, the best set of parameters is selected as the parameters of the reconstruction model, so that the reconstruction model can learn the mesoscopic structure of the rock to be reconstructed and achieve the purpose of reconstructing the mesoscopic structure in the image, and can ensure the accuracy of the reconstructed image obtained by the reconstruction model.

[0095] It should be noted that the data layer, model, etc. involved in the above embodiments are all software programs and are applied to the execution entity for processing the images of the rocks to be reconstructed.

[0096] It should also be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In practical applications, all the above possible implementation manners can be combined arbitrarily to form the possible embodiments of the present application, which will not be elaborated herein one by one.

[0097] Based on the same inventive concept, the embodiments of the present application also provide an electronic device, including a processor and a memory. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the mesoscopic structure reconstruction method based on deep learning of rock images in any one of the above embodiments.

[0098] In an exemplary embodiment, an electronic device is provided, such as Figure 6 shown. Figure 6 As shown, the electronic device 600 includes: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present application.

[0099] The processor 601 may be a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0100] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0101] The memory 603 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0102] The memory 603 is used to store the computer program code for executing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the computer program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.

[0103] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0104] Based on the same inventive concept, the embodiments of this application also provide a storage medium, in which a computer program is stored. Among them, the computer program is set to execute the method for mesoscopic structure reconstruction based on deep learning of rock images in any of the foregoing embodiments when running.

[0105] Those skilled in the art can clearly understand that the specific working processes of the above-described systems, devices, and modules can refer to the corresponding processes in the foregoing method embodiments. For the sake of brevity, they will not be described in detail here.

[0106] Those of ordinary skill in the art can understand that: the technical solution of this application can essentially or all or part of this technical solution be embodied in the form of a software product. This computer software product is stored in a storage medium, which includes several program instructions for causing an electronic device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of this application when running the program instructions. And the foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program code.

[0107] Alternatively, all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions (such as an electronic device such as a personal computer, a server, or a network device, etc.). The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the electronic device, the electronic device executes all or part of the steps of the methods described in the embodiments of this application.

[0108] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that within the spirit and principle of the present application, it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of the present application.

Claims

1. A mesoscopic structure reconstruction method based on deep learning of rock images, characterized in that, Including: Obtain an image of the rock to be reconstructed; After dividing the image into blocks and mapping it to a space of a specified dimension, multiple subspaces are obtained, and global features are obtained based on the long-range dependence relationship between the multiple subspaces; Extract the grain boundaries and microcracks in the image to obtain local features. The grain boundary is the boundary of the crystal in the image, the crystal is a component of the rock to be reconstructed, and the microcrack is the gap between two adjacent crystals; Fuse the global features and the local features to obtain a multi-scale joint embedding vector; Reconstruct the contour of the grain boundary according to the multi-scale joint embedding vector to obtain a reconstructed image.

2. The method according to claim 1, wherein The reconstructing the contour of the grain boundary according to the multi-scale joint embedding vector to obtain a reconstructed image includes: Obtain the contour offset of the grain boundary; Adjust the multi-scale joint embedding vector according to the contour offset to obtain an input image; Obtain the reconstructed image according to the input image and a standard image. The standard image refers to an image showing the mineral composition, crystal arrangement, and pore distribution of the rock to be reconstructed.

3. The method according to claim 2, wherein The adjusting the multi-scale joint embedding vector according to the contour offset to obtain an input image includes: , wherein, is the input image, is the multi-scale joint embedding vector, is the image of the rock to be reconstructed, is the preset weight corresponding to the multi-scale joint embedding vector, F( ) is the deformable convolution function, and the deformable convolution function is used to adjust the multi-scale joint embedding vector according to the contour offset to obtain the input image, is the contour offset.

4. The method according to claim 2, characterized in that The obtaining the reconstructed image according to the input image and a standard image includes: Use a generator to obtain at least one candidate image according to the input image; Use a discriminator to obtain the similarity between the candidate image and the standard image according to the candidate image; Select the candidate image with the highest similarity as the reconstructed image.

5. The method according to claim 1, characterized in that After obtaining the image of the rock to be reconstructed, the method further includes: Use a Gaussian filtering algorithm to filter out the interference signals in the image; Use a histogram equalization method to enhance the contrast of the image.

6. A mesoscopic structure reconstruction system based on deep learning of rock images, for performing the method according to any one of claims 1 to 5, characterized in that, Including: A data acquisition layer for obtaining an image of the rock to be reconstructed; A data extraction layer including a first extraction sub-layer and a second extraction sub-layer arranged in parallel. The first extraction sub-layer is used to extract the global features of the image, and the second extraction sub-layer is used to extract the local features of the image; A data processing layer is respectively connected to the first extraction sub-layer and the second extraction sub-layer, and is used to fuse the global features and the local features to obtain the multi-scale joint embedding vector; A data generation layer for obtaining a reconstructed image according to the multi-scale joint embedding vector.

7. The system according to claim 6, wherein The first extraction sub-layer uses a ViT model to extract the global features of the image. The ViT model is used to perform the following operations: Linear embedding and marking: Divide the image into blocks and map it to a space of a specified dimension to obtain multiple subspaces; Position encoding: After capturing the spatial order of the multiple subspaces, generate a spatial sequence; Weight adjustment: Adjust the weight of each subspace according to the long-range dependence relationship between the multiple subspaces in the spatial sequence; Extract global features.

8. The system according to claim 6, wherein The second extraction sub-layer uses a MobileNetV3 model to extract the local features of the image; The MobileNetV3 model at least includes a separable convolutional network, a non-linear activation network, and a residual network.

9. The system according to any one of claims 7 or 8, characterized in that, The data processing layer uses a cross-modal attention mechanism to fuse the global features and the local features; The calculation formula of the cross-modal attention mechanism is as follows: , Among them, G represents the global feature, L represents the local feature, and both G and L are dimensional vectors, A is the multi-scale joint embedding vector, is the function of the cross-modal attention mechanism.

10. The system according to claim 6, wherein The system further includes a model training layer, which is used to train a reconstruction model. The reconstruction model is obtained by combining the data acquisition layer, the data extraction layer, the data processing layer, and the data generation layer; The model training layer includes: A first training sub-layer, which is used to construct a loss function for constraining the reconstruction model; A second training sub-layer, which uses a five-fold cross-validation method to evaluate the accuracy of the reconstructed image output by the reconstruction model; A third training sub-layer: adjusts the parameters of the reconstruction model and selects the best set of parameters as the parameters of the reconstruction model.

Citation Information

Patent Citations

  • Deep electrical tomography method and system based on residual self-attention connection

    CN116740213A

  • Brain tumor image-oriented three-dimensional reconstruction detection method and system

    CN118279302A

Cited By

  • Mineral particle identification and grading method based on digital image and deep learning algorithm

    CN120526229A