Three-dimensional reconstruction method for material microstructure of multi-mode diffusion model
The three-dimensional microstructure is reconstructed through the multimodal diffusion model, which solves the problems of insufficient microdetail control, generation efficiency and multimodal data integration in the existing technology, and realizes high-precision and rapid three-dimensional reconstruction, which is suitable for resource-constrained environments.
Patent Information
- Application Number
- CN202510081365.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-24
AI Technical Summary
The existing three-dimensional microstructure reconstruction methods have significant shortcomings in microdetail control, generation efficiency and multimodal data integration, resulting in insufficient material behavior prediction, difficulty in adapting to complex microstructure changes, and high computing resource consumption.
A three-dimensional reconstruction method for material microstructures of multimodal diffusion model is proposed. By pre-processing the three-dimensional original structure, two-dimensional slices and mask slices are generated, and inputting them to the forward diffusion module to increase noise, outputting pure noise two-dimensional slices to update the multi-modal diffusion model, and then reconstructing the three-dimensional structure with the reverse denoising module to update the model.
This method can accurately reconstruct the complex microstructure inside the material, improve the accuracy and reliability of microstructure generation, enhance the resolution of the reconstruction model, ensure the authenticity of the samples, and reduce the consumption of computing resources.
Smart Images

Figure CN120198577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional reconstruction, and in particular to a method for three-dimensional reconstruction of the microstructure of materials using a multimodal diffusion model. Background Art
[0002] Most traditional microstructure reconstruction methods rely on Representative Volume Element (RVE) simulations. Although these models can reveal the relationship between microstructure and macroscopic properties, they face challenges in effectively transferring microscopic information to the macroscopic scale. The main problem is that they fail to fully consider the diversity and randomness of the microstructure, resulting in inaccurate predictions of material behavior. Due to the significant differences in microscopic characteristics among different materials, traditional methods lack sufficient flexibility to adapt to complex microstructure changes, thus limiting their application in the design of new materials. Existing deep learning-based generation methods often lack precise control over details and multiphase structures during the three-dimensional structure generation process, leading to deviations between the generated results and the real structure in terms of details. These methods often rely on single-modal information and ignore the integration of multi-source data, further limiting the accuracy of the model in dealing with complex structures. In addition, single-modal input methods usually cannot fully express the spatial resolution during three-dimensional reconstruction, and the generated microstructures may be incoherent or deviated. In the three-dimensional reconstruction of microstructures, although diffusion models have high accuracy, they still face certain challenges. The DDPM model requires multiple iterations to generate high-quality three-dimensional structures, which not only takes a long time but also has high requirements for computing resources and is not suitable for resource-constrained environments. In addition, the model is still insufficient in simulating structural anisotropy, especially when dealing with materials with significant anisotropy such as alloys and composites. In summary, there is still significant room for improvement in existing three-dimensional microstructure reconstruction methods in terms of microscopic detail control, generation efficiency, and multi-modal data integration. Such deficiencies limit their application in materials science. Especially in scenarios that require high-precision and rapid reconstruction, more efficient and accurate three-dimensional reconstruction models are urgently needed to make up for the deficiencies of existing technologies. Summary of the Invention
[0003] In view of the problems existing in the above-mentioned existing technologies, the present invention is proposed.
[0004] To achieve the above object, the present invention provides the following technical solution: A method for three-dimensional reconstruction of the microstructure of materials using a multimodal diffusion model, including,
[0005] Preprocess the three-dimensional original structure to obtain two-dimensional slices and mask slices;
[0006] Input the three-dimensional original structure, two-dimensional slices, and mask slices into the forward diffusion module to add noise, output pure-noise two-dimensional slices, and update the multimodal diffusion model;
[0007] The reverse denoising module of the multi-modal diffusion model processes the pure noise two-dimensional slice, the three-dimensional original structure, and the mask slice to obtain the three-dimensional structure and update the multi-modal diffusion model.
[0008] As a further solution of the present invention: The steps of preprocessing the three-dimensional original structure to obtain the two-dimensional slice and the mask slice include:
[0009] Obtain the three-dimensional original structure from the open-source dataset;
[0010] Slice along the X, Y, and Z directions of the three-dimensional original structure respectively to obtain two-dimensional slices;
[0011] Form a mask slice by splicing the segmentation mask with the two-dimensional slice.
[0012] As a further solution of the present invention: The steps of inputting the three-dimensional original structure, the two-dimensional slice, and the mask slice into the forward diffusion module to add noise, output the pure noise two-dimensional slice, and update the multi-modal diffusion model include:
[0013] Take the three-dimensional original structure and the mask slice as constraint conditions and input them into the forward diffusion module;
[0014] The forward diffusion module gradually introduces Gaussian noise to the input two-dimensional slice to obtain a noisy two-dimensional slice;
[0015] Input the noisy two-dimensional slice and the time step into the U-Net model of the forward diffusion module to estimate the noise and obtain the pure noise two-dimensional slice;
[0016] Calculate the loss between the introduced Gaussian noise and the noise estimated by the Unet model, and update the U-Net model by gradient.
[0017] As a further solution of the present invention: When the forward diffusion module gradually introduces Gaussian noise to the input two-dimensional slice, the probability distribution formula at time step t is,
[0018]
[0019] where q(x t |x t-1 ) represents the probability distribution of the state x t-1 at time step t given the state x t at t-1, I represents the identity matrix, β t represents the amount of noise introduced at time step t, and N represents the normal distribution;
[0020] The noise distribution at time step t can be derived from the noise at the initial moment:
[0021]
[0022] Among them, q(x t |x0) represents the probability distribution of the image x at time step t given the original image x0, I represents the identity matrix, β t represents the amount of noise introduced at time step t, x0 represents the original two-dimensional slice, and N represents the normal distribution. t As a further solution of the present invention: the formula for the forward noise addition process of the diffusion model is
[0023]
[0024]
[0025] where x t represents the noisy two-dimensional slice at time step t, x0 represents the original two-dimensional slice, represents the product of all α i from time step 1 to t, which represents the cumulative retention ratio of the original image information from the beginning to time step t. ∈ represents the noise sampled from the standard normal distribution N(0, I), where I is the identity matrix, indicating that the noise is independent in all dimensions and has the same variance.
[0026] As a further solution of the present invention: the formula for the forward noise addition process of the multi-modal diffusion model is
[0027]
[0028] where x b,t represents the masked feature information of the slice at time step t, x b represents the two-dimensional image, represents the product of all αi from time step 1 to t, which represents the cumulative retention ratio of the original image information from the beginning to time step t. ∈ represents the noise sampled from the standard normal distribution N(0, I), where I is the identity matrix.
[0029] As a further solution of the present invention: the loss function used for training Unet is
[0030] Loss = |||∈ - ∈ θ (x t ,t)|| 2
[0031] where x t represents the pure noise two-dimensional slice at time step t, ∈ represents the noise sampled from the standard normal distribution N(0, I), t represents time, and ∈ θ (x t ,t) represents the noise in the case of time step t.
[0032] As a further aspect of the present invention: The reverse denoising module of the multi-modal diffusion model processes the pure noise two-dimensional slice, the three-dimensional original structure, and the mask slice to obtain a three-dimensional structure. The steps of updating the multi-modal diffusion model include:
[0033] Taking the three-dimensional original structure and the mask slice as constraint conditions and inputting them into the reverse denoising module;
[0034] Inputting the pure noise two-dimensional slice and the time step into the updated U-Net model to obtain the predicted current noise;
[0035] Inputting the predicted current noise into the reverse denoising formula of the multi-modal diffusion model to obtain the denoised two-dimensional slice;
[0036] Repeating the above steps for iterative denoising;
[0037] Recombining the denoised two-dimensional slices to obtain a three-dimensional structure.
[0038] As a further aspect of the present invention: The denoising formula is
[0039]
[0040] where x t-1 represents the two-dimensional slice at time step t-1; x t represents the noise-added two-dimensional slice at time step t; ∈ θ (x t , t) represents the noise added at time step t; σ t represents the variance at each time step t; z represents a variable; (1-α t ) is the noise introduced at time step t, I is the identity matrix, and t represents the time feature; α t represents a coefficient; represents the product of all α i from time step 1 to t.
[0041] As a further aspect of the present invention: The formula for the reverse denoising process of the multi-modal diffusion model is:
[0042]
[0043] where x b,t-1 represents the noise-added image combined with mask feature information at time step t-1, x b,t represents the noise-added image combined with mask feature information at time step t, ∈ θ (x t , t) represents the noise at time step t; σ t represents the variance at each time step t; z represents the randomly sampled noise used to introduce the randomness of data diffusion; (1-α t) is the noise introduced at time step t, I is the identity matrix, t represents the time feature; α t represents the coefficient; represents the product of all α from time step 1 to t, and b represents the two-dimensional slice. i
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows: The method for three-dimensional reconstruction of the material microstructure of the multi-modal diffusion model introduces a denoising diffusion implicit model that can transform two-dimensional slices into three-dimensional structures, and precisely reconstructs the complex microstructure inside the material through the denoising diffusion implicit model; The image segmentation mask technology is introduced into the reconstruction of the material microstructure, enhancing the resolution of the reconstruction model and ensuring the authenticity of the samples; At the model input stage, the microstructure segmentation mask and time step are introduced, and the information of different modalities is used to constrain the generation of three-dimensional structures, improving the accuracy and reliability of microstructure generation. Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0046] Figure 1 is the training flow chart of the Unet network.
[0047] Figure 2 is the reverse denoising process.
[0048] Figure 3 is the overall network framework diagram of the inverse structure design algorithm based on the multi-modal conditional diffusion model. Among them, (a) overall architecture, (b) attention mechanism module, (c) residual module.
[0049] Figure 4 is the multi-modal conditional diffusion model based on the U-net network architecture diagram.
[0050] Figure 5 is the 3D visualization of Beadpack data and the comparison of experimental statistical results. (a-c) respectively represent the original three-dimensional structure, the three-dimensional structure generated by the Slicegan method, and the three-dimensional structure generated by this study. (d-f) respectively represent the average porosity distribution measured by the layer-by-layer slices perpendicular to the X, Y, and Z planes. (g-i) respectively represent the statistical results of S2(r), C2(r), and L(r) of the original, Slicegan, and this method models.
[0051] Figure 6 For the 3D visualization of Berea standstone data and the comparison of experimental statistical results, (a-c) represent the original three-dimensional structure, the three-dimensional structure generated by the Slicegan method, and the three-dimensional structure generated in this study, respectively. (d-f) represent the average porosity distribution measured by layer-by-layer slicing perpendicular to the X, Y, and Z planes, respectively. (g-i) represent the statistical results of S2(r), C2(r), and L(r) of the original, Slicegan, and this method models, respectively. Detailed implementation manners
[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific implementation manners of the present invention will be given in conjunction with the drawings of the specification.
[0053] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0054] Secondly, the present invention will be described in detail with reference to the schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views showing the device structure will be enlarged locally out of the general scale, and the schematic diagrams are only examples and should not limit the scope of protection of the present invention herein. In addition, in actual production, three-dimensional spatial dimensions including length, width, and depth should be included.
[0055] Furthermore, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" appearing in different places in this specification does not all refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.
[0056] Embodiment 1
[0057] As Figures 1 to 4 shown, the present invention provides a technical solution: a method for three-dimensional reconstruction of the microstructure of materials using a multimodal diffusion model, including,
[0058] S1: Preprocess the three-dimensional original structure to obtain two-dimensional slices and masked slices;
[0059] This step is used to obtain two-dimensional slices and masked slices. The segmentation mask provides boundary information of different phases in the material; the segmentation mask is spliced with the two-dimensional slice to form a masked slice, which can help the model capture the global and local features of the material, so as to maintain these key microscopic features during the noise addition and denoising processes.
[0060] S2: Input the three-dimensional original structure, two-dimensional slices, and mask slices into the forward diffusion module to add noise, output pure-noise two-dimensional slices, and update the multi-modal diffusion model;
[0061] It should be noted that step S2 is a forward diffusion process. Specifically, the input of the three-dimensional original structure and mask slices in this step is used to constrain the generation of the three-dimensional structure. This process is for forward diffusion, and in the forward process, noise is gradually added until it becomes pure noise. Specifically, at this stage, Gaussian noise is gradually introduced into the two-dimensional slices. After adding noise multiple times (1000 times), pure-noise two-dimensional slices that conform to the normal distribution are finally generated. This process is similar to changing a two-dimensional slice image into a pure-noise two-dimensional slice image, making it have a certain degree of randomness.
[0062] S3: The reverse denoising module of the multi-modal diffusion model processes the pure-noise two-dimensional slices, three-dimensional original structure, and mask slices to obtain the three-dimensional structure and update the multi-modal diffusion model.
[0063] It should be noted that step S3 is a reverse denoising process. Specifically, the input of the three-dimensional original structure and mask slices is used to constrain the generation of the three-dimensional structure. This process is reverse denoising. In the reverse process, the noise in the forward process is gradually removed, so as to achieve the purpose of reconstructing the three-dimensional structure. Specifically, the model uses pure-noise two-dimensional slices as the input, removes the noise through the reverse diffusion process, and gradually reconstructs a three-dimensional structure that is highly consistent with the three-dimensional original structure. To improve the reconstruction accuracy, this solution integrates the segmentation mask, time encoding, etc. into multi-modal data (specifically, multi-modal data refers to data from different data sources or different types, such as images and texts, audio and texts, videos and audios), and guides the reverse denoising process of the diffusion model. The specific process of its guidance is that the segmentation mask provides clear boundary information about different phases in the two-dimensional slices, helping the model maintain the integrity and accuracy of the structure during the denoising process; the time encoding, as a condition, can help the model perform more accurate noise estimation and removal at each time step, and more effectively restore the original image.
[0064] It should be further noted that the segmentation mask and time encoding are converted into feature vectors, and then concatenated and fused with the two-dimensional slice feature vectors at the feature level; in the skip connections of the U-Net, the features of the segmentation mask and time encoding are passed together with the image features.
[0065] As Figure 4 shown, the multi-modal conditional diffusion model is based on the U-net network architecture diagram. (The upper part (purple box) shows the forward diffusion and reverse denoising processes of the diffusion model; the lower part (green box) shows how to reconstruct the three-dimensional structure from the segmentation mask and two-dimensional slices; starting from the three-dimensional structure, two-dimensional slices are extracted and the segmentation mask is generated, and the U-Net network is used for noise estimation, and the original three-dimensional structure is gradually restored from the noisy image).
[0066] Furthermore, the steps of preprocessing the three-dimensional original structure to obtain two-dimensional slices and mask slices include:
[0067] S11: Obtain the three-dimensional original structure from the open-source dataset;
[0068] It should be noted that the three-dimensional original structure specifically refers to a three-dimensional image, which can be modeled and organized into a dataset through modeling software. The open-source dataset refers to data information that is publicly and freely available for anyone to use.
[0069] S12: Perform slicing along the X, Y, and Z directions of the three-dimensional original structure respectively to obtain two-dimensional slices;
[0070] S13: Form mask slices by splicing the segmentation mask with the two-dimensional slices.
[0071] Furthermore, the steps of inputting the three-dimensional original structure, two-dimensional slices, and mask slices into the forward diffusion module to add noise, output pure-noise two-dimensional slices, and update the multi-modal diffusion model include:
[0072] S21: Input the three-dimensional original structure and mask slices as constraint conditions into the forward diffusion module;
[0073] S22: The forward diffusion module gradually introduces Gaussian noise to the input two-dimensional slices to obtain noise-added two-dimensional slices;
[0074] Specifically, this process can be understood as a Markov chain. By increasing the noise level until reaching a preset maximum noise threshold, the original image is gradually transformed into a noise image conforming to the normal distribution. The diffusion time step T is usually set to 1000 to ensure that the increase in noise is orderly and controllable at each time step.
[0075] It should be noted that when the forward diffusion module gradually introduces Gaussian noise to the input two-dimensional slices, the probability distribution formula at time step t is
[0076]
[0077] where q(x t |x t-1 ) represents the probability distribution of the state x t-1 at time step t given the state x t at t-1. I represents the identity matrix, β t represents the amount of noise introduced at time step t, N represents the normal distribution, and the variance sequence {β1, β2...β t} gradually increases as the diffusion time step T increases, ensuring that the amount of noise introduced at each step exceeds the noise increment of the previous step;
[0078] Therefore, the noise distribution at time step t can be derived from the noise at the initial time:
[0079]
[0080] where q(x t |x0) represents the probability distribution of the image x at time step t given the original image x0, I represents the identity matrix, βt represents the amount of noise introduced at time step t, and the variance sequence {β1, β2... β t} gradually increases as the diffusion time step T increases, and x0 represents the original two-dimensional slice. t} increases gradually as the diffusion time step T increases, and x0 represents the original two-dimensional slice.
[0081] Specifically, if we want to sample from an arbitrary Gaussian distribution with mean μ and variance β, we can first sample ε from a standard Gaussian distribution (mean 0, variance 1), and then μ + β·ε is equivalent to the result of sampling from an arbitrary Gaussian distribution. This model can directly represent the image xt at any time step as a function of the original image x0. This model can directly represent the image x t at any time step as a function of the original image x0, specifically as the formula for the forward noise addition process of the diffusion model.
[0082] Applying the reparameterization trick, the formula for the forward noise addition process of the diffusion model is
[0083]
[0084] where x t represents the noisy two-dimensional slice at time step t, x0 represents the original two-dimensional slice, represents the product of all α i from time step 1 to t, which represents the cumulative retention ratio of the original image information from the start to time step t. ∈ represents the noise sampled from the standard normal distribution N(0, I), where I is the identity matrix, indicating that the noise is independent in all dimensions and has the same variance.
[0085] S23: Input the noisy two-dimensional slice and the time step into the U-Net model of the forward diffusion module to estimate the noise and obtain a pure noise two-dimensional slice;
[0086] It should be noted that, for example Figure 3As shown, the model adopts an improved U-Net architecture, introducing 3D convolution, residual blocks, attention mechanism, and temporal embedding into the basic framework, allowing for the fusion of multi-modal data to specifically process 3D material microstructures; the encoder part extracts features through multiple convolutional blocks and residual modules, while introducing attention modules and residual modules to enhance the selectivity for key information; in the decoding stage, features of different scales are combined through skip connections to ensure the effective transmission of information during reconstruction. As Figure 3 shown, (a) Overall architecture: The two-dimensional slices of the input 3D structure are downsampled to extract features, gradually reducing the spatial dimension and increasing the feature dimension, and then upsampled to gradually restore the spatial dimension. At the same time, the features obtained from downsampling are combined with the features obtained from upsampling through skip connections to retain more detailed information. Finally, the reconstructed 3D structure is output; (b) Attention mechanism module: The feature map after downsampling is input, normalized and processed by 3D convolution, then the feature map is adjusted, normalized by the Softmax function to obtain attention weights, and finally the final output is obtained through weighted summation and residual connection; (c) The input of the residual module is first normalized, and then the SiLU activation function is used to increase non-linearity. Next, the data extracts features through a convolutional layer. Some features are regularized by downsampling, and another part adds temporal information through the temporal embedding layer. After the two are concatenated, they are transformed through a linear layer and finally added to the original input to form the output.
[0087] Specifically, in the U-net architecture, convolutional blocks and residual modules are organized in the encoder and decoder, and features are extracted and reconstructed through the downsampling and upsampling processes; the encoder reduces the spatial dimension of the data through multiple downsamplings while increasing the number of channels, thereby extracting higher-level features; the decoder maps these features back to the original spatial dimension through upsampling and skip connections, and combines multi-scale features from the encoder to achieve accurate reconstruction.
[0088] Among them, the process of extracting features by the convolutional block is as follows:
[0089] 1. Initialization: The convolutional block starts with a convolutional layer with a 3×3×3 convolutional kernel, which is used to extract preliminary features from the input 3D original structure, two-dimensional slices, and mask slices;
[0090] 2. Normalization layer: After the convolutional layer, there is a normalization layer, which helps to stabilize the training process and accelerate the convergence speed;
[0091] 3. Activation function: After the normalization layer is the activation function;
[0092] 4. Convolutional layer: Finally, the convolutional block contains a 3×3×3 convolutional layer for further feature extraction.
[0093] Among them, the process of the residual module extracting features is as follows:
[0094] 1. Core convolution block: Each residual module contains two core convolution blocks, and each block contains a normalization layer, an activation function layer, and a 3×3×3 convolution layer.
[0095] 2. Temporal embedding layer: To utilize temporal information, each convolution block embeds a temporal embedding layer, which incorporates information from the temporal dimension into the feature map.
[0096] 3. Feature fusion: The design of the residual module allows the addition of the input features and the features after convolution processing. This structure helps the model maintain the continuity of feature information when dealing with deep networks and prevents the problem of gradient vanishing.
[0097] It should be noted that the noisy two-dimensional slice and the time step are input into the U-Net model of the forward diffusion module to estimate the noise. The noisy two-dimensional slice and the time step are used as multi-modal information. The formula for the forward noise addition process of the multi-modal diffusion model is
[0098]
[0099] where x b,t represents the masked feature information of the slice at time step t, x b represents the two-dimensional image, represents the product of all αi from time step 1 to t, which represents the cumulative retention ratio of the original image information from the beginning to time step t. ∈ represents the noise sampled from the standard normal distribution N(0, I), where I is the identity matrix.
[0100] S24: Calculate the loss between the introduced Gaussian noise and the noise estimated by the Unet model, and update the U-Net model by gradient (as Figure 1 shown).
[0101] It should be noted that the loss function used for training Unet is
[0102] Loss = |||∈ - ∈ θ (x, t)|| 2
[0103] where x t represents the pure noise two-dimensional slice at time step t, ∈ represents the noise sampled from the standard normal distribution N(0, I), t represents time, and ∈ θ (x t , t) represents the noise in the case of time step t.
[0104] Further, the reverse denoising module of the multi-modal diffusion model processes the pure noise two-dimensional slice, the three-dimensional original structure, and the mask slice to obtain the three-dimensional structure. The steps to update the multi-modal diffusion model include:
[0105] S31: Use the three-dimensional original structure and the mask slice as constraint conditions and input them into the reverse denoising module;
[0106] S32: Input the pure noise two-dimensional slice and the time step into the updated U-Net model to obtain the predicted current noise;
[0107] S33: Input the predicted current noise into the reverse denoising formula of the multi-modal diffusion model to obtain the denoised two-dimensional slice;
[0108] S34: Repeat the above steps for iterative denoising;
[0109] S35: Recombine the denoised two-dimensional slices to obtain the three-dimensional structure. The specific process is that each two-dimensional slice has a number, and the denoised two-dimensional slices are recombined in order to form the three-dimensional structure.
[0110] It should be noted that the reverse denoising formula is
[0111]
[0112] where x t-1 represents the two-dimensional slice at time step t-1; x t represents the noisy two-dimensional slice at time step t; ∈ θ (x t , t) represents the noise added at time step t; σ t represents the variance at each time step t; z represents a variable; (1-α t ) is the noise introduced at time step t. As time step t increases, this value increases, meaning the contribution of the noise increases. I is the identity matrix, t represents the time feature; α t represents a coefficient; represents the product of all α i from time step 1 to t.
[0113] It should be noted that in order to achieve an efficient denoising process by comprehensively using multi-modal information, the formula for the reverse denoising process of the multi-modal diffusion model is:
[0114]
[0115] where x b,t-1 represents the noisy image combined with mask feature information at time step t-1, x b,t represents the noisy image combined with mask feature information at time step t, ∈ θ (x t, t) represents the noise at time step t; σ t represents the variance at each time step t; z represents the noise of random sampling, which is used to introduce the randomness of data diffusion; (1-α t ) is the noise introduced by time step t. As time step t increases, this value increases, which means that the contribution of noise increases. I is the unit matrix, and t represents the time feature; α t represents the coefficient; represents all α from time step 1 to t i The product of , b represents a two-dimensional slice.
[0116] Among them, in the diffusion model framework constructed in this scheme, σ t It represents a learnable parameter in the model, representing the variance at each time step t; this parameter is obtained through careful adjustment during the training process of the model to optimize the performance of the model; the variable z refers to the noise component randomly sampled during the diffusion process of the model, which is used to introduce randomness in data diffusion.
[0117] Through a trained multimodal U-Net model (attached Figure 1 ), which is specially designed to handle diffusion processes, and its input is According to the inverse denoising formula of the multimodal diffusion model, the goal of model training is to extract the noise from the image x t Subtract the estimated noise from the t , t), in order to restore a state close to the original image; this method allows the model to learn how to reverse the diffusion process, thereby recovering accurate microstructural information in noisy images (Appendix Figure 2 ).
[0118] like Figure 2 The specific process of reverse denoising is shown below:
[0119] 1. From Gaussian noise x 1000 start.
[0120] 2. Use Unet to predict the current noise ∈ θ (x t ,t).
[0121] 3. Apply the above equation to generate x step by step t-1 .
[0122] 4. Repeat this process until you reach x0, the denoised image.
[0123] Specifically, the reverse process uses an optimized U-Net network architecture. In the encoding stage, it is initialized by a convolutional block with a 3×3×3 convolutional kernel to extract the basic features of the input data. These features are then passed to the residual module to enhance the expressive ability and ensure that the model maintains good feature capture performance in the deep network. The features pass through the downsampling layer to reduce the data dimension and computational amount, improving the computational efficiency. After each downsampling, a residual module is connected immediately to maintain the continuity of the feature information and avoid feature loss. In addition, during the further downsampling process, the model not only strengthens the features through the residual module but also introduces an attention module to improve the sensitivity and selectivity of the model to key features. The residual module contains two core convolutional blocks, and each convolutional block consists of a normalization layer, an activation function layer, and a 3×3×3 convolutional layer. These components work together to optimize feature processing and enhance the nonlinear expressive ability of the model. In addition, to make full use of the time information, a time embedding layer is embedded in each convolutional block to integrate the time dimension information into the feature map. By adding the output of the second convolutional layer in the residual module to the output of the time embedding layer, the model can process time series data more efficiently. This structural design not only makes full use of the time information but also provides support for constructing a deep network and the orderly training of the model.
[0124] To further improve the training efficiency and the robustness of the model, skip connections are introduced before each downsampling. The skip connections are used to integrate multi-scale feature information, ensure that the features are not lost in the deep network, and promote the smooth flow of gradients during training. This strategy not only enhances the model's ability to capture detailed features but also improves the overall stability. After the third and fourth downsamplings in the encoding stage, the combination of the residual module and the attention module further strengthens the model's ability to identify key features.
[0125] After the encoding stage is completed, the model enters the middle layer module, which consists of two residual modules and an attention module to further integrate the deep feature information and ensure the accuracy of the decoding process. In the decoding stage, the model contains four upsampling modules, and the multi-scale features in the encoder are combined with the decoder through skip connections. This design ensures the smooth transmission of information throughout the network and improves the quality and accuracy of the decoder when reconstructing features, ultimately achieving high-precision three-dimensional structure reconstruction.
[0126] The training of the model adopts an iterative optimization strategy. Under different noise levels, through multiple iterative trainings, the model learns how to accurately reconstruct the microstructure of the material in a complex noise environment. At the same time, the model parameters are continuously optimized through error feedback to ensure that the model maintains high precision and consistency in the generated three-dimensional structure.
[0127] This technical solution introduces a denoising diffusion implicit model (DDIM) that can transform two-dimensional slices into three-dimensional structures. Through the denoising diffusion implicit model (DDIM), the complex microstructure inside the material can be accurately reconstructed, and the physical properties of the material can be predicted. Secondly, the image segmentation mask technology can be introduced into the reconstruction of the material microstructure to effectively distinguish each different phase in the material, enhance the resolution of the reconstruction model, and ensure the authenticity of the samples. Thirdly, the microstructure segmentation mask and time step (time-encoded features) are introduced in the model input stage, and information of different modalities is used to constrain the generation of three-dimensional structures, improving the accuracy and reliability of microstructure generation. Finally, the standard U-net network architecture is improved, and a three-dimensional U-net network architecture that can reconstruct the three-dimensional material microstructure is developed. This architecture can optimize the feature processing process and enhance the nonlinear expression ability of the model through residual modules and attention modules.
[0128] Example 2
[0129] The difference between this example and the previous example is that in order to verify the effectiveness of the multi-modal denoising diffusion method proposed in the present invention in reconstructing three-dimensional microstructures, a comparative experimental analysis is carried out on two materials.
[0130] (1) Beadpack material
[0131] A reconstruction experiment from two-dimensional to three-dimensional was carried out on the Beadpack material. Figure 5 (a-c) Details show the original microstructure, the three-dimensional structure generated by the Slicegan algorithm, and the three-dimensional structure reconstructed using the model; it can be observed that the structure generated by the model of this solution is very close to the original structure at the microscopic level, which clearly demonstrates the excellent performance of the method of this solution in accurately reconstructing the complex material microstructure.
[0132] To deeply explore the performance of the micro-3ddpm model in reconstructing the pore structure inside the Beadpack sample, a simulation analysis of the slice porosity perpendicular to the X-axis, Y-axis, and Z-axis was carried out, and the relevant results are as Figure 5 (d-f) shown.
[0133] According to the analysis results, the average porosity in the three axial directions is 0.287; the porosity curves perpendicular to the X-axis and Y-axis are smooth and consistent, indicating that the reconstructed Beadpack material has a relatively high porosity structure uniformity in these two directions, demonstrating the consistency of the material in these directions; while in the Z-axis direction, the fluctuation range of the porosity curve is small, and the deviation from the overall mean is also small, further proving that the reconstructed material exhibits good isotropy in three-dimensional space; in addition, to more comprehensively evaluate the performance of the model, the average two-point correlation function, linear path function, and two-point clustering function of the three orthogonal planes were calculated and analyzed, and their results were compared with the structures generated by the Slicegan algorithm and the original material structure. The detailed data are as Figure 5 (g-i) shows.
[0134] From Figure 5 (g), it can be seen that the two-point correlation function curve of the original material almost completely coincides with the curve generated by the model of this scheme, indicating that the method of this scheme can not only accurately reconstruct the three-dimensional morphology of the material, but also capture and reproduce the complex spatial relationships and structural characteristics existing in the original material; the results of the linear path function (as shown in Figure 5 (h)) further show that the generation model of this scheme also reaches a level similar to that of the original material in capturing the spatial distribution between different phases, indicating that the generated three-dimensional structure performs well in terms of the connectivity of the material; Figure 5 (i) shows the analysis results of the two-point clustering function. The high coincidence of the curves indicates that the model of this scheme maintains a high structural fidelity during the reconstruction process, especially in terms of the aggregation structure and particle distribution, ensuring the consistency of the reconstructed structure with the original material.
[0135] The experimental results verify the effectiveness and accuracy of the micro-3ddpm model from multiple aspects; the model can generate high-quality three-dimensional material microstructures and maintain structural consistency with the original material during the reconstruction process; the model can accurately reproduce the micro-geometric morphology of the material and effectively capture the spatial statistical characteristics of the material; generally speaking, the model proposed in this invention shows excellent ability in reconstructing material microstructures and provides an important reference basis for the future application of generative models in material research.
[0136] (2) Berea standstone
[0137] In this experiment, Berea sandstone was selected as a case to verify the effectiveness of the model; Figure 6(a-c) show the real three-dimensional structure, the three-dimensional structure generated by Slicegan, and the three-dimensional structure of 64×64×64 generated by the model of this solution. To more clearly compare the generated results, this solution conducts slice comparisons in different XYZ directions for the real three-dimensional structure, the structure generated by Slicegan, and the three-dimensional structure generated by the model of this solution.
[0138] The experimental results show that the two-dimensional slices generated by the model of this solution are closer to the real structure; in addition, the average porosity of the three axes is measured, and the result is 0.211.
[0139] Figure 6 (d-f) show the porosity simulation analysis perpendicular to the X, Y, and Z directions. The results show that the porosity changes slightly in the three directions, indicating that the pore structure of the material has good consistency in all directions, reflecting the isotropic characteristics of the Berea sandstone material; this further confirms that the structure generated by the model of this method has high authenticity.
[0140] Figure 6 (g-i) show the average results of each statistical function; the results of the two-point correlation function show that the curve distributions of the model of this solution and Slicegan are similar, but the curve generated by the model of this solution is closer to the real data, showing a high degree of similarity; in addition, the generated curves of the linear path function and the two-point clustering function are closer to the real samples than the Baseline. From the above porosity simulation analysis and comparison of various indicators, the model of this solution shows superiority in the performance of structure reconstruction.
[0141] Importantly, it should be noted that the construction and arrangement of the present application shown in multiple different exemplary embodiments are merely illustrative. Although only a few embodiments are described in detail in this disclosure, those who refer to this disclosure should easily understand that many modifications are possible without substantially departing from the novel teachings and advantages of the subject matter described in this application. For example, the dimensions, scales, structures, shapes and proportions of various elements, as well as parameter values (e.g., temperature, pressure, etc.), installation arrangements, use of materials, color, orientation changes, etc. For example, an element shown as integrally formed may be composed of multiple parts or elements, the position of the element may be inverted or otherwise changed, and the nature, number or position of discrete elements may be altered or changed. Therefore, all such modifications are intended to be included within the scope of the present invention. The order or sequence of any process or method steps may be changed or reordered according to alternative embodiments. In the claims, any "means-plus-function" clause is intended to cover the structures that perform the functions described herein, and not only structural equivalents but also equivalent structures. Other substitutions, modifications, changes and omissions may be made in the design, operating conditions and arrangement of the exemplary embodiments without departing from the scope of the present invention. Therefore, the present invention is not limited to specific embodiments, but extends to various modifications that still fall within the scope of the appended claims.
[0142] In addition, in order to provide a concise description of the exemplary embodiments, not all features of the actual embodiments may be described, i.e., those features that are not relevant to the currently considered best mode of implementing the present invention or those features that are not relevant to the implementation of the present invention.
[0143] It should be understood that in the development of any actual implementation, as in any engineering or design project, a large number of specific implementation decisions may be made. Such development efforts may be complex and time-consuming, but for those of ordinary skill in the art who benefit from this disclosure, without excessive experimentation, the development efforts will be a routine task of design, manufacturing and production.
[0144] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention may be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model, characterized by: include, Preprocessing the three-dimensional original structure to obtain two-dimensional slices and mask slices; Input the 3D original structure, 2D slices and mask slices into the forward diffusion module to add noise, output pure noise 2D slices and update the multimodal diffusion model; The inverse denoising module of the multimodal diffusion model processes the pure noise 2D slices, 3D original structures and mask slices to obtain the 3D structure and update the multimodal diffusion model.
2. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 1, characterized in that: The steps of preprocessing the original three-dimensional structure to obtain two-dimensional slices and mask slices include: Obtaining 3D raw structures in open source datasets; Slice the original three-dimensional structure along the X, Y and Z directions to obtain two-dimensional slices; Masked slices are formed by concatenating the segmentation mask with the 2D slice.
3. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 1 or 2, characterized in that: The steps of inputting the 3D original structure, the 2D slices and the mask slices into the forward diffusion module to add noise, outputting the pure noise 2D slices and updating the multimodal diffusion model include: The 3D original structure and mask slices are used as constraints and input into the forward diffusion module; The forward diffusion module gradually introduces Gaussian noise into the input two-dimensional slices to obtain noisy two-dimensional slices; The noisy 2D slice and time step are input into the U-Net model of the forward diffusion module to estimate the noise and obtain a pure noise 2D slice; Calculate the loss of the introduced Gaussian noise and the noise estimated by the Unet model, and gradient update the U-Net model.
4. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 3, characterized in that: In the forward diffusion module, Gaussian noise is gradually introduced into the input two-dimensional slices. The probability distribution formula at time step t is: Among them, q(x t |x t-1 ) means that at a given t-1 state x t-1 Under the condition that the state x at time step t t The probability distribution of t represents the amount of noise introduced at time step t, and N represents the normal distribution; The noise distribution at time step t can be derived from the noise at the initial moment: Among them, q(x t |x0) means that given the original image x0, the image x at time step t t The probability distribution of t represents the amount of noise introduced at time step t, x0 represents the original two-dimensional slice, and N represents the normal distribution.
5. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 4, characterized in that: The forward noise addition process formula of the diffusion model is: Among them, x t represents the noisy 2D slice at time step t, x0 represents the original 2D slice, represents all α from time step 1 to t i The product of represents the cumulative retention ratio of the original image information from the beginning to time step t, ∈ represents the noise sampled from the standard normal distribution N(0,I), where I is the identity matrix, indicating that the noise is independent in all dimensions and has the same variance.
6. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 5, characterized in that: The forward noise addition process formula of the multimodal diffusion model is: Among them, x b,t Represents the mask feature information of the slice at time step t, x b represents a two-dimensional image, represents the product of all αi from time step 1 to t, which represents the cumulative retention ratio of the original image information from the beginning to time step t, ∈ represents the noise sampled from the standard normal distribution N(0,I), where I is the identity matrix.
7. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 6, characterized in that: The loss function used to train Unet is, Loss=|||∈-∈ θ (x t ,t)|| 2 Among them, x t represents a two-dimensional slice of pure noise at time step t, ∈ represents the noise sampled from the standard normal distribution N(0,I), t represents time, ∈ θ (x t , t) represents the noise at time step t.
8. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 7, characterized in that: The inverse denoising module of the multimodal diffusion model processes the pure noise two-dimensional slices, the three-dimensional original structure and the mask slices to obtain the three-dimensional structure. The steps of updating the multimodal diffusion model include: The 3D original structure and mask slices are used as constraints and input into the inverse denoising module; The pure noise 2D slice and time step are input into the updated U-Net model to obtain the predicted current noise; The predicted current noise is input into the inverse denoising formula of the multimodal diffusion model to obtain a denoised two-dimensional slice; Repeat the above steps to iteratively denoise; Reconstruct the denoised 2D slices to obtain the 3D structure.
9. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 8, characterized in that: The denoising formula is: Among them, x t-1 represents a two-dimensional slice at time step t-1; x t represents a noisy 2D slice at time step t; ∈ θ (x t , t) represents the noise added at time step t; σ t represents the variance at each time step t; z represents the variable; (1-α t ) is the noise introduced by time step t, I is the unit matrix, and t represents the time feature; α t represents the coefficient; represents all α from time step 1 to t i The product of .
10. The method for three-dimensional reconstruction of material microstructure of a multimodal diffusion model according to claim 9, characterized in that: The formula for the inverse denoising process of the multimodal diffusion model is: Among them, x b,t-1 represents the noisy image combined with mask feature information at time step t-1, x b,t Represents the noisy image combined with mask feature information at time step t, ∈ θ (x t , t) represents the noise at time step t; σ t represents the variance at each time step t; z represents the noise of random sampling, which is used to introduce the randomness of data diffusion; (1-α t ) is the noise introduced by time step t, I is the unit matrix, and t represents the time feature; α t represents the coefficient; represents all α from time step 1 to t i The product of , b represents a two-dimensional slice.
Citation Information
Cited By
Vertebral bone tissue segmentation method combined with structure-guided conditional diffusion network
CN121121125A