Alveolar bone defect image restoration method based on global features
By combining the global semantic repair model and the high-frequency detail repair model, the problems of consistency and inter-layer continuity in the image repair of alveolar bone defects are solved, more accurate alveolar bone repair surgery planning is achieved, and the effect and recovery quality of alveolar bone repair surgery are improved.
Patent Information
- Application Number
- CN202510786055.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-16
AI Technical Summary
The existing technology for image restoration of alveolar bone defects has a limited receptive field of convolutional neural networks, resulting in poor consistency in the content of the restoration results, and serious inter-layer discontinuity problems during three-dimensional image processing, which affects the accuracy of medical decision-making.
A global semantic restoration model and a high-frequency detail restoration model are combined to perform global semantic prediction and LAMA network detail restoration through the ViT encoder. The model parameters are optimized by combining inter-layer sliding window sampling and multi-objective loss function to achieve continuity and consistency restoration of three-dimensional alveolar bone images.
It improves the content consistency and inter-layer continuity of alveolar bone repair images, provides more accurate guidance for alveolar bone repair surgery planning, and improves surgical results and recovery quality.
Smart Images

Figure SMS_4 
Figure SMS_9 
Figure SMS_12
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a method, device and storage medium for restoring an alveolar bone defect image based on global features. Background Art
[0002] Alveolar bone restoration surgery, which involves placing implants in the maxillary and mandibular bones to replace lost tooth roots and serve as a fixed support for artificial teeth, has become a key area of modern dental care and is widely used for patients with tooth loss. Alveolar bone defects are a common challenge in alveolar bone restoration surgery, primarily caused by periodontal disease, trauma, tumor resection, genetics, and systemic diseases. In addition to impacting dental implant surgery, alveolar bone defects can also lead to reduced chewing efficiency, speech impairments, and altered facial contours.
[0003] With the continuous advancement of image processing and artificial intelligence technologies, generative AI has gained increasing importance in the field of medical image processing. Medical images can predict the normal morphology of a patient's diseased area, providing a crucial basis for surgical planning and disease detection. Image restoration technology is a key tool in computer vision for image editing and analysis. In this context, image restoration can predict the normal morphology of the diseased area, specifically the normal morphology of the alveolar bone defect, based on the non-diseased area, serving as a primary reference for surgical planning.
[0004] Although deep learning-based image restoration technology has achieved certain results, it still has some limitations. For one thing, the direct training process using convolutional neural networks as the model architecture is limited by the receptive field of convolutional neural networks, and the model is insufficient in understanding the global features of the image, resulting in problems with the content consistency of the restoration results. For another thing, when processing three-dimensional medical images, existing methods typically split the 3D image into two-dimensional images and then perform layer-by-layer restoration. This processing method can cause serious inter-layer discontinuities. For example, when several two-dimensional cross-sectional layers are restored and then stacked into three-dimensional volume data, the restored image will show obvious quality defects such as burrs and artifacts when viewed from the coronal and sagittal planes. These problems can interfere with doctors' observation and judgment, affecting the accuracy of medical decisions. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and provide a method for image restoration of alveolar bone defects based on global features, comprising the following steps: S1: Acquire the patient's three-dimensional CBCT image data, obtain the alveolar bone area and perform standardization processing to obtain a pre-processed image; S2: The global semantic restoration model masks some image blocks of the preprocessed image to obtain random masked image blocks, and the ViT encoder of the global semantic restoration model performs global semantic prediction on the random masked image blocks based on the non-masked image blocks to generate a reconstructed image and a global attention feature map; S3: Inputting the reconstructed image and the global attention feature map into a high-frequency detail restoration model, obtaining multi-scale features from the reconstructed image according to the LAMA network, fusing the multi-scale features with the global attention feature map to obtain high-frequency features, performing Fourier transform on the high-frequency features through a Fourier convolution module to obtain transformed features, and fusing the transformed features with the multi-scale features to generate a restored image; S4: applying a three-dimensional continuity constraint to the inpainted image based on an inter-layer sliding window sampling strategy; S5: Optimize model parameters by using a joint multi-objective loss function, and update models including the global semantic restoration model and the high-frequency detail restoration model according to the model parameters, wherein the joint multi-objective loss function includes L1 reconstruction loss, perceptual loss, style loss, adversarial loss, and inter-layer consistency loss.
[0006] Preferably, in step S2, the global semantic restoration model masks part of the image blocks of the preprocessed image to obtain random masked image blocks, further comprising: The global semantic restoration model divides the preprocessed image into non-overlapping image blocks, randomly masks 75% of the image blocks to obtain the randomly masked image blocks, and the remaining 25% of the image blocks are the non-masked image blocks.
[0007] Preferably, in step S2, the ViT encoder of the global semantic restoration model performs global semantic prediction on the random masked image block according to the non-masked image block to generate a reconstructed image and a global attention feature map, including: S21: the ViT encoder linearly projects the non-masked image block to obtain a non-masked position representation, and the multi-layer attention obtains a masked position representation according to the non-masked position representation; S22: The ViT decoder superimposes the non-masked position representation and the masked position representation of the masked image block to synthesize an image sequence representation, obtains the reconstructed image through the transposed attention layer and the MLP, and splices the output of the attention head into the global attention feature map; S23: Supervising the pixel-level difference between the reconstructed image and the randomly masked image block in the three-dimensional CBCT image data through an L2 loss function.
[0008] Preferably, in step S21, the ViT encoder linearly projects the non-masked image block to obtain a non-masked position representation, and the multi-layer attention obtains a masked position representation according to the non-masked position representation, including: The ViT encoder obtains a query matrix, a key matrix, and a value matrix according to the non-masked image block. The query matrix , the bond matrix and the value matrix The calculation formula is as follows: in, , , is the learnable parameter matrix, is the normalized feature of the non-masked image block; The attention score is calculated based on the query matrix, the key matrix and the value matrix, and the formula is as follows: in, is the bond matrix The transposed matrix of The attention score is normalized by softmax to obtain the attention weight, the formula is as follows: in, is the key vector dimension; Calculating a weighted value matrix according to the attention weights and the value matrix; Adding the weighted value matrix to the non-masked image block to obtain an intermediate feature; After the intermediate features are processed by LayerNorm, they are input into MLP for nonlinear transformation to obtain output features; The output feature is added to the intermediate feature to obtain the mask position representation.
[0009] Preferably, in step S3, the reconstructed image obtains multi-scale features according to the LAMA network, and the multi-scale features are fused with the global attention feature map to obtain high-frequency features, including: Inputting the reconstructed image into the LAMA network to extract the multi-scale features; Adding the multi-scale features and the global attention feature map element by element to obtain a fused feature; The high-frequency features are obtained through the convolution layer.
[0010] Preferably, in step S3, the high-frequency features are Fourier transformed by a Fourier convolution module to obtain transformed features, including: According to the Fourier convolution module, the original spatial data of the high-frequency feature is converted to the frequency domain and Fourier transformed to obtain a complex tensor with a dimension of ; The real and imaginary parts of the complex tensor are concatenated along the channel dimension to obtain a real tensor with the dimension of ; In the frequency domain, the real tensor is input to the convolution layer to extract frequency domain features, and the output tensor is obtained through the batch normalization layer and the ReLU activation function layer. The dimension of the output tensor is ; The output tensor is re-divided into real and imaginary parts, and the frequency domain data is converted into the spatial domain using inverse Fourier transform to obtain the transformation feature, with a dimension of .
[0011] Preferably, in step S4, applying a three-dimensional continuity constraint to the repaired image based on an inter-layer sliding window sampling strategy includes: S41: starting from the starting position of the repaired image, sliding the window stepwise in three dimensions including the X-axis, the Y-axis, and the Z-axis according to the window size according to the set sliding step length, traversing the entire repaired image; S42: When the window covers a certain local area, extract the voxel features including grayscale value and texture feature in the window, and calculate the feature statistics including mean, variance and covariance for each layer of the window; S43: For adjacent layers, that is, windows along the Z-axis direction, the similarity measurement of inter-layer features is calculated based on the feature statistics. If the inter-layer similarity measurement result exceeds the set continuity threshold, it means that there is an inter-layer discontinuity problem in the current window area. For the window area determined to be discontinuous, a three-dimensional interpolation algorithm including trilinear interpolation and nearest neighbor interpolation is used to adjust the voxel value, and return to step S41 for resampling until the continuity measurement of all inter-layer areas meets the continuity threshold to complete the optimization of all window areas.
[0012] Preferably, in step S5, the joint multi-objective loss function includes L1 reconstruction loss, perceptual loss, style loss, adversarial loss and inter-layer consistency loss, including: The L1 reconstruction loss calculates the average of the absolute differences between the inpainted image and the 3D CBCT image at the pixel level, and then updates the model parameters using a gradient descent algorithm so that the pixel values of the inpainted image gradually approach the pixel values of the real image. The calculation formula is as follows: in, is the true pixel value of the 3D CBCT image, is the predicted pixel value of the restored image; The perceptual loss compares the high-level feature differences between the restored image and the 3D CBCT image based on a pre-trained network. , the calculation formula is: in, Represents the pre-trained network The feature map of the layer, For the The number of channels of the layer feature map, For the The height of the layer feature map, For the The width of the layer feature map, y is the feature map of the three-dimensional CBCT image, is the feature map of the restored image; The style loss is used to calculate the Gram matrix of the feature map of the restored image and the 3D CBCT image to measure the style. The calculation formula is: in, For the Gram matrix of the feature map of the three-dimensional CBCT image, For the The Gram matrix of the feature map of the layer repair image; The adversarial loss consists of a generator and a discriminator. The generator is responsible for generating the repaired image, and the discriminator determines whether the input image is a three-dimensional CBCT image or the generated repaired image. The adversarial loss formula of the generator is for: in, is the output of the discriminator for the 3D CBCT image, is the repaired image generated by the generator based on the noise z; The inter-layer consistency loss is obtained by calculating the restoration results of the overlapping group and superimposing the restoration results of all overlapping groups. The overlapping group is a group of overlapping layers that is divided into a plurality of overlapping groups according to the step size. Adjacent layers have continuity. The calculation formula of the inter-layer consistency loss is: in, is the restoration result output by the i-th layer.
[0013] Based on the same concept, the present invention also provides a computer device including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor performs the steps of the alveolar bone defect image repair method based on global features as described in any one of the embodiments.
[0014] Based on the same concept, the present invention also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the alveolar bone defect image repair method based on global features as described in any one of the embodiments.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention masks some image blocks of the preprocessed image through a global semantic restoration model to obtain random masked image blocks. The ViT encoder performs global semantic prediction on the random masked image blocks based on the non-masked image blocks to generate a reconstructed image to first repair the basic content. Then, the high-frequency detail restoration model is input to supplement the high-frequency details, completing the restoration task of the alveolar bone CBCT image in steps.
[0016] The present invention utilizes a global semantic repair model with global attention to provide a powerful global feature prior for the convolutional neural network, expands the receptive field of the corresponding model, and can utilize global relevant information and integrate long-range information including the contralateral side to guide the repair of specific locations, thereby greatly improving the content consistency of the original convolutional network.
[0017] The present invention combines a multi-objective loss function to optimize model parameters. The multi-objective loss function includes L1 reconstruction loss, perceptual loss, style loss, adversarial loss and inter-layer consistency loss. This ensures that the comprehensive restoration requirements of content, style, intensity, and inter-layer continuity are met during the training image restoration process, and fully guarantees the consistency and rationality of anatomical structure information, effectively solving the problem of difficulty in ensuring the comprehensive consistency of image restoration. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those skilled in the art by reading the following detailed description of the preferred embodiment.The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the invention.
[0019] Figure 1 This is a flow chart of the alveolar bone defect image restoration method based on global features of the present invention; Figure 2 This is the overall framework diagram of the masked self-encoding of the present invention; Figure 3 This is a diagram of the semantic repair model architecture of the present invention; Figure 4 This is a diagram of the multi-layer attention mechanism of the present invention; Figure 5 This is the overall architecture diagram of the repair network of the present invention; Figure 6 This is the inter-layer consistency loss diagram of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Obviously, the embodiments described are part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.
[0021] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0022] First embodiment At present, alveolar bone repair surgery is generally used to repair alveolar bone defects. However, there is no unified and universal standard for the classification of alveolar bone defects, the selection and planning of the best alveolar bone repair surgery plan. Based on the "restoration-oriented" principle of implant treatment, doctors need to choose a reasonable surgical plan to reconstruct alveolar bone defects. Clinicians are required to master the indications of various bone augmentation technologies and provide patients with the correct alveolar bone repair surgery plan. Alveolar bone repair surgery requires steps including surgical planning based on medical imaging, incision and exposure of bone surface, filling of bone implant materials, and suturing. Among these, determining the volume and range of bone augmentation based on X-rays or CBCT images will have a huge practical impact on the patient's surgical results and postoperative recovery, and is a very critical surgical step.
[0023] However, before this, experience-based surgical planning often required doctors to have a lot of clinical experience, and it was difficult to determine the specific size and range of the incremental volume, which posed a great challenge to the program setting and surgical planning of alveolar bone repair surgery.
[0024] See also Figure 1As shown, the alveolar bone defect image restoration method based on global features provided in this embodiment is inspired by the idea of masked autoencoder and other technologies. A global semantic restoration model and a high-frequency detail restoration model based on masked autoencoder are constructed. This method realizes three-dimensional image restoration of alveolar bone for the first time and is a robust restoration method for CBCT-based surgical planning. It can provide doctors with important guidance including results of bone defect-related areas. Specifically, it includes the following steps: Before performing alveolar bone repair surgery, doctors generally analyze the patient's alveolar bone condition through panoramic X-rays or CBCT images, and based on this, complete the corresponding surgical planning.
[0025] S1: Acquire the patient's three-dimensional CBCT image data, acquire the alveolar bone area and standardize it to obtain a preprocessed image. Specifically, in this embodiment, this embodiment selects a CBCT that can provide a three-dimensional situation of the alveolar bone as input, completes image restoration on this basis, inputs the patient's image, and completes the generation of a surgical reference target for the alveolar bone part to determine the situation of the alveolar bone defect.
[0026] Preferably, obtaining three-dimensional CBCT image data of the patient, obtaining the alveolar bone region and performing standardization processing to obtain a pre-processed image, includes: Three-dimensional CBCT image data of the alveolar bone region is acquired through CBCT scanning. The three-dimensional CBCT image data is input into a trained segmentation model to predict and output the segmentation results of the alveolar bone region. The segmentation model is trained by first collecting a large number of labeled CBCT images (the alveolar bone region has been labeled) and then using machine learning algorithms (such as random forests and support vector machines) or deep learning models (such as UNet and 3DUNet) to learn the characteristic patterns of the alveolar bone. Map the grayscale values of all pixels in the image to a specific interval (such as [0, 1] or [0, 255]), calculate the minimum and maximum grayscale values in the image, and convert each pixel using the formula (pixel value - minimum value) / (maximum value - minimum value) to ensure consistent grayscale scales across patient images. Normalize the grayscale values to a normal distribution with a fixed mean and standard deviation (e.g., mean 0, standard deviation 1). First, calculate the mean and standard deviation of the image grayscale values. Then, transform each pixel using the formula (pixel value - mean) / standard deviation to eliminate grayscale offsets caused by device differences.
[0027] Because the spatial resolution (voxel size) of images scanned by different CBCT devices may vary, they need to be standardized to a standard resolution (e.g., 1 mm × 1 mm × 1 mm). Based on the voxel size of the original image and the target resolution, a scaling factor is calculated, and the images are interpolated and resampled to ensure consistent spatial spacing across all images. All images are aligned to the same anatomical coordinate system, typically using specific anatomical landmarks on the patient's head (e.g., nasion and auricular points) as a reference. Image position and orientation are adjusted through rotation and translation to make images from different patients spatially comparable. Images are then scaled or cropped to maintain consistent physical dimensions or voxel size across all images for ease of subsequent analysis and model training. For example, images can be cropped into fixed-size regions of interest (ROIs) to remove extraneous background information and ensure geometric consistency for subsequent processing.
[0028] In the first stage, the global semantic restoration model, also known as the ViT (Vision Transformers) model, is trained based on the mask autoregressive strategy. The model's global information interaction capability is utilized to enable it to learn global feature priors and prepare for the next stage.
[0029] S2: The global semantic restoration model masks some image blocks of the preprocessed image to obtain random masked image blocks. The ViT encoder of the global semantic restoration model performs global semantic prediction on the random masked image blocks based on the unmasked image blocks to generate a reconstructed image and a global attention feature map. Specifically, in this embodiment, the purpose of the first stage is to obtain global understanding and ensure content consistency. In clinical scenarios, common restoration planning methods often originate from the symmetry of human anatomy and use normal information from the contralateral side to guide the restoration of the current diseased side. However, in real-world situations, on the one hand, due to the influence of many factors (congenital incomplete symmetry, postnatal development, and changes in the position of surrounding tissues caused by disease), it is difficult for the human body to achieve a completely symmetrical structure, making it difficult to obtain accurate planning based on symmetry methods. On the other hand, symmetry is only a special form of long-range attention. Global attention proposes using different global elements (tokens) for pairwise interaction, which can better learn the relationship between elements and fully utilize the different features of the entire 3D image. Therefore, we use the mainstream mask autoregressive task to guide the ViT (Vision Transformers) model to obtain global understanding information as a prior.
[0030] See also Figure 2 As shown, in step S2, the global semantic restoration model masks some image blocks of the preprocessed image to obtain random masked image blocks, further comprising: The global semantic restoration model divides the preprocessed image into non-overlapping image blocks, randomly masks 75% of the image blocks to obtain randomly masked image blocks, and the remaining 25% of the image blocks are non-masked image blocks. Specifically, in this embodiment, during the ViT training process, the preprocessed image is divided into non-overlapping blocks (patches) of 16×16 pixels, and 75% of the blocks are randomly masked (for example, 64 blocks are retained out of 256 blocks), which is significantly higher than the 15% mask rate of BERT. The model is restored based on the remaining 25% of non-masked image blocks, and the L2 loss is calculated between the restored results and the original image blocks, which prompts the training network to learn the relationship between different global image blocks and understand the global scene to guide the next stage.
[0031] See also Figure 2 、 Figure 3 and Figure 4 As shown, in step S2, the ViT encoder of the global semantic restoration model performs global semantic prediction on the random masked image block based on the non-masked image block to generate a reconstructed image and a global attention feature map, including: S21: The ViT encoder performs linear projection on the non-masked image block to obtain a non-masked position representation. Multi-layer attention is used to obtain a masked position representation (positional embedding) based on the non-masked position representation. Specifically, in this embodiment, the ViT encoder consists of 12 layers of Transformer modules. Each Transformer module contains a specific operation flow for processing and extracting features of the input visible block, completing global token interaction, and adding a learnable position encoding to each non-masked image block to obtain a masked position representation for preserving spatial topology information. S22: The ViT decoder superimposes the non-masked position representation and the masked position representation (mask token) of the masked image block to synthesize the image sequence representation, obtains the reconstructed image through the transposed attention layer and the MLP, and splices the output of the attention head into a global attention feature map. Specifically, in this embodiment, the masked position representation is spliced with the non-masked image block features extracted by the ViT encoder to form a complete feature sequence. The spliced feature sequence is input into the 4-layer Transformer decoder. Based on the non-masked position representation and the masked position representation, the content of the random masked image block is predicted, thereby generating a preliminary repaired image. S23: Supervise the pixel-level difference between the reconstructed image and the randomly masked image patches in the 3D CBCT image data through the L2 loss function.
[0032] See also Figure 4As shown, in step S21, the ViT encoder linearly projects the non-masked image block to obtain a non-masked position representation, and the multi-layer attention obtains a masked position representation based on the non-masked position representation, including: The ViT encoder obtains the query matrix, key matrix, value matrix, and query matrix based on the non-masked image block. , bond matrix Sum Matrix The calculation formula is as follows: in, , , is the learnable parameter matrix, is the normalized feature of the non-masked image block; The attention score is calculated based on the query matrix, key matrix and value matrix. The formula is as follows: in, is the bond matrix Specifically, in this embodiment, in order to complete the global feature interaction, multi-head attention is first used to allow different tokens to interact and calculate the similarity, that is, the attention score; The attention score is softmax normalized to obtain the attention weight, the formula is as follows; in, is the key vector dimension; Calculate the weighted value matrix based on the attention weights and value matrix; The weighted value matrix is added to the non-masked image block to obtain an intermediate feature. Specifically, in this embodiment, the weighted value matrix is added to the non-masked image block to obtain the intermediate feature to implement a residual connection, thereby avoiding gradient explosion in a deep network. After the intermediate features are processed by LayerNorm, they are input into MLP for nonlinear transformation to obtain output features. Specifically, in this embodiment, MLP is used to transform the features to increase the expressive power of the global semantic repair model. The output features are added to the intermediate features to obtain the masked position representation and realize the quadratic residual connection.
[0033] Based on the information about the non-masked area, the model predicts the normal morphology of the masked area. Images of the corresponding area in a healthy individual are used as a gold standard for supervision, allowing the model to learn the basic morphology of a normal person. During inference, the model infers the lesion data, masks the lesion area, and estimates the normal morphology within the lesion area based on information from other non-lesion areas.
[0034] In the second stage, in order to further restore image details, especially the lost high-frequency features, the LAMA network is used, and the global feature prior of the first stage is used to complete further repair through the convolutional network and the additional inserted Fourier convolution module to restore more image details. Among them, the attention feature map output by the decoder in the first stage is used as an additional input before the Fourier convolution network to provide global features, further enhance the network's receptive field, and ensure content consistency. At the same time, the model uses inter-layer continuity consistency loss to greatly improve the inter-layer continuity and ensure a reasonable three-dimensional structure.
[0035] S3: The reconstructed image and the global attention feature map are input into the high-frequency detail restoration model. The reconstructed image obtains multi-scale features according to the LAMA network. The multi-scale features and the global attention feature map are fused to obtain high-frequency features. The high-frequency features are Fourier transformed through the Fourier convolution module to obtain transformation features. The transformation features are fused with the multi-scale features to generate a restored image. Specifically, in this embodiment, the Fourier convolution module used further increases the receptive field of the downstream restoration network.
[0036] Preferably, in step S3, the reconstructed image obtains multi-scale features according to the LAMA network, and the multi-scale features are fused with the global attention feature map to obtain high-frequency features, including: The reconstructed image is input into the LAMA network to extract multi-scale features. Specifically, in this embodiment, when the reconstructed image is input into the encoder of the LAMA network, the network extracts features at different levels. Since the convolution kernel size, step size, and receptive field at different levels are different, the extracted features can reflect the information of the image at different scales. The multi-scale features and the global attention feature map are added element by element to obtain a fused feature. Specifically, in this embodiment, starting from the first element of the multi-scale feature map and the global attention feature map, the elements at corresponding positions are extracted in sequence (for example, the element at the i-th row, j-th column, and k-th channel in the multi-scale feature map is extracted, and the element at the same position (i-th row, j-th column, and k-th channel) in the global attention feature map is extracted at the same time), and the two extracted elements at corresponding positions are added to obtain a new element value. This new element value will be used as the element value at the same position in the fused feature map. The above operation is performed on each position in the feature map in a left-to-right and top-to-bottom order until all elements of the multi-scale feature map and the global attention feature map are traversed. In this way, the element-by-element addition of the two feature maps is completed, and a fused feature map is obtained. High-frequency features are obtained through the convolution layer. Specifically, in this embodiment, the designed convolution kernel is slid on the fused feature map. Each time it slides to a position, the convolution kernel is element-wise multiplied with the feature map area corresponding to the position, and then all products are added to obtain the feature value at the position after convolution. During the convolution process, padding is used to control the size of the output feature map, and a convolution operation is performed on each channel of the fused feature map. Then, the convolution results of all channels are added to obtain the final output value.
[0037] See also Figure 5 As shown, in step S3, the high-frequency features are Fourier transformed by the Fourier convolution module to obtain transformed features, including: According to the Fourier convolution module, the original spatial data of high-frequency features is converted to the frequency domain and Fourier transformed to obtain a complex tensor with a dimension of ; The real and imaginary parts of the complex tensor are concatenated along the channel dimension to obtain a real tensor with the dimension ; In the frequency domain, the real tensor is input to the convolution layer to extract the frequency domain features, and the output tensor is obtained after the batch normalization layer and the ReLU activation function layer. The dimension of the output tensor is ; The output tensor is re-divided into real and imaginary parts, and the inverse Fourier transform is used to convert the frequency domain data into the spatial domain to obtain the transformation features, with a dimension of Specifically, in this embodiment, analysis and processing are performed in the frequency space of the feature map, which can capture more high-frequency details and expand the receptive field of the convolutional network, thereby further improving the restoration effect.
[0038] S4: Based on the inter-layer sliding window sampling strategy, three-dimensional continuity constraints are imposed on the repaired image.
[0039] Preferably, in step S4, a three-dimensional continuity constraint is imposed on the repaired image based on an inter-layer sliding window sampling strategy, including: S41: starting from the starting position of the repaired image, the window is gradually slid along the three dimensions including the X-axis, the Y-axis, and the Z-axis according to the window size according to the set sliding step size, traversing the entire repaired image; S42: When the window covers a certain local area, extract the voxel features including grayscale value and texture feature in the window, and calculate the feature statistics including mean, variance and covariance for each layer of the window; S43: For adjacent layers, that is, windows along the Z-axis, the similarity measurement of inter-layer features is calculated based on the feature statistics. If the inter-layer similarity measurement result exceeds the set continuity threshold, it means that there is an inter-layer discontinuity problem in the current window area. For the window area determined to be discontinuous, a three-dimensional interpolation algorithm including trilinear interpolation and nearest neighbor interpolation is used to adjust the voxel value, and return to step S41 for resampling until the continuity measurement of all inter-layer areas meets the continuity threshold to complete the optimization of all window areas.
[0040] S5: Optimize model parameters using a joint multi-objective loss function. Update models including the global semantic restoration model and the high-frequency detail restoration model based on the model parameters. The joint multi-objective loss function includes L1 reconstruction loss, perceptual loss, style loss, adversarial loss, and inter-layer consistency loss.
[0041] Preferably, in step S5, the joint multi-objective loss function includes L1 reconstruction loss, perceptual loss, style loss, adversarial loss and inter-layer consistency loss, including: The L1 reconstruction loss calculates the average of the absolute differences between the inpainted image and the 3D CBCT image at the pixel level, and then updates the model parameters using the gradient descent algorithm so that the pixel values of the inpainted image gradually approach the pixel values of the real image. The calculation formula is as follows: in, is the true pixel value of the 3D CBCT image, is the predicted pixel value of the restored image. Specifically, in this embodiment, the restored image needs to be as consistent as possible with the real image, so the consistency of the loss needs to be guaranteed; Perceptual loss compares the high-level feature differences between the restored image and the 3D CBCT image based on the pre-trained network , the calculation formula is: in, Represents the pre-trained network The feature map of the layer, For the The number of channels of the layer feature map, For the The height of the layer feature map, For the The width of the layer feature map, y is the feature map of the three-dimensional CBCT image, It is a feature map of the restored image. Specifically, in this embodiment, the image clarity and perceptual reconstruction effect are guaranteed by calculating the perceptual distance; The style loss is used to calculate the Gram matrix of the feature map of the restored image and the 3D CBCT image to measure the style. The calculation formula is: in, For the Gram matrix of the feature map of the three-dimensional CBCT image, For the The Gram matrix of the feature map of the layer-restored image. Specifically, in this embodiment, the Gram matrix is used to calculate the distance between the style vectors of the image to ensure that it is still within the modality style of CBCT and meets basic reconstruction requirements; The adversarial loss consists of a generator and a discriminator. The generator is responsible for generating the repaired image, and the discriminator determines whether the input image is a 3D CBCT image or a generated repaired image. The adversarial loss formula of the generator is for: in, is the output of the discriminator for the 3D CBCT image, is the restored image generated by the generator based on the noise z. Specifically, in this embodiment, a zero-sum game is achieved between the generator and the discriminator through a generative adversarial network, thereby ensuring that the image generated by the generator is so realistic that the discriminator cannot distinguish it, achieving the desired reconstruction effect in a high-dimensional space. The inter-layer consistency loss is obtained by calculating the restoration results of the overlapping groups and superimposing the restoration results of all overlapping groups. The overlapping groups are formed by dividing the 3D CBCT image into several overlapping groups with several overlapping layers according to the step size. Adjacent layers have continuity. The calculation formula of the inter-layer consistency loss is: in, The restoration result outputted by the i-th layer is shown in FIG1 . Specifically, in this embodiment, the step size is 1. The three-dimensional CBCT image data is divided into multiple overlapping three-layer groups along the axial direction (Z axis) (taking three layers to be restored each time as an example, the section numbers are 0, 1, 2, 3, 4, etc.), and different layers are numbered. In the first restoration process, the restoration results of the first group: layer 0, layer 1, and layer 2 are kept consistent (continuous). Then, in the second restoration process, the restoration results of the second group: layer 1, layer 2, and layer 3 are kept consistent (continuous). When the original sequence numbers are constrained to be layer 1 and layer 2, the restoration results of the two results are kept consistent (continuous). When the consistency between the two results is achieved, the overall structural continuity is transferred from layer 0 to layer 3 in the entire sequence. The third group: the repair results of layer 2, layer 3, and layer 4 are kept consistent (continuous). When the original sequence number is also layer 2 and layer 3, when the consistency between the two results is achieved, the overall structural continuity is transferred from layer 1 to layer 4. Since the overall structural continuity has been transferred from layer 0 to layer 3, that is, the overall structural continuity is transferred from layer 0 to layer 4, the entire sequence gradually accumulates with the increasing number of repairs and the use of consistency constraints, achieving consistency and continuity in the overall three-dimensional structure. Joint multi-objective loss function It combines L1 reconstruction loss, perceptual loss, style loss, adversarial loss and inter-layer consistency loss according to certain weights. The formula is: in, 、 、 、 and are the weights of each loss function.
[0042] The inter-layer consistency reconstruction loss is used to ensure the three-dimensional continuity of the restoration results. In each restoration process, multiple layers of two-dimensional data (2.5D data) are used as input to the convolution. The image restoration results of the multiple layers of two-dimensional data are output simultaneously through transposed convolution or upsampling. The attention module captures the dependencies between layers and ensures the continuity between the results of different restorations for the same layer through the consistency constraints with the real continuous multi-layer image data. That is, the inter-layer continuity of the restoration results of this part of 2.5D data is guaranteed, and then the continuity transfer is completed through the discriminant loss. In addition, since multiple layers of two-dimensional data are input each time, a certain layer of the original three-dimensional data will appear in different positions of the 2.5D data in different sampling times, such as Figure 6 As shown in the figure, when constraining the consistency of the original results of the same layer in the results obtained by different sampling and restoration, the continuous restoration results of different times can be continuously transmitted with the same layer as the medium, thereby ensuring the consistency of the entire 3D restoration result.
[0043] This embodiment uses deep learning technology to construct a new medical image restoration algorithm, and for the first time achieves fully automatic restoration results of alveolar bone CBCT films. The research results of this embodiment will provide effective support for precise preoperative planning, assist in determining the bone reconstruction goals of alveolar bone repair surgery, and improve the clinical efficacy of alveolar bone repair surgery.
[0044] Second embodiment In some embodiments of the present application, a computer device is also provided, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the alveolar bone defect image repair method based on global features in the first embodiment of the present invention.
[0045] The present invention also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, enables the one or more processors to execute the steps of the alveolar bone defect image repair method based on global features in the first embodiment of the present invention.
[0046] It can be understood that, for the aforementioned alveolar bone defect image restoration method based on global features, if it is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0047] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0048] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for restoring alveolar bone defects based on global features, characterized in that: The following steps are involved: S1: Acquire the patient's three-dimensional CBCT image data, obtain the alveolar bone area and perform standardization processing to obtain a pre-processed image; S2: The global semantic restoration model masks some image blocks of the preprocessed image to obtain random masked image blocks, and the ViT encoder of the global semantic restoration model performs global semantic prediction on the random masked image blocks based on the non-masked image blocks to generate a reconstructed image and a global attention feature map; S3: Inputting the reconstructed image and the global attention feature map into a high-frequency detail restoration model, obtaining multi-scale features from the reconstructed image according to the LAMA network, fusing the multi-scale features with the global attention feature map to obtain high-frequency features, performing Fourier transform on the high-frequency features through a Fourier convolution module to obtain transformed features, and fusing the transformed features with the multi-scale features to generate a restored image; S4: applying a three-dimensional continuity constraint to the inpainted image based on an inter-layer sliding window sampling strategy; S5: Optimize model parameters by using a joint multi-objective loss function, and update models including the global semantic restoration model and the high-frequency detail restoration model according to the model parameters, wherein the joint multi-objective loss function includes L1 reconstruction loss, perceptual loss, style loss, adversarial loss, and inter-layer consistency loss.
2. The alveolar bone defect image restoration method based on global features according to claim 1, characterized in that: In step S2, the global semantic restoration model masks some image blocks of the preprocessed image to obtain random masked image blocks, further comprising: The global semantic restoration model divides the preprocessed image into non-overlapping image blocks, randomly masks 75% of the image blocks to obtain the randomly masked image blocks, and the remaining 25% of the image blocks are the non-masked image blocks.
3. The alveolar bone defect image restoration method based on global features according to claim 2, characterized in that: In step S2, the ViT encoder of the global semantic restoration model performs global semantic prediction on the random masked image block according to the non-masked image block to generate a reconstructed image and a global attention feature map, including: S21: the ViT encoder linearly projects the non-masked image block to obtain a non-masked position representation, and the multi-layer attention obtains a masked position representation according to the non-masked position representation; S22: The ViT decoder superimposes the non-masked position representation and the masked position representation of the masked image block to synthesize an image sequence representation, obtains the reconstructed image through the transposed attention layer and the MLP, and splices the output of the attention head into the global attention feature map; S23: Supervising the pixel-level difference between the reconstructed image and the randomly masked image block in the three-dimensional CBCT image data through an L2 loss function.
4. The alveolar bone defect image restoration method based on global features according to claim 3, characterized in that: In step S21, the ViT encoder linearly projects the non-masked image block to obtain a non-masked position representation, and multi-layer attention obtains a masked position representation based on the non-masked position representation, including: The ViT encoder obtains a query matrix, a key matrix, and a value matrix according to the non-masked image block. The query matrix , the bond matrix and the value matrix The calculation formula is as follows: in, , , is the learnable parameter matrix, is the normalized feature of the non-masked image block; The attention score is calculated based on the query matrix, the key matrix and the value matrix, and the formula is as follows: in, is the bond matrix The transposed matrix of The attention score is normalized by softmax to obtain the attention weight, the formula is as follows: in, is the key vector dimension; Calculating a weighted value matrix according to the attention weights and the value matrix; Adding the weighted value matrix to the non-masked image block to obtain an intermediate feature; After the intermediate features are processed by LayerNorm, they are input into MLP for nonlinear transformation to obtain output features; The output feature is added to the intermediate feature to obtain the mask position representation.
5. The alveolar bone defect image restoration method based on global features according to claim 4, characterized in that: In step S3, the reconstructed image obtains multi-scale features according to the LAMA network, and the multi-scale features are fused with the global attention feature map to obtain high-frequency features, including: Inputting the reconstructed image into the LAMA network to extract the multi-scale features; Adding the multi-scale features and the global attention feature map element by element to obtain a fused feature; The high-frequency features are obtained through the convolution layer.
6. The alveolar bone defect image restoration method based on global features according to claim 5, characterized in that: In step S3, the high-frequency features are Fourier transformed by a Fourier convolution module to obtain transformed features, including: According to the Fourier convolution module, the original spatial data of the high-frequency feature is converted to the frequency domain and Fourier transformed to obtain a complex tensor with a dimension of ; The real and imaginary parts of the complex tensor are concatenated along the channel dimension to obtain a real tensor with the dimension of ; In the frequency domain, the real tensor is input to the convolution layer to extract frequency domain features, and the output tensor is obtained through the batch normalization layer and the ReLU activation function layer. The dimension of the output tensor is ; The output tensor is re-divided into real and imaginary parts, and the frequency domain data is converted into the spatial domain using inverse Fourier transform to obtain the transformation feature, with a dimension of .
7. The alveolar bone defect image restoration method based on global features according to claim 6, characterized in that: In step S4, applying a three-dimensional continuity constraint to the repaired image based on the inter-layer sliding window sampling strategy includes: S41: starting from the starting position of the repaired image, sliding the window stepwise in three dimensions including the X-axis, the Y-axis, and the Z-axis according to the window size according to the set sliding step length, traversing the entire repaired image; S42: When the window covers a certain local area, extract the voxel features including grayscale value and texture feature in the window, and calculate the feature statistics including mean, variance and covariance for each layer of the window; S43: For adjacent layers, that is, windows along the Z-axis direction, the similarity measurement of inter-layer features is calculated based on the feature statistics. If the inter-layer similarity measurement result exceeds the set continuity threshold, it means that there is an inter-layer discontinuity problem in the current window area. For the window area determined to be discontinuous, a three-dimensional interpolation algorithm including trilinear interpolation and nearest neighbor interpolation is used to adjust the voxel value, and return to step S41 for resampling until the continuity measurement of all inter-layer areas meets the continuity threshold to complete the optimization of all window areas.
8. The alveolar bone defect image restoration method based on global features according to claim 1, characterized in that: In step S5, the joint multi-objective loss function includes L1 reconstruction loss, perceptual loss, style loss, adversarial loss and inter-layer consistency loss, including: The L1 reconstruction loss calculates the average of the absolute differences between the inpainted image and the 3D CBCT image at the pixel level, and then updates the model parameters using a gradient descent algorithm so that the pixel values of the inpainted image gradually approach the pixel values of the real image. The calculation formula is as follows: in, is the true pixel value of the ith pixel in the 3D CBCT image, is the i-th predicted pixel value of the repaired image; The perceptual loss compares the high-level feature differences between the restored image and the 3D CBCT image based on a pre-trained network. , the calculation formula is: in, Represents the pre-trained network The feature map of the layer, For the The number of channels of the layer feature map, For the The height of the layer feature map, For the The width of the layer feature map, y is the feature map of the three-dimensional CBCT image, is the feature map of the restored image; The style loss is used to calculate the Gram matrix of the feature map of the restored image and the 3D CBCT image to measure the style. The calculation formula is: in, For the Gram matrix of the feature map of the three-dimensional CBCT image, For the The Gram matrix of the feature map of the layer repair image; The adversarial loss consists of a generator and a discriminator. The generator is responsible for generating the repaired image, and the discriminator determines whether the input image is a three-dimensional CBCT image or the generated repaired image. The adversarial loss formula of the generator is for: in, is the output of the discriminator for the 3D CBCT image, is the repaired image generated by the generator based on the noise z; The inter-layer consistency loss is obtained by calculating the restoration results of the overlapping group and superimposing the restoration results of all overlapping groups. The overlapping group is a group of overlapping layers that is divided into a plurality of overlapping groups according to the step size. Adjacent layers have continuity. The calculation formula of the inter-layer consistency loss is: in, is the restoration result output by the i-th layer.
9. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes a module of the alveolar bone defect image repair method based on global features as described in any one of claims 1 to 8.
10. A storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the alveolar bone defect image restoration method based on global features according to any one of claims 1 to 8.