Three-dimensional medical image anomaly detection method and system based on self-supervised learning

Through self-supervised learning, the method of generating pseudo-anomaly images and constructing multi-scale feature embedding and symmetric decoding reconstruction is solved, and the problem of serious dependence on labeled data in three-dimensional medical image abnormality detection is realized, and high-precision label-free abnormality detection is suitable for real clinical scenarios.

CN120259787AActive Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510736680.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The existing three-dimensional medical imaging abnormality detection methods rely on a large amount of labeled data and are difficult to effectively apply under conditions without labels or weak labels. Especially when abnormal samples are scarce or complex distribution, the model performance is unstable.

Method used

Using a self-supervised learning strategy, a label-free self-supervised training framework is constructed by generating multi-scale feature embedding and symmetric decoding and reconstruction of pseudo-exception images and normal images, and a label-free self-supervised training framework is used to generate high-simulation false anomaly regions, and the model's anomaly detection capability is improved through Huber loss and feature distillation mechanisms.

Benefits of technology

Without manual labeling, the model's response ability and detection accuracy to abnormal areas are significantly improved, and the system's universality and clinical implementation in different patient and equipment environments are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259787A_ABST
    Figure CN120259787A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional medical image anomaly detection method and system based on self-supervised learning, and the method comprises the following steps: respectively constructing three independent feature embedding networks according to three original feature maps with different scales; converting the three original feature maps into one-dimensional embedding vectors by using a feature embedding network; calculating the sum of mean square errors between the embedding vectors of the pseudo abnormal image and the normal image under each scale; performing parameter updating on the feature embedded network by using an Adam optimizer; constructing a decoder which is symmetrical to the encoder in structure; the three embedded vectors are expanded through a full connection layer and remodeled respectively, three low-resolution feature maps are generated, and the spatial resolution of the low-resolution feature maps is the same as that of the layer3 feature map; fusing the three low-resolution feature maps to obtain a layer3 reconstruction feature map; and an up-sampling module of the decoder performs cascade reconstruction on the layer-by-layer reconstructed feature pattern to obtain a layer-1 reconstructed feature pattern and a layer-2 reconstructed feature pattern which are consistent with the resolution of the layer-1 feature pattern and the resolution of the layer-2 feature pattern respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image analysis, and particularly to a three-dimensional medical image anomaly detection method and system based on self-supervised learning. Background Art

[0002] In the field of medical image analysis, anomaly detection technology is one of the current research and application focuses, and is widely used in the early detection and auxiliary diagnosis of diseases. This technology aims to identify abnormal regions that are significantly different from normal anatomical structures in three-dimensional medical images such as CT and MRI. Since abnormal regions usually exhibit characteristics such as sparse distribution, various shapes, and blurred boundaries, existing supervised learning-based methods often rely on a large amount of high-quality labeled data in applications. However, in the actual clinical environment, it is difficult to obtain abnormal samples and the cost of manual annotation is relatively high, resulting in difficulties in popularizing and applying such methods under unlabeled or weakly labeled conditions.

[0003] In recent years, self-supervised learning methods have received increasing attention in the field of anomaly detection. Such methods enable the model to automatically learn the potential structural features of the data itself by constructing prior tasks that do not rely on manual annotation, such as reconstruction, contrast learning, or context prediction. In the scenario of medical image analysis, self-supervised learning can utilize the consistency features among normal samples, enabling the model to identify potential abnormal regions that deviate from the normal distribution during the inference stage. This unsupervised modeling idea has obvious advantages in the absence of abnormal samples, and is particularly suitable for problem scenarios where lesion annotation is scarce in real clinical practice.

[0004] With the development of three-dimensional imaging technology, three-dimensional convolutional neural networks (3D CNNs) have gradually become the mainstream architecture for medical volume data processing because they can simultaneously model spatial and depth dimension features. Compared with traditional two-dimensional slice-based methods, 3D CNNs have more advantages in maintaining the continuity of the context structure and have shown higher accuracy and stability in lesion detection of three-dimensional structure organs such as the lungs and brain.

[0005] However, currently, most three-dimensional anomaly detection methods in medical images are still implemented based on the supervised approach. Anomaly localization in three-dimensional space usually requires pixel-level mask annotation as the training target, which is extremely laborious in the actual data preparation process and limits the applicable range of the model under unlabeled conditions. To reduce label dependence, existing technologies have introduced mechanisms such as mask modeling and region occlusion reconstruction, and simulate abnormal regions by constructing pseudo-labels or local masking to guide the model to learn the difference features between abnormal and normal structures. However, most of these methods still need to rely on a certain proportion of abnormal regions or auxiliary supervision signals, and there is still a risk of unstable model performance under conditions of scarce samples or complex distributions.

[0006] In particular, in the field of three-dimensional medical image anomaly detection, research on adopting self-supervised learning mechanisms and achieving label-free modeling is still relatively limited. How to construct a stable self-supervised learning framework suitable for three-dimensional volume data and effectively improve the model's response ability and detection accuracy for abnormal regions without any manual annotation is still one of the main challenges faced in this technical direction currently. Summary of the Invention

[0007] In view of the above-mentioned defects of the prior art, the present invention provides a method and system for three-dimensional medical image anomaly detection based on self-supervised learning. By introducing a label-free self-supervised training strategy and utilizing the internal consistency and multi-scale reconstruction mechanism of normal samples, accurate detection of abnormal regions in three-dimensional medical images is achieved, solving problems such as the severe dependence of traditional methods on labeled data and the difficulty in obtaining abnormal samples.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, a method for three-dimensional medical image anomaly detection based on self-supervised learning includes the following steps: Step S1: Obtain and preprocess three-dimensional medical image data to generate normal images; add noise to the normal images to generate pseudo-abnormal images; Step S2: Input the normal images and the pseudo-abnormal images into a three-dimensional convolutional neural network encoder for multi-scale feature extraction to obtain three original feature maps with different scales. The original feature maps are sorted from high to low resolution as the layer1 feature map, the layer2 feature map, and the layer3 feature map; Step S3: According to the three original feature maps with different scales, respectively construct three independent feature embedding networks; use the feature embedding networks to convert the three original feature maps into one-dimensional embedding vectors; each feature embedding network processes the original feature map corresponding to its scale; calculate the sum of the mean square errors between the embedding vectors of the pseudo-abnormal images and the normal images at each scale; use the Adam optimizer to update the parameters of the feature embedding networks; Step S4: Construct a decoder symmetric to the encoder structure; use a fully connected layer to expand and reshape the three embedding vectors respectively to generate three low-resolution feature maps. The spatial resolution of the low-resolution feature maps is the same as that of the layer3 feature map; Fuse the three low-resolution feature maps to obtain a layer3 reconstruction feature map; the upsampling module of the decoder cascades and reconstructs the layer3 reconstruction feature map level by level to obtain a layer1 reconstruction feature map and a layer2 reconstruction feature map with resolutions consistent with the layer1 feature map and the layer2 feature map respectively.

[0009] Preferably, in the step S1, the preprocessing includes performing size normalization, pixel intensity normalization, spatial resampling processing, and denoising processing on the three-dimensional medical image data.

[0010] Preferably, in the step S1, during the process of adding noise to the normal image, Perlin noise and Worley noise are mixed and superimposed to generate a perturbed texture; during the noise fusion process, the contribution degrees of the two noises to the texture and the boundary are controlled through a weighting parameter.

[0011] Preferably, in the step S2, for the network structure of the three-dimensional convolutional neural network, the first layer adopts a 7×7×7 three-dimensional convolutional layer with a stride of 2 for extracting low-level texture features; after the 7×7×7 three-dimensional convolutional layer, a batch normalization layer and a ReLU activation function are connected, and then a 3×3×3 three-dimensional max pooling layer is introduced.

[0012] Preferably, in the step S3, each of the feature embedding networks includes: a group of 1×1×1 three-dimensional convolutional layers, a spatial pooling layer connected after the 1×1×1 three-dimensional convolutional layer, and a fully connected layer finally connected.

[0013] Preferably, in the step S4, the Huber loss function is used to maintain the structural consistency of the reconstruction process by minimizing the structural differences between the reconstructed feature maps at each scale and the original feature maps.

[0014] Preferably, it further includes step S5, calculating the cosine similarity between the reconstructed feature maps at each scale and the original feature maps respectively and fusing them to generate an anomaly score map.

[0015] In a second aspect, a three-dimensional medical image anomaly detection system based on self-supervised learning includes: an acquisition and preprocessing module, a pseudo-anomaly generation module, a three-dimensional feature extraction module, a feature embedding module, and a symmetric decoding and reconstruction module; the system is used to implement the steps of the method as described in the first aspect; The acquisition and preprocessing module is used to obtain and preprocess the three-dimensional medical image to generate the normal image; the pseudo-anomaly generation module is used to generate the pseudo-anomaly image according to the normal image; the three-dimensional feature extraction module inputs the normal image and the pseudo-anomaly image into the encoder of the three-dimensional convolutional neural network for multi-scale feature extraction to obtain three original feature maps with different scales; the feature embedding module respectively constructs three independent feature embedding networks according to the three original feature maps with different scales and uses the feature embedding networks to convert the three original feature maps into one-dimensional embedding vectors; the symmetric decoding and reconstruction module is used to construct a decoder symmetric to the encoder structure; and uses the decoder to generate the low-resolution feature map and further generate the reconstructed feature map.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: Different from the prior art that only samples a single noise or local occlusion to construct pseudo-anomalies and mostly reconstructs in a single scale, the present invention proposes to fuse Perlin noise and Worley noise to generate a highly realistic pseudo-anomaly region with delicate texture and natural boundaries; at the same time, a three-scale decoding path completely symmetric to the 3D ResNet-50 encoder structure is constructed to achieve cross-level structure alignment, significantly enhancing the model's ability to distinguish multi-modal anomalies.

[0017] Different from the prior art that simply splices or single-maps multi-scale features, the present invention designs a self-supervised framework of "multi-scale embedding + distillation": for the features of layer1, layer2, and layer3, feature embedding networks are respectively constructed to generate compact embedding vectors, and a robust optimization based on the Huber loss and a feature distillation mechanism with the frozen encoder features as the teacher and the decoded output as the student are introduced to ensure cross-scale semantic consistency during the reconstruction process, greatly improving the detection accuracy of occult or complex boundary lesions.

[0018] Different from the prior art that relies on empirical thresholds or manual parameter tuning, the present invention proposes a cosine similarity distribution based on multi-scale reconstruction error and an F1-driven cross-validation strategy to automatically determine the voxel-level and image-level classification boundaries, realizing threshold self-calibration under unsupervised conditions, and greatly enhancing the universality and clinical feasibility of the system in different patient and device environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is the overall architecture diagram of the three-dimensional medical image anomaly detection method based on self-supervised learning in Embodiment 1 of the present invention; Figure 2 It is the overall architecture diagram of 3D ResNet-50 in Embodiment 1 of the present invention; Figure 3 It is the architecture diagram of the bottleneck layer in 3D ResNet-50 in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0020] In order to make the technical means, creative features, achieved purposes and effects of the invention easy to understand, the present invention will be further described below in conjunction with specific drawings. However, the present invention is not limited to the following implemented cases.

[0021] It should be noted that the structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the efficacy that the present invention can produce and the purpose that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention. Embodiment

[0022] Such as Figure 1 A three-dimensional medical image anomaly detection method based on self-supervised learning as shown, includes the following steps: Step S1: Obtain and preprocess three-dimensional medical image data to generate normal images; add noise to the normal images to generate pseudo-anomaly images.

[0023] Step S1-1: Obtain three-dimensional medical image data and perform preprocessing.

[0024] First, obtain three-dimensional medical image data from a public medical image database or a hospital imaging system. The image data may include image data in formats such as MRI (Magnetic Resonance Imaging), CT (Computed Tomography), etc. Then perform standardization processing on the original image data, that is, preprocessing, to generate normal images. The preprocessing includes size normalization, pixel intensity normalization, and spatial resampling operations. Among them, size normalization is used to adjust all images to a unified spatial resolution to eliminate differences caused by different devices or imaging conditions; pixel intensity normalization adjusts the image gray values to a unified interval through linear stretching or Z-score normalization methods to facilitate subsequent model training; spatial resampling interpolates the image to a fixed voxel spacing to ensure the consistency of three-dimensional structural features. In addition, optionally, perform denoising processing on the image, such as using algorithms such as Gaussian filtering and median filtering to improve the image quality. The three-dimensional medical image after preprocessing, that is, the normal image, will be used as input for the subsequent self-supervised training stage.

[0025] Step S1-2: Apply a forged anomaly data enhancement strategy based on Perlin noise and Worley noise to the data.

[0026] To achieve the self-supervised learning goal, by introducing controllable forged anomaly regions in the normal images, the model is made to learn to recover the original normal anatomical structure from the interference regions, thereby enhancing its ability to identify real anomaly regions in the inference stage. In this embodiment, a fusion generation mechanism of Perlin noise and Worley noise is used to enhance the image to generate pseudo-anomaly regions with complex textures and natural boundaries.

[0027] In the specific implementation process, first, based on the random sampling method, several local voxel regions are selected from the normal images as pseudo-anomaly injection regions. In each selected region, Perlin noise and Worley noise are mixed and superimposed to generate a perturbed texture. Perlin noise is used to construct a multi-scale and continuously varying texture structure, which can simulate the delicate and dense tissue features presented in real lesions; Worley noise is used to introduce spatial distribution inhomogeneity, generating features with blurred boundaries or local cavities, and strengthening the uncertain appearance of the lesion shape. By introducing a weighting parameter in the noise fusion process, the contribution degrees of the two noises to the texture and boundaries are controlled to achieve a perturbation effect closer to the real abnormal appearance.

[0028] The fused noise is superimposed on the normal image to form a pseudo-abnormal image. Finally, the pseudo-abnormal image and its corresponding normal image are input into the model, enabling the model to learn the prior features of the normal tissue structure in the reconstruction task, so as to accurately detect real abnormal regions in the inference stage.

[0029] Step S2: Use a three-dimensional convolutional neural network for multi-scale feature extraction.

[0030] To effectively extract the structural semantic information in three-dimensional medical images, the present invention uses a three-dimensional convolutional neural network as an encoder to perform multi-scale feature extraction on the input image. Preferably, in this embodiment, a 3D ResNet-50 network is used as the three-dimensional convolutional neural network. Figure 2 This is the overall architecture diagram of 3D ResNet-50 in the embodiment of the present invention, including a 3D convolutional layer, batch normalization, ReLU activation layer, and bottleneck layer. Figure 3 This is the architecture diagram of the bottleneck layer in 3D ResNet-50 in the embodiment of the present invention. The 3D ResNet-50 model adapts to the dimensional characteristics of three-dimensional medical images by introducing three-dimensional convolutional operations on the basis of the standard two-dimensional ResNet structure. Its network weights remain frozen during the training process and do not participate in gradient updates, only being used as a feature extractor.

[0031] In terms of the network structure, the first layer uses a 7×7×7 three-dimensional convolutional layer with a stride of 2 to extract low-level texture features. This convolutional layer is followed by a batch normalization layer and a ReLU activation function to enhance the non-linear expression ability of the network and simultaneously suppress the problem of gradient vanishing. Immediately afterwards, a 3×3×3 three-dimensional max pooling layer is introduced for preliminary downsampling to reduce the spatial resolution and expand the receptive field.

[0032] The encoder backbone consists of residual modules in 4 stages (Stage 1 to Stage 4) in sequence. Each stage is composed of multiple stacked 3D residual units, and a cross-layer connection mechanism is adopted to achieve the layer-by-layer accumulation of features and semantic fusion. In the stages from Stage 1 to Stage 3, the network outputs three encoded feature maps of different scales respectively, namely the high-resolution feature map layer1, the medium-resolution feature map layer2, and the low-resolution feature map layer3, which correspond to the structural and semantic information at different spatial resolutions.

[0033] In this embodiment, the normal image and the pseudo-abnormal image are respectively input into the 3D ResNet-50 encoder with shared weights to obtain their corresponding three groups of multi-scale feature maps. Since this encoder does not participate in parameter update, the extracted feature representations will be passed as static inputs to the subsequent trainable feature embedding module for feature alignment and reconstruction target realization.

[0034] Step S3: Feature embedding of normal images and pseudo-abnormal images and optimization of multi-scale feature loss.

[0035] In this embodiment, to achieve the structural restoration of the perturbed region in the pseudo-abnormal image, the multi-scale features output by the encoder need to be further processed by a trainable feature mapping. The present invention constructs corresponding feature embedding networks for the three groups of feature maps of the high-resolution feature map layer1, the medium-resolution feature map layer2, and the low-resolution feature map layer3 output by the encoder, respectively, for mapping the features into compact embedding vectors.

[0036] Specifically, the feature embedding network includes the following structural units: a group of 1×1×1 3D convolutional layers for compressing the channel dimension while keeping the spatial structure unchanged; then obtaining the global feature response of the whole image through a spatial pooling operation; and finally connecting a fully-connected layer to map the pooled result into a one-dimensional embedding vector of a fixed length. The parameters of the above three sub-networks are independent and act on feature maps of different scales respectively to retain the semantic feature distributions at each scale.

[0037] In the training stage, the normal image and its corresponding pseudo-abnormal image are respectively input into the encoder to obtain three groups of feature maps with consistent scales, and then the embedding vectors of layer1, layer2, and layer3 are respectively extracted through their respective feature embedding networks. Based on this, the mean square error (MSE) between the embedding vectors of the pseudo-abnormal image and the normal image at each scale is calculated as the feature reconstruction error at this scale.

[0038] The present invention uses the sum of the above-mentioned multi-scale MSE as the training loss function, and uses the Adam optimizer to update the parameters of the feature embedding network. By minimizing the distance in the feature space between the normal image and the pseudo-abnormal image, the model is guided to learn the ability to recover the normal structure from the perturbed region. As the training iterates, the embedding network gradually improves the ability to repair the perturbed features, thereby enhancing the accuracy and robustness of the overall model in identifying real abnormal regions during the inference stage.

[0039] Step S4: Three-dimensional feature reconstruction and structure restoration based on the symmetric decoding structure.

[0040] In this embodiment, to achieve the feature reconstruction of the normal image, a three-dimensional decoder network symmetric to the encoder structure is constructed to reconstruct the multi-scale feature vectors after feature embedding layer by layer. First, the embedding vectors are respectively expanded and reshaped to obtain their respective low-resolution feature maps. Further, they are fused to obtain the fused low-resolution feature map as the initial input of the decoder. Subsequently, the decoder gradually restores the spatial resolution through cascaded upsampling modules and reconstructs the structural details of the image.

[0041] Specifically, the fully connected layer is used to expand and reshape the three embedding vectors of layer1, layer2, and layer3 output to obtain a feature map with the same spatial resolution as the low-resolution feature map layer3 of the encoder. Further, the three feature maps obtained by expansion and reshaping are fused. The fusion can adopt average fusion or weighted average fusion to obtain the layer3 reconstruction feature map and use it as the input of the decoder. The decoder includes multiple hierarchical decoding units, and each decoding unit consists of an upsampling operation and a three-dimensional convolutional layer for restoring the spatial structure and extracting layer-by-layer semantic features. The upsampling can adopt transposed convolution, trilinear interpolation plus convolution, or other spatial scale restoration methods. During the decoding process, the layer2 reconstruction feature map and the layer1 reconstruction feature map symmetric to the encoder can be obtained, and finally three reconstruction feature maps are obtained, with the resolution consistent with the encoder output.

[0042] To improve the structural consistency of the decoder learning, a feature distillation mechanism is introduced to guide the reconstruction process. Specifically: the original feature maps of layer1, layer2, and layer3 extracted by the frozen encoder network are used as the teacher signals; the reconstruction feature maps output by the decoder are used as the student prediction results; the Huber loss function is used for supervised optimization between the two. By minimizing the structural difference between the reconstruction feature and the original feature at each scale, the decoder is guided to learn the normal representation modeled by the encoder while maintaining the spatial structure, thereby enhancing the model's ability to recover normal features from pseudo-abnormal images.

[0043] Step S5: Abnormal scoring mechanism and discrimination strategy based on multi-scale structural error.

[0044] In the inference stage, an encoder, a trained feature embedding network, and a decoder are used to perform feature extraction, embedding mapping, and structure reconstruction operations on the three-dimensional medical image to be measured, and anomaly scoring is performed based on the reconstruction error. In the specific implementation process, first, the input image is fed into the encoder to extract its multi-scale feature representation, and corresponding multi-scale embedding vectors are generated through the feature embedding module. Subsequently, the embedding vectors are extended and reshaped and then input into the decoder network to obtain three reconstructed feature maps, corresponding to the spatial resolutions of the original layer1, layer2, and layer3, respectively.

[0045] For each scale, the structural error between the reconstructed feature map output by the decoder and the original feature map of the corresponding layer of the encoder is calculated. The error reflects the degree of structural deviation in the embedding space through cosine similarity. The above multi-scale errors are fused by addition to generate an anomaly score map, which can be used to reflect the voxel-level anomaly probability distribution. In practical applications, the anomaly score map can also be statistically aggregated to generate image-level or sample-level anomaly scoring metrics for overall risk assessment.

[0046] To determine the discrimination threshold, a cross-validation method is used to traverse the candidate threshold range on the validation set, and the optimal threshold is selected as the final classification criterion with the F1-score as the evaluation index. During the inference process, when the reconstruction error of a certain area in the image exceeds the threshold, it can be determined that the area is a possible abnormal area; when the global score of the entire image exceeds the threshold, it can be judged that the sample has a structural anomaly.

[0047] Step 6: System implementation of model deployment optimization and efficient inference process.

[0048] To facilitate the deployment and operation of the method of the present invention in the actual environment, a variety of optimization strategies are adopted.

[0049] In the training stage, the mixed-precision training function in the PyTorch framework is used, which can reduce the video memory occupancy and speed up the training speed.

[0050] To adapt to different types of devices, channel pruning and quantization processing of the model are supported to reduce the model size and computational resource requirements, facilitating operation on embedded devices or edge computing devices.

[0051] In the inference stage, a sliding window method is used to process large-size images. The images are input into the model in blocks to improve the operation efficiency. At the same time, asynchronous processing processes for the GPU and CPU are supported, enabling data reading, model calculation, and result output to be carried out simultaneously, further improving the processing speed.

[0052] To verify the effectiveness of the proposed 3D medical image anomaly detection method based on self-supervised learning, 580 brain CT image data were retrospectively collected as the experimental dataset, including 290 normal samples and 290 abnormal samples diagnosed by professional doctors. All image data underwent a unified preprocessing process, including operations such as size normalization, pixel intensity standardization, and spatial resampling, to ensure the consistency of data in spatial resolution and gray-scale distribution.

[0053] Subsequently, the dataset was divided into a training set and a validation set at a ratio of 8:2. Among them, the training set only contains normal samples for the learning of the model in the self-supervised training stage of the present invention; while the validation set contains both normal and abnormal samples for evaluating the model's ability to identify abnormal regions without labeled supervision. Since there are relatively few existing unsupervised or self-supervised anomaly detection methods for 3D medical images, several representative 3D models were introduced as comparison baselines in this experimental design to comprehensively evaluate the performance advantages of the method of the present invention. Among the 3D models introduced in this study, 3D ResNet-18 adopts a standard supervised training process, and 3D STFPM and SSMCTB adopt self-supervised learning methods.

[0054] The experiment used three indicators, namely accuracy, precision, and F1-score, to comprehensively evaluate the detection effects of each model on the validation set. The experimental results are shown in Table 1: Table 1 Experimental Results

[0055] It can be seen from the experimental results that the method of the present invention is superior to other control methods in the three core indicators of accuracy, precision, and F1-score. Especially in terms of the F1-score, it reaches 0.7097, which is about 3.4% higher than that of the traditional supervised method 3D ResNet-18, showing better anomaly detection performance. It is worth noting that the method of the present invention still achieves a performance higher than that of the supervised method without relying on abnormal samples to participate in the training, fully verifying the effectiveness and robustness of the proposed self-supervised multi-scale structure alignment mechanism.

[0056] In summary, the method of the present invention not only shows excellent accuracy, but also has stronger generalization ability and practicality, and is suitable for popularization and use in real clinical scenarios.

[0057] Only certain exemplary embodiments of the present invention have been described by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A three-dimensional medical image anomaly detection method based on self-supervised learning, characterized in that, It includes the following steps: Step S1: Obtain and preprocess three-dimensional medical image data to generate a normal image; add noise to the normal image to generate a pseudo-abnormal image; Step S2: Input the normal image and the pseudo-abnormal image into a three-dimensional convolutional neural network encoder for multi-scale feature extraction to obtain three original feature maps with different scales. The original feature maps are sorted from high to low in resolution as the layer1 feature map, the layer2 feature map, and the layer3 feature map; Step S3: According to the three original feature maps with different scales, respectively construct three independent feature embedding networks; use the feature embedding networks to convert the three original feature maps into one-dimensional embedding vectors; each feature embedding network processes the original feature map corresponding to its scale; calculate the sum of the mean square errors between the embedding vectors of the pseudo-abnormal image and the normal image at each scale; use the Adam optimizer to update the parameters of the feature embedding networks; Step S4: Construct a decoder symmetric to the encoder structure; use a fully connected layer to expand and reshape the three embedding vectors respectively to generate three low-resolution feature maps. The spatial resolution of the low-resolution feature maps is the same as that of the layer3 feature map; Fuse the three low-resolution feature maps to obtain a layer3 reconstructed feature map; the upsampling module of the decoder cascades and reconstructs the layer3 reconstructed feature map level by level to obtain a layer1 reconstructed feature map and a layer2 reconstructed feature map with resolutions consistent with the layer1 feature map and the layer2 feature map respectively.

2. The method for abnormal detection of three-dimensional medical images based on self-supervised learning according to claim 1, wherein In step S1, the preprocessing includes performing size normalization, pixel intensity normalization, spatial resampling processing, and denoising processing on the three-dimensional medical image data.

3. The three-dimensional medical image abnormality detection method based on self-supervised learning according to claim 1, wherein, In step S1, during the process of adding noise to the normal image, Perlin noise and Worley noise are used for mixed superposition to generate a perturbation texture; During the noise fusion process, the contribution degrees of the two noises to the texture and the boundary are controlled through a weighting parameter.

4. The method for abnormal detection of three-dimensional medical images based on self-supervised learning according to claim 1, wherein, In step S2, for the network structure of the three-dimensional convolutional neural network, the first layer uses a 7×7×7 three-dimensional convolutional layer with a stride of 2 to extract low-level texture features; after the 7×7×7 three-dimensional convolutional layer, a batch normalization layer and a ReLU activation function are connected, and then a 3×3×3 three-dimensional max pooling layer is introduced.

5. The three-dimensional medical image abnormality detection method based on self-supervised learning according to claim 2, wherein In step S3, each feature embedding network includes: a group of 1×1×1 three-dimensional convolutional layers, a spatial pooling layer connected after the 1×1×1 three-dimensional convolutional layer, and a fully connected layer finally connected.

6. The method for three-dimensional medical image anomaly detection based on self-supervised learning according to claim 1, wherein, In step S4, the Huber loss function is used to maintain the structural consistency of the reconstruction process by minimizing the structural differences between the reconstructed feature maps and the original feature maps at each scale.

7. The method for abnormal detection of three-dimensional medical images based on self-supervised learning according to claim 6, wherein It further includes step S5: Calculate the cosine similarity between the reconstructed feature maps at each scale and the original feature maps respectively and fuse them to generate an abnormal score map.

8. A three-dimensional medical image anomaly detection system based on self-supervised learning, characterized in that, It includes: An acquisition and preprocessing module, a pseudo-abnormality generation module, a three-dimensional feature extraction module, a feature embedding module, and a symmetric decoding and reconstruction module; The system is used to implement the steps of the method described in Claim 1; The acquisition and preprocessing module is used to acquire and preprocess the three-dimensional medical image to generate the normal image; The pseudo-abnormality generation module is used to generate the pseudo-abnormality image according to the normal image; The three-dimensional feature extraction module inputs the normal image and the pseudo-abnormality image into the encoder of the three-dimensional convolutional neural network for multi-scale feature extraction to obtain three original feature maps with different scales; the feature embedding module respectively constructs three independent feature embedding networks according to the three original feature maps with different scales and uses the feature embedding networks to convert the three original feature maps into one-dimensional embedding vectors; the symmetric decoding and reconstruction module is used to construct the decoder symmetric to the encoder structure; the decoder is used to generate the low-resolution feature map and then generate the reconstructed feature map.

Citation Information

Patent Citations

  • Medical image anomaly detection method and terminal based on unsupervised learning

    CN114155237A

  • Deep learning-based crowd abnormal behavior real-time detection system and method

    CN115620227A

  • Image anomaly detection method based on feature reconstruction and distribution loss

    CN116597255A

  • Image anomaly detection method based on self-supervised learning and knowledge distillation

    CN117934425A

  • Tumor radiotherapy reaction prediction method, system and program product based on three-dimensional image and residual network

    CN120032857A