Lung CT anomaly detection method based on consistent reverse distillation
By introducing a jump connection between the teacher encoder and the student decoder, and designing distillation loss and unified regularization loss based on hard global cosine similarity, the problem of insufficient consistency in the characteristic expression of lung CT image abnormality detection methods in normal regions in the prior art is solved, and more efficient abnormality detection accuracy is achieved.
Patent Information
- Application Number
- CN202510158826.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-06
AI Technical Summary
The existing knowledge-distillation-based abnormal detection method for lung CT images cannot show consistency in the normal area when processing lung CT images, resulting in improved detection performance.
A consistent reverse distillation method based on global regularization and jump connection is proposed. By introducing jump connections between the teacher encoder and the student decoder, the consistency of the internal features of the image is enhanced; distillation loss based on hard global cosine similarity is designed to reduce the error detection rate; unified regularization loss is introduced to optimize the consistent expression of feature between images and between samples.
By paying attention to consistent expressions within the image and between batches, the error detection rate of normal areas is reduced, the sensitivity to abnormal areas is enhanced, and the accuracy of lung CT abnormality detection is significantly improved.
Smart Images

Figure CN120107178A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and in particular relates to a lung CT abnormality detection method based on consistency reverse distillation. Background Art
[0002] According to statistics from the World Health Organization (WHO), pneumonia and other lower respiratory tract infections were the most deadly infectious diseases and the fourth leading cause of death worldwide from 2010 to 2019. For the diagnosis and treatment of such diseases, imaging examinations are an indispensable part, which usually include X-rays and computer tomography (CT). Among them, lung CT, with its high sensitivity and lesion synchronization, can provide doctors with intuitive indicators for evaluating changes in the condition. Despite this, its lesion abnormality detection is heavily dependent on the visual observation of professional doctors, resulting in an inefficient detection process and some false detections mainly caused by subjective errors. Therefore, the research on abnormal lung lesion detection using deep learning methods is of great significance to improving the efficiency of medical imaging diagnosis and has become a research hotspot in the field of computer vision.
[0003] Although lung CT scans play a key role in the diagnosis of lower respiratory tract infections, the variable nature of lesions (such as size, shape, and type) makes image analysis a huge challenge. The data imbalance problem, due to the low incidence of rare lesions, exacerbates this difficulty, and manual annotation of abnormal images is resource-intensive. To address these issues, the research community has explored a variety of solutions, including but not limited to memory-based anomaly detection methods, stream-based anomaly detection methods, and knowledge distillation-based anomaly detection methods, aiming to improve diagnostic efficiency and accuracy. In particular, knowledge distillation-based methods have shown potential in unsupervised anomaly detection due to their excellent knowledge transfer capabilities.
[0004] The performance of the existing lung CT image abnormality detection method based on knowledge distillation needs to be further improved. This is because lung CT images have more subtle abnormal details, which makes the knowledge distillation-based method unable to show the consistency of feature expression in normal areas. Summary of the invention
[0005] The purpose of the present invention is to overcome the defects of the prior art, provide a lung CT anomaly detection method based on consistency reverse distillation, and propose a consistency reverse distillation method based on global regularization and jump connection for anomaly detection in lung CT images. First, in order to enhance the consistency of internal features of the image, a jump connection is introduced between the teacher encoder and the student decoder. This enables the student decoder to utilize the multi-scale features of the image, thereby enhancing the consistent distribution of shallow semantics within the same image. Then, a distillation loss based on hard global cosine similarity is designed. This loss mechanism guides the student decoder and the teacher encoder to achieve consistency in deep semantics within the same image, while effectively reducing the false detection rate of difficult-to-identify samples. In addition, a unified regularization loss is introduced to optimize the distance distribution between images in the same batch and the feature consistency expression between different samples. Finally, anomaly detection is performed by comparing the feature differences between the student decoder and the teacher encoder.
[0006] To achieve the above objectives, the technical solution of the present invention is: a lung CT anomaly detection method based on consistency reverse distillation, and a consistency reverse distillation method based on global regularization and jump connection is proposed for anomaly detection in lung CT images.
[0007] In one embodiment of the present invention, the method is implemented as follows:
[0008] In order to enhance the consistency of internal features of the image, a skip connection is introduced between the teacher encoder and the student decoder, so that the student decoder can utilize the multi-scale features of the image, thereby enhancing the consistent distribution of shallow semantics within the same image;
[0009] Design a distillation loss based on hard global cosine similarity to guide the student decoder and the teacher encoder to achieve deep semantic consistency within the same image, while effectively reducing the false detection rate of hard-to-recognize samples;
[0010] Introducing a unified regularization loss to optimize the distance distribution between images in the same batch and the feature consistency expression between different samples;
[0011] Anomaly detection is performed by comparing the feature differences between the student decoder and the teacher encoder.
[0012] In one embodiment of the present invention, the method constructs a jump connection reverse distillation structure based on an encoder-decoder, which consists of two parts: a frozen pre-trained teacher encoder T E and a trainable student decoder S D ; First, for a given input image, a neural network fully pre-trained on ImageNet is used as the teacher encoder T E To extract the multi-scale features of the image; then, the student decoder SD Through the distillation loss based on hard global cosine similarity and uniform regularization loss Maximizing with Teacher Coder T E The feature mapping similarity between them is used to reconstruct the multi-scale features of the teacher encoder; among them, in the teacher encoder T E and Student Decoder S D Skip connections are introduced between them to prevent normal information loss in multi-scale learning.
[0013] In one embodiment of the present invention, a neural network fully pre-trained on ImageNet is used as the teacher encoder T E To extract the multi-scale features of the image, the specific implementation is as follows:
[0014] Use a neural network pre-trained on ImageNet as the teacher encoder T E , during the training process, T E All parameters of will be frozen; to match T E The multi-scale features of the student decoder S D The structure is symmetrical but with T E The “reverse” one aims to imitate the teacher encoder T E behavior during training; this imitation is multi-scale, covering both local low-level information and global semantic information; for a given normal image Where h is the height, w is the width, and c is the number of channels, which is input into the teacher encoder T E Thus, the multi-scale feature map of layer i is obtained where h i , w i , c i denote the height, width and number of channels of the i-th layer of the encoder respectively; due to the introduction of the skip connection strategy, the student decoder S D The corresponding j-layer multi-scale feature map The definition is as follows:
[0015]
[0016] Use the first 4 layers of ResNet as the teacher encoder T E And reconstruct the output of the first three layers, so i∈ { 1,2,3,4 } ,j∈1,2,3 } ; When i=j, the feature mapping dimensions of the teacher encoder and the student decoder correspond, and the similarity metric is used and To perform knowledge distillation and obtain a two-dimensional anomaly graph M j , defined as follows:
[0017]
[0018] Among them, ||·|| represents l 2 Normalized, M j The larger the value, the greater the abnormal probability of the corresponding position; due to the introduction of the skip connection strategy, The information in will help the student decoder better reconstruct the shallow features of normal images during the testing phase, thereby increasing the feature consistency expression with the teacher encoder.
[0019] In one embodiment of the present invention, the distillation loss based on hard global cosine similarity The details are as follows:
[0020] The anomaly detection method based on knowledge distillation measures the pixel-by-pixel local cosine distance. and The similarity between them is minimized to minimize the two-dimensional abnormal map M on the normal image. j , defined as follows:
[0021]
[0022] in, However, as a binary classification task, anomaly detection focuses on global semantic information, which is the key to judging whether a lung CT image is abnormal. Although the skip connection strategy is introduced, its reconstruction of shallow features ignores the consistent expression of high-level semantic features. In self-supervised contrastive learning, simply using global average pooling may cause all feature points to be confused, thus losing the ability to distinguish between normal and abnormal. Therefore, while retaining and Under the premise of the corresponding relationship, global attention is introduced and the two-dimensional feature map is mapped by flattening operation. Convert to a one-dimensional vector Therefore, the distillation loss based on global cosine similarity is defined as follows:
[0023]
[0024] in, Represents the Flatten operation, exist During the optimization process, the change of one feature will change the entire feature representation and further affect the global cosine distance between the teacher and student networks. Furthermore, we consider how to focus on the areas that are difficult to identify to reduce the false detection rate. In abnormal lung CT images, the identification of subtle areas located at the boundary between abnormal and normal areas is usually a major challenge. These areas are easily misidentified and lead to false detection. Therefore, a hard mining strategy is proposed for areas with small average global cosine distance. Gradient discarding, hard mining The definition is as follows:
[0025]
[0026] Among them, sg(·) is the stop gradient operation, μ(M j ) represents the average global cosine distance within a batch, σ(M j ) is the standard deviation, α H is a hyperparameter that controls the dropout rate; therefore, the final distillation loss based on hard global cosine similarity is The definition is as follows:
[0027]
[0028] In one embodiment of the present invention, the unified regularization loss The details are as follows:
[0029] First, suppose a batch contains N images. For an image x b (b∈N) corresponds to the feature z b , and its corresponding local distance distribution is recorded as Where r represents the radius of the local distance distribution, that is, Represents z b The probability that the relative distance to the corresponding features of other images in the batch is less than r; secondly, let is the uniform distance distribution under the same batch, which is regarded as a one-dimensional uniform distribution on the interval [0, ι]; since z b The nearby local fixed dimension LID can reflect the characteristics of the entire local distribution without obvious errors; therefore, the Euclidean distance paradigm is used as the estimator of LID; specifically, the pairwise Euclidean distance between a batch of sample features is calculated to estimate LID, which is defined as follows:
[0030]
[0031] Where k represents z b The number of nearest neighbors, ω k Yes b The distance to the kth nearest neighbor feature, μ k Yes b The average distance to all k nearest neighbor features; since is a one-dimensional uniform distribution, so Equal to 1; Fisher-Rao metric is used as the optimal solution for measuring distribution distance to unify the distance distribution and local distance distribution The relationship between is considered as maximizing the asymptotic Fisher-Rao distance between the two distributions, or minimizing the negative logarithm of the geometric mean within the local distance distribution for uniform regularization, which is defined as follows:
[0032]
[0033] The overall optimization objective is defined as minimizing the following loss:
[0034]
[0035] in, represents any distillation loss, here is β Reg is a balancing hyperparameter.
[0036] In one embodiment of the present invention, the method obtains the corresponding feature map from the teacher-student network and accumulates it pixel by pixel using a bilinear upsampling operation to construct a final scoring function, which is defined as follows:
[0037]
[0038] Among them, Ψ represents the upsampling operation;
[0039] Finally, use S A The maximum value among them is taken as the abnormality score of the entire image.
[0040] An embodiment of the present application further provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any of the method steps described above.
[0041] An embodiment of the present application further provides a computer-readable storage medium, which includes a computer program. When the computer program is executed on an electronic device, the computer program is used to enable the electronic device to execute any of the method steps described above.
[0042] An embodiment of the present application also provides a computer program product, including a computer program, which is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes any of the method steps described above.
[0043] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a consistent reverse distillation method based on global regularization and jump connection. First, by focusing on the consistent expression within the image and between batches, the expression consistency of the model in normal areas is enhanced, the false detection rate of normal areas is effectively reduced, and the sensitivity of the model to abnormal areas is enhanced. Secondly, by introducing the strategy of jump connection and hard global cosine similarity, the local and global semantic information is fully utilized to enhance the consistent expression within the image. Since lung CT images have fixed anatomical structure information, finally, the consistent expression between batches is achieved by guiding the uniform distance distribution under the same batch and the consistency of the distance distribution of each sample in the batch. The experimental results on a real lung CT medical image dataset show that compared with several state-of-the-art methods, the method of the present invention significantly improves the accuracy of lung CT abnormality detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a structural diagram of the method of the present invention.
[0045] Figure 2 This is the GradCAM visualization result of the method of the present invention on the SC2CT dataset. DETAILED DESCRIPTION
[0046] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0047] The present invention provides a lung CT abnormality detection method based on consistency reverse distillation, and proposes a consistency reverse distillation method based on global regularization and jump connection for abnormality detection of lung CT images. The method is specifically implemented as follows:
[0048] In order to enhance the consistency of internal features of the image, a skip connection is introduced between the teacher encoder and the student decoder, so that the student decoder can utilize the multi-scale features of the image, thereby enhancing the consistent distribution of shallow semantics within the same image;
[0049] Design a distillation loss based on hard global cosine similarity to guide the student decoder and the teacher encoder to achieve deep semantic consistency within the same image, while effectively reducing the false detection rate of hard-to-recognize samples;
[0050] Introducing a unified regularization loss to optimize the distance distribution between images in the same batch and the feature consistency expression between different samples;
[0051] Anomaly detection is performed by comparing the feature differences between the student decoder and the teacher encoder.
[0052] The following is the specific implementation process of the present invention.
[0053] 1 Method Structure
[0054] The method proposed by the present invention is summarized as follows Figure 1 This method constructs a jump connection reverse distillation structure based on encoder-decoder, which consists of two parts: a frozen pre-trained teacher encoder T E and a trainable student decoder S D First, for a given input image, a neural network fully pre-trained on ImageNet is used as the teacher encoder T E To extract the multi-scale features of the image. Then, the student decoder S D Through the distillation loss based on hard global cosine similarity and uniform regularization loss Maximizing with Teacher Coder T E The feature mapping similarity between them is used to reconstruct the multi-scale features of the teacher encoder. E and Student Decoder S D Skip connections are introduced between them to prevent normal information loss in multi-scale learning.
[0055] 1.1 Reverse distillation framework based on local consistency
[0056] The present invention uses a neural network based on ImageNet pre-training as the teacher encoder T E , during the training process, T E All parameters of will be frozen. To match T E The multi-scale features of the student decoder S D The structure is symmetrical but with T E The “reverse” one aims to imitate the teacher encoder T E behavior during training. This imitation is multi-scale, covering both local low-level information (such as color, edges, etc.) and global semantic information. For a given normal image Where h is the height, w is the width, and c is the number of channels, which is input into the teacher encoder T E Thus, the multi-scale feature map of layer i is obtained where h i , w i , c i Represent the height, width and number of channels of the i-th layer of the encoder respectively. Due to the introduction of the skip connection strategy, the student decoder S D The corresponding j-layer multi-scale feature map The definition is as follows:
[0057]
[0058] Use the first 4 layers of ResNet as the teacher encoder T EAnd reconstruct the output of the first three layers, so the factor i∈1,2,3,4 in (1) } ,j∈ { 1,2,3 } When i=j, the feature mapping dimensions of the teacher encoder and the student decoder correspond, and the similarity metric is used and To perform knowledge distillation and obtain a two-dimensional anomaly graph M j , defined as follows:
[0059]
[0060] Among them, ||·|| represents l 2 Normalized, M j A larger value indicates a higher probability of abnormality at that location. The information in will help the student decoder better reconstruct the shallow features of normal images during the testing phase, thereby increasing the feature consistency expression with the teacher encoder.
[0061] 1.2 Distillation loss based on hard global cosine similarity
[0062] Existing anomaly detection methods based on knowledge distillation generally measure the local cosine distance of each pixel. and The similarity between them is minimized to minimize the two-dimensional abnormal map M on the normal image. j , defined as follows:
[0063]
[0064] in, However, as an anomaly detection task is a binary classification task, the key to judging whether a lung CT image is abnormal is to focus on global semantic information. Although the skip connection strategy is introduced, its reconstruction of shallow features ignores the consistent expression of high-level semantic features. In self-supervised contrastive learning, it is common to use global average pooling (i.e., a component of the projection head) to reduce the dimension of the last layer of feature maps, which provides the possibility of introducing globality. But simply using global average pooling may cause all feature points to be confused, thereby losing the ability to distinguish between normal and abnormal. Therefore, while retaining and Under the premise of the corresponding relationship, global attention is introduced and the two-dimensional feature map is mapped using Flatten operation. Convert to a one-dimensional vector Therefore, the distillation loss based on global cosine similarity is defined as follows:
[0065]
[0066] in, Represents the Flatten operation, exist During the optimization process, the change of one feature will change the entire feature representation and further affect the global cosine distance between the teacher and student networks. Furthermore, we consider how to focus on the more difficult to identify areas to reduce the false detection rate. In abnormal lung CT images, the identification of subtle areas located at the boundary between abnormal and normal areas is usually a major challenge. These areas are easily misidentified and lead to false detection. Therefore, a hard mining strategy is proposed for areas with small average global cosine distance. Gradient discarding, hard mining The definition is as follows:
[0067]
[0068] Among them, sg(·) is the stop-gradient operation (Stop-Gradient), μ(M j ) represents the average global cosine distance within a batch, σ(M j ) is the standard deviation, α H is a hyperparameter that controls the dropout rate. Therefore, the final distillation loss based on hard global cosine similarity is The definition is as follows:
[0069]
[0070] 1.3 Distillation consistency between batches
[0071] Since lung CT images have fixed anatomical structure information, images in a dataset should have uniform common information. Therefore, we should consider how to achieve consistent expression of the teacher-student network within a batch. First, assume that a batch contains N images. For a certain image x b (b∈N) corresponds to the feature z b , and its corresponding local distance distribution is recorded as where r represents the radius of the local distance distribution. In other words, Represents z b The probability that the relative distance to the corresponding features of other images in the batch is less than r. is the uniform distance distribution in the same batch, which can be regarded as a one-dimensional uniform distribution on the interval [0, ι]. b The local intrinsic dimensionality (LID) nearby can reflect the characteristics of the entire local distribution without obvious errors. Therefore, a simple Euclidean distance paradigm can be used as an estimator of LID. Specifically, the pairwise Euclidean distance between a batch of sample features is calculated to estimate LID, which is defined as follows:
[0072]
[0073] Where k represents z b The number of nearest neighbors, ω k Yes b The distance to the kth nearest neighbor feature, μ k Yes b The average distance to all k nearest neighbor features. is a one-dimensional uniform distribution, so Equal to 1. Fisher-Rao metric is used as the optimal solution to measure the distribution distance, thereby unifying the distance distribution and local distance distribution The relationship between can be viewed as maximizing the asymptotic Fisher-Rao distance between the two distributions, or minimizing the negative logarithm of the geometric mean within the local distance distribution for uniform regularization, which is defined as follows:
[0074]
[0075] Combined with formula (6), the overall optimization goal can be defined as minimizing the following loss:
[0076]
[0077] in, represents any distillation loss, here is β Reg is a balancing hyperparameter.
[0078] 1.4 Computation and Visualization
[0079] When an abnormal image is input in the test phase, the teacher encoder can still identify the abnormality well due to its pre-trained nature. However, since the student decoder only learns normal information in the training phase, it cannot effectively reconstruct the abnormal features. Therefore, anomaly detection can be performed by measuring the difference between the student decoder and the teacher encoder. According to formula (6), the corresponding feature map can be obtained from the teacher-student network, and it is accumulated pixel by pixel using a bilinear upsampling operation to construct the final scoring function, which is defined as follows:
[0080]
[0081] Among them, Ψ represents the upsampling operation. Finally, use S A The maximum value among them is taken as the abnormality score of the entire image.
[0082] 2 Experiments
[0083] Dataset A was collected from multiple hospitals in location A and contains 2482 lung CT images of 240 patients, of which 1252 are positive CT images of infected patients and 1230 are negative CT images of uninfected patients.
[0084] The B dataset was collected at Medical Center X in Location B. The dataset contains all original CT scans of 377 patients. During data preprocessing, images lacking clear information were excluded to ensure dataset quality, such as some normal images of closed lungs that do not carry information. The final dataset contains 1,053 normal images from 274 patients and 666 diseased images from 68 patients.
[0085] The public datasets used were all produced and edited by the original authors, and the corresponding ethical statements were provided for research purposes. During the experiments, all images were resized to 256×256 pixels. According to the standard protocol for anomaly detection, only negative (normal) images were used during training. Therefore, these datasets were reorganized and divided into 60% training set, 20% validation set, and 20% test set.
[0086] To maximize the fairness of the experimental comparison, all methods use ResNet18 pre-trained in the ImageNet classification task as the teacher encoder model and freeze all parameters during training. They uniformly perform 100 epochs of iterative training, and the training batch size is set to 32. The network of the student decoder is symmetrical with the teacher encoder but trainable. The optimizer uses a learning rate of 5×10 -3 , beta is (0.5, 0.999) Adam. The number of jump connections is set to 2. α in formula (3-5) H It increases linearly from -3 to 1 in the first tenth iteration and remains at 1 in subsequent iterations. The number of nearest neighbors k is uniformly set to 20. For formula (3-9), β of data set A Reg 5×10 -3 , β of the B dataset Reg Then it is 1×10 -3 .
[0087] In order to reduce the impact of randomness, the proposed method takes the average AUROC value and FPR@TPR95 value of three times as the evaluation index of quantitative model performance. AUROC (Area Under the Receiver Operating Characteristic Curve, the area under the receiver operating characteristic curve, that is, the area under the ROC curve) is a commonly used metric for anomaly detection classification, which defines the false positive rate FPR (False Positive Rate, FPR) as the X-axis and the true positive rate TPR (True Positive Rate, TPR) as the Y-axis, as follows:
[0088]
[0089] Among them, TP stands for True Positive, i.e., normal samples predicted by the model to be normal; FP stands for False Positive, i.e., normal samples predicted by the model to be abnormal; TN stands for True Negative, i.e., abnormal samples predicted by the model to be abnormal; and FN stands for False Negative, i.e., abnormal samples predicted by the model to be normal. When the AUROC value is closer to 100%, it means that the classifier can better classify normal and abnormal samples. FPR@TPR95 reflects the probability that an abnormal sample is misclassified as normal when TPR is as high as 95%. The lower the value, the better, and it can be used to evaluate the false detection rate of the model.
[0090] 2.1 Qualitative comparison
[0091] In order to qualitatively compare the lesion detection effect, anomaly localization visualization was performed. Since it is difficult to obtain pixel-by-pixel labels of local abnormalities on lung CT medical images, GradCAM visualization was used to highlight the areas that affect the anomaly detection decision. Figure 2 The abnormality detection results of 9 lung CT images under the SC2CT dataset are shown. These localization results almost meet the doctors' concerns about abnormal areas on lung CT images.
[0092] 2.2 Quantitative comparison
[0093] Table 1 shows the quantitative evaluation results on the two datasets, and the best results are bolded. On the SC2CT dataset, the present invention achieves the highest AUROC value of 86.47% and the lowest FPR@TPR95 value of 61.64%. As the first anomaly detection method combining knowledge distillation and encoder-decoder, RD4AD achieves a poor AUROC value. RD++, ReContrast, and Skip-ST all improve the performance of RD4AD through different strategies. Compared with these methods, the present invention performs better, and the AUROC values of the second and third places are 2.57% and 4.44% higher, respectively, and the corresponding false detection rate is reduced. The reason is that the present invention realizes the feature consistency expression of the teacher-student network from within and between images. In addition, compared with traditional knowledge distillation schemes such as STFPM and CDO, the proposed method also achieves better results.
[0094] Table 1 Quantitative comparison of the proposed method with the most advanced algorithm on real data sets
[0095]
[0096] On the B dataset, the present invention also achieved the highest AUROC value of 90.81% and the lowest FPR@TPR95 value of 43.23%. Skip-ST and ReContrast took the second and third place respectively, but their AUROC values were still 3.71% and 4.93% lower than those of the present invention, and their FPR@TPR95 values were 11.06% and 1.06% higher than those of the present invention. In addition, compared with STFPM and CDO, two anomaly detection methods based on traditional teacher-student networks, the structural similarity between the teacher and student networks easily leads to overfitting in the training process, thus affecting the performance improvement. The present invention solves the above problems through an encoder-decoder structure based on jump connections.
[0097] An embodiment of the present application further provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any of the method steps described above.
[0098] An embodiment of the present application further provides a computer-readable storage medium, which includes a computer program. When the computer program is executed on an electronic device, the computer program is used to enable the electronic device to execute any of the method steps described above.
[0099] An embodiment of the present application also provides a computer program product, including a computer program, which is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes any of the method steps described above.
[0100] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0101] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0102] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0104] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. A lung CT abnormality detection method based on consistency reverse distillation, characterized in that: A consistent back distillation method based on global regularization and skip connections is proposed for abnormality detection in lung CT images.
2. The lung CT abnormality detection method based on consistency reverse distillation according to claim 1, characterized in that: The method is implemented as follows: In order to enhance the consistency of internal features of the image, a skip connection is introduced between the teacher encoder and the student decoder, so that the student decoder can utilize the multi-scale features of the image, thereby enhancing the consistent distribution of shallow semantics within the same image; Design a distillation loss based on hard global cosine similarity to guide the student decoder and the teacher encoder to achieve deep semantic consistency within the same image, while effectively reducing the false detection rate of hard-to-recognize samples; Introducing a unified regularization loss to optimize the distance distribution between images in the same batch and the feature consistency expression between different samples; Anomaly detection is performed by comparing the feature differences between the student decoder and the teacher encoder.
3. The lung CT abnormality detection method based on consistency reverse distillation according to claim 1 or 2, characterized in that: The method constructs a jump connection reverse distillation structure based on encoder-decoder, which consists of two parts: a frozen pre-trained teacher encoder T E and a trainable student decoder S D ; First, for a given input image, a neural network fully pre-trained on ImageNet is used as the teacher encoder T E To extract the multi-scale features of the image; then, the student decoder S D Through the distillation loss based on hard global cosine similarity and uniform regularization loss Maximizing with Teacher Coder T E The feature mapping similarity between them is used to reconstruct the multi-scale features of the teacher encoder; among them, in the teacher encoder T E and Student Decoder S D Skip connections are introduced between them to prevent normal information loss in multi-scale learning.
4. The lung CT abnormality detection method based on consistency reverse distillation according to claim 3 is characterized in that: Use a neural network fully pre-trained on ImageNet as the teacher encoder T E To extract the multi-scale features of the image, the specific implementation is as follows: Use a neural network pre-trained on ImageNet as the teacher encoder T E , during the training process, T E All parameters of will be frozen; to match T E The multi-scale features of the student decoder S D The structure is symmetrical but with T E "Reverse", whose goal is to imitate the teacher encoder T E behavior during training; this imitation is multi-scale, covering both local low-level information and global semantic information; for a given normal image Where h is the height, w is the width, and c is the number of channels, which is input into the teacher encoder T E Thus, the multi-scale feature map of layer i is obtained where h i , w i , c i denote the height, width and number of channels of the i-th layer of the encoder respectively; due to the introduction of the skip connection strategy, the student decoder S D The corresponding j-layer multi-scale feature map The definition is as follows: Use the first 4 layers of ResNet as the teacher encoder T E And reconstruct the output of the first three layers, so i∈{1,2,3,4},j∈{1,2,3} in the above formula; when i=j, the feature mapping dimensions of the teacher encoder and the student decoder correspond, and the similarity metric is used and To perform knowledge distillation and obtain a two-dimensional anomaly graph M j , defined as follows: Among them, ||·|| represents l2 normalization, M j The larger the value, the greater the abnormal probability of the corresponding position; due to the introduction of the skip connection strategy, The information in will help the student decoder better reconstruct the shallow features of normal images during the testing phase, thereby increasing the feature consistency expression with the teacher encoder.
5. The lung CT abnormality detection method based on consistency reverse distillation according to claim 4, characterized in that: Distillation loss based on hard global cosine similarity The details are as follows: The anomaly detection method based on knowledge distillation measures the pixel-by-pixel local cosine distance. and The similarity between them is minimized to minimize the two-dimensional abnormal map M on the normal image. j , defined as follows: in, However, as a binary classification task, anomaly detection focuses on global semantic information, which is the key to judging whether a lung CT image is abnormal. Although the skip connection strategy is introduced, its reconstruction of shallow features ignores the consistent expression of high-level semantic features. In self-supervised contrastive learning, simply using global average pooling may cause all feature points to be confused, thus losing the ability to distinguish between normal and abnormal. Therefore, while retaining and Under the premise of the corresponding relationship, global attention is introduced and the two-dimensional feature map is mapped by flattening operation. Convert to a one-dimensional vector Therefore, the distillation loss based on global cosine similarity is defined as follows: in, Represents the Flatten operation, exist During the optimization process, the change of one feature will change the entire feature representation and further affect the global cosine distance between the teacher and student networks. Furthermore, we consider how to focus on the areas that are difficult to identify to reduce the false detection rate. In abnormal lung CT images, the identification of subtle areas located at the boundary between abnormal and normal areas is usually a major challenge. These areas are easily misidentified and lead to false detection. Therefore, a hard mining strategy is proposed for areas with small average global cosine distance. Gradient discarding, hard mining The definition is as follows: Among them, sg(·) is the stop gradient operation, μ(M j ) represents the average global cosine distance within a batch, σ(M j ) is the standard deviation, α H is a hyperparameter that controls the dropout rate; therefore, the final distillation loss based on hard global cosine similarity is The definition is as follows:
6. The lung CT abnormality detection method based on consistency reverse distillation according to claim 1, characterized in that: Uniform regularization loss The details are as follows: First, suppose a batch contains N images. For an image x b (b∈N) corresponds to the feature z b , and its corresponding local distance distribution is recorded as Where r represents the radius of the local distance distribution, that is, Represents z b The probability that the relative distance to the corresponding features of other images in the batch is less than r; secondly, let is the uniform distance distribution under the same batch, which is regarded as a one-dimensional uniform distribution on the interval [0, ι]; since z b The nearby local fixed dimension LID can reflect the characteristics of the entire local distribution without obvious errors; therefore, the Euclidean distance paradigm is used as the estimator of LID; specifically, the pairwise Euclidean distance between a batch of sample features is calculated to estimate LID, which is defined as follows: Where k represents z b The number of nearest neighbors, ω k Yes b The distance to the kth nearest neighbor feature, μ k Yes b The average distance to all k nearest neighbor features; since is a one-dimensional uniform distribution, so Equal to 1; Fisher-Rao metric is used as the optimal solution for measuring distribution distance to unify the distance distribution and local distance distribution The relationship between is considered as maximizing the asymptotic Fisher-Rao distance between the two distributions, or minimizing the negative logarithm of the geometric mean within the local distance distribution for uniform regularization, which is defined as follows: The overall optimization objective is defined as minimizing the following loss: in, represents any distillation loss, here is β Reg is a balancing hyperparameter.
7. The method for detecting abnormalities in lung CT based on consistency reverse distillation according to claim 5, characterized in that: The method obtains the corresponding feature map from the teacher-student network and accumulates it pixel by pixel using a bilinear upsampling operation to construct the final scoring function, which is defined as follows: Among them, Ψ represents the upsampling operation; Finally, use S A The maximum value among them is taken as the abnormality score of the entire image.
8. An electronic device comprising a processor and a memory, wherein: The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method steps according to any one of claims 1 to 7.
9. A computer-readable storage medium, comprising a computer program, wherein when the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the method steps as claimed in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, wherein the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the method steps as described in any one of claims 1 to 7.
Citation Information
Cited By
OCT fundus image anomaly detection method, system and device and medium
CN120876448A
Medical large model knowledge distillation method and system for medical image segmentation
CN121392536A