Semi-supervised multi-view feature alignment pulmonary nodule CT detection algorithm
The semi-supervised multi-view feature alignment algorithm integrates convolutional and Transformer networks to enhance lung nodule detection in CT images, improving precision and speed by leveraging the strengths of both models.
Patent Information
- Application Number
- CN202510499315.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
AI Technical Summary
The existing lung nodule detection algorithm in lung CT images has the problem of insufficient global modeling capabilities of convolutional neural networks and the problem of Transformer being unfriendly to small targets, resulting in a high rate of missed or misdetected detection, which is difficult to meet the needs of lung nodule detection.
The pulmonary nodule CT detection algorithm with semi-supervised multi-view feature alignment is adopted, combined with the object detector based on convolutional neural network and Transformer, and the advantages of the two types of models are integrated through semi-supervised learning technology to build a teacher-student network, and use gradient descent to optimize weight parameters to achieve feature alignment and knowledge transfer, and improve detection accuracy and speed.
The accuracy and speed of lung nodules detection in lung CT images are improved, and the detection performance is optimized, especially in complex scenarios.
Smart Images

Figure CN120318211A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pulmonary nodule detection in lung CT images, and particularly relates to a semi-supervised multi-view feature alignment lung nodule CT detection algorithm. Background Art
[0002] The detection of pulmonary nodules in lung CT images is a core link in the early screening and diagnosis of lung cancer and has important clinical value. As one of the most lethal malignant tumors globally, the five-year survival rate of lung cancer is closely related to early detection. Research shows that the postoperative survival rate of patients with early-stage lung cancer with a diameter less than 2 cm can reach over 80%, while that of advanced-stage patients is less than 20%. The popularization of thin-slice high-resolution CT technology has significantly improved the detection rate of pulmonary nodules, but at the same time, it has also put forward higher requirements for image analysis. Pulmonary nodules usually present as focal round shadows with a diameter ≤ 3 cm, and their morphological, density, and edge features are highly heterogeneous. For example, ground-glass nodules, part-solid nodules, and calcified nodules show significant differences in gray scale and texture in CT images and are easily confused with artifacts such as blood vessel cross-sections and inflammation, resulting in problems such as a missed diagnosis rate (about 20%-30%) and inconsistent subjective interpretations in manual film reading.
[0003] Most existing pulmonary nodule detection algorithms in lung CT images adopt object detection models based on convolutional neural networks, such as Faster RCNN or YOLO series detectors. Convolutional neural networks utilize the locality and parameter sharing characteristics of convolutional operations, with relatively optimized computational complexity, suitable for running on resource-constrained devices. At the same time, they are good at extracting low-level features (such as edges and textures) from local regions and gradually combining them into high-level semantic information, suitable for processing visual tasks. However, the receptive field of convolutional operations is limited to a fixed local area, which may lead to insufficient capture of global context information. Convolutional neural networks have limited ability to model the global relationships between objects, especially in complex scenarios, which may lead to missed detections or false detections. In recent years, object detectors based on Transformer have gradually emerged in the public view. Transformer is based on the self-attention mechanism and can simultaneously focus on all pixels or regions in the image, suitable for capturing global relationships and context information between objects. Transformer detectors like DETR do not require manual design of anchor boxes or region proposal networks, simplifying the model structure. However, Transformer requires a long training time and has a high dependence on large-scale data and pre-training. Therefore, using any single type of object detector alone cannot well meet the needs of pulmonary nodule detection in lung CT images, and how to combine the two types of detection models for use is a problem to be solved. Summary of the Invention
[0004] To solve the above technical problems existing in the prior art, the present invention proposes a semi-supervised multi-view feature alignment lung nodule CT detection algorithm. Aiming at the disadvantages that the object detector based on convolutional neural network has a complex structure and lacks global modeling ability, and the object detector based on Transformer is not friendly to small targets, the semi-supervised learning technology is used to integrate the advantages of the two types of models, effectively improving the detection accuracy of lung nodules in pulmonary CT images, and obtaining faster detection speed and higher detection accuracy.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A semi-supervised multi-view feature alignment lung nodule CT detection algorithm, comprising the following steps:
[0007] Step S1: Collect and preprocess the open-source pulmonary CT image dataset, perform noise addition processing in three degrees of low, medium, and high according to the noise intensity, divide it into labeled and unlabeled parts, and divide it into a training set, a validation set, and a test set according to the deep learning paradigm;
[0008] Step S2: Construct a lung nodule detection network composed of a semi-supervised multi-view feature alignment network, a convolutional-based teacher network, and a Transformer-based student network;
[0009] Step S3: Use the training set of labeled data to train the teacher network, and guide the student network to train with the combination of labeled and unlabeled data through the teacher network. The teacher model in the semi-supervised network initially learns the features of the pulmonary CT image dataset, and then the trained teacher model guides the training of the student network under the labeled dataset and the unlabeled dataset; during the training process, perform gradient descent on the training error, complete the learning of trainable weight parameters, and obtain the trained lung nodule target detection model in the pulmonary CT image, and obtain the target detection model;
[0010] Step S4: Adjust the hyperparameters of the model through the validation set, send the validation set into the lung nodule detection model in the pulmonary CT image trained in Step S3, further estimate the generalization error, and adjust the hyperparameters of the lung nodule target detection model in the pulmonary CT image;
[0011] Step S5: Use the optimized model to detect lung nodules in the test set and the pulmonary CT image to be detected, use the lung nodule detection model in the pulmonary CT image after the hyperparameter adjustment and optimization is completed, complete the detection of the pictures in the pulmonary CT image dataset, evaluate the test results, and then use the qualified lung nodule detection model in the pulmonary CT image to detect the target detection dataset of the pulmonary CT image to be recognized.
[0012] Preferably, the preprocessing in Step S1 includes:
[0013] Normalize all images and unify the input size;
[0014] The division ratio of labeled data to unlabeled data is 2:8, and the division ratio of the training set, validation set, and test set is 6:2:2.
[0015] Preferably, the convolutional-based teacher network in step S2 includes:
[0016] A data augmentation module for rotating, flipping, and cropping the input images;
[0017] A feature extraction module that divides the input image into four pieces and reduces the computational complexity through channel concatenation;
[0018] A feature fusion module that fuses multi-scale features through upsampling and downsampling;
[0019] A classification and regression module for predicting the class and location information of the detection boxes.
[0020] Preferably, the Transformer-based student network in step S2 includes:
[0021] A feature extraction module composed of ResNet34, which outputs three layers of feature maps;
[0022] An encoder that extracts global features based on the last layer of feature maps and performs upsampling fusion;
[0023] A decoupled feature prediction module that contains parallel decoders corresponding to the prediction of large, medium, and small detection boxes respectively.
[0024] Preferably, the semi-supervised multi-view feature alignment network in step 2 realizes knowledge transfer through the feature map alignment of the teacher network and the student network, specifically:
[0025] The i-th layer of feature map s of the student network i and the i-th layer t of the teacher network i and the upsampled feature map t of the i+1-th layer i+1 Calculate the consistency loss;
[0026] The semi-feature alignment loss function of a certain layer of feature map is:
[0027] l i =l(s i ,t i )+l(s i ,up(t i+1 ))
[0028] The total loss function is the sum of the losses of each layer:
[0029] L=∑(loss(si ,t i )+loss(s i ,up(t i+1 )))
[0030] Among them, l represents the l1 loss function, and up represents bilinear interpolation upsampling.
[0031] Preferably, the operations of the feature fusion module include:
[0032] Upsample the high-level features to the low-level resolution step by step to transfer semantic information;
[0033] Downsample the low-level features to the high-level resolution step by step to fuse detailed information.
[0034] Preferably, the parallel decoder of the decoupled feature prediction module interacts with the encoder features through a cross-attention mechanism to generate query vectors and predict detection boxes of different sizes respectively and outputs the final results through the classification and regression modules.
[0035] Preferably, the training of the teacher network and the student network adopts cross-model class alignment, and multi-view feature interaction is performed between the three-layer feature maps t3, t4, and t5 of the teacher network and the three-layer feature maps s3, s4, and s5 of the student network.
[0036] Preferably, the classification module multiplies the class probability by the detection box position accuracy, and the regression module assigns weights according to the detection box size to enhance the prediction ability of targets of different scales.
[0037] Preferably, the test results of the algorithm are evaluated by the detection classification score. Under the same noise conditions, its detection accuracy is higher than that of YOLOv5 and the original RTDETR model.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. Introduce a semi-supervised learning method to combine the advantages of object detectors based on convolutional neural networks and object detectors based on Transformers, solve the respective disadvantages of the two types of models, and obtain an efficient semi-supervised multi-view feature alignment backbone network, which greatly improves the detection performance of the object detector based on Transformers as the student network.
[0040] 2. Refer to the prediction module of the object detector based on convolutional neural networks and propose a decoupled feature prediction module. This prediction model is built based on the DETR series detectors, decomposes the serial decoder structure into a parallel decoder structure, and is used to predict detection targets of different sizes respectively.
[0041] 3. In the semi-supervised learning stage, to ensure the efficiency of knowledge transfer between the teacher model and the student model, a multi-perspective feature alignment module is adopted to enhance the interaction and information transfer between different feature layers. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used in conjunction with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0043] Figure 1 is a flowchart of an algorithm for detecting lung nodules in CT images of the lungs with multiple noises from multiple perspectives in a semi-supervised manner according to an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of constructing a semi-supervised multi-perspective feature alignment network according to an embodiment of the present invention.
[0045] Figure 3 is a schematic diagram of constructing a decoupled feature prediction module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the present invention claimed, but only represents the selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0047] As Figure 1 - Figure 3 shown, an algorithm for detecting lung nodules in CT images of the lungs with semi-supervised multi-perspective feature alignment is proposed in this embodiment.
[0048] An algorithm for detecting lung nodules in CT images of the lungs with semi-supervised multi-perspective feature alignment includes the following steps:
[0049] Step S1: Collect and preprocess an open-source lung CT image dataset, perform noise addition processing at three levels of low, medium, and high according to the noise intensity, divide it into labeled and unlabeled parts, and divide it into a training set, a validation set, and a test set according to the deep learning paradigm;
[0050] Step S2: Construct a lung nodule detection network composed of a semi-supervised multi-perspective feature alignment network, a convolutional-based teacher network, and a Transformer-based student network;
[0051] Step S3: Train the teacher network using the training set of labeled data, and guide the student network to train by combining labeled and unlabeled data through the teacher network. Use gradient descent to optimize the weight parameters to obtain the object detection model;
[0052] Step S4: Adjust the hyperparameters of the model through the validation set;
[0053] Step S5: Use the optimized model to detect pulmonary nodules in the test set and the lung CT images to be detected.
[0054] Furthermore, in step S1, the existing open-source lung CT image dataset is preprocessed according to the semi-supervised learning paradigm, including adding noise to the collected open-source lung CT image dataset. The original dataset is added with noise at three levels of low noise, medium noise, and high noise according to the noise intensity, and then divided into two parts: labeled and unlabeled. Finally, it is divided into a training set, a validation set, and a test set according to the deep learning paradigm. For the input images, all images are normalized to unify the input size to adapt to the input requirements of the teacher model and the student model, which is particularly important for semi-supervised learning to ensure the consistency between pseudo-labeled images and labeled images in different batches.
[0055] Among them, the division ratio of labeled data to unlabeled data is 2:8, and the division ratio of the training set, the validation set, and the test set is 6:2:2.
[0056] Furthermore, the convolution-based teacher network in step S2 includes a data augmentation module, a feature extraction module, a feature fusion module, and a classification and regression module;
[0057] Among them, the data augmentation module is used to perform rotation, flipping, and cropping operations on the input image to achieve data augmentation of the input image, so as to increase the diversity of the input image and enrich the dataset;
[0058] The feature extraction module is used to extract the feature information of the input image. Its main function is to extract low-level and high-level features in the image. First, the input image is divided into four pieces, and the resolution of the feature map is reduced by channel splicing, thereby reducing the computational amount. Subsequently, the feature map is divided into two parts in the channel dimension, and only one part is processed in the backbone network; by separating and fusing partial features, the feature extraction module can effectively reduce the computational amount and prevent the repetition of gradient information, improving the learning ability of the model; after passing through the feature extraction module, the input image will obtain multiple feature maps, and generally the last three layers of feature maps t3, t4, and t5 are sent to the feature fusion module;
[0059] The role of the feature fusion module is to fuse features at different levels to help the model better handle targets of different scales. Upsampling is performed step by step from the high-level features t3, t4, and t5 to the low-level features to obtain t`3, t`4, and t`5, transmitting rich semantic information to the low-level features, which is helpful for small target detection. Subsequently, downsampling is performed step by step from the low-level features to the high-level features, bringing the detailed information at the bottom layer to the high-level features, making the high-level features more conducive to the classification and localization of targets.
[0060] The classification and regression module is mainly used to predict the category and location information of the detection box. The classification module multiplies the category probability by the location accuracy of the detection box, and the regression module assigns weights according to the size of the detection box to enhance the prediction ability for targets of different scales.
[0061] Specifically, the classification module is responsible for predicting the category of the target within the detection box, outputting the probability distribution of each category, and selecting the category corresponding to the maximum probability as the output category y t ; The regression module is responsible for predicting the coordinates and size of the detection box, and describes the position and size of the detection box through multiple regression parameters (such as the center coordinates, width, and height). Both the classification network and the regression network are composed of linear layers. The regression network converts the feature information output by the decoder into position information, and the classification network converts the feature information output by the decoder into category information. To ensure the correlation between the position information and the category information, the category information predicted by the classification network will be multiplied by the accuracy of the corresponding detection box position information, and the regression network will assign corresponding weights according to the size of the detection box during prediction to enhance the prediction ability of the regression network for detection boxes of different scales.
[0062] Furthermore, the Transformer-based student network in step S2 includes a feature extraction module, an encoder, and a decoupled feature prediction module composed of ResNet34;
[0063] Among them, the feature extraction module composed of ResNet34 is mainly composed of a convolutional neural network. After downsampling the input image, it outputs three feature maps, and usually takes the last three feature maps s3, s4, and s5 as the feature information of the image;
[0064] The encoder extracts global features based on the last layer of the feature map and performs upsampling fusion. For the input of the encoder, only the last layer of the feature map s5 is used here to reduce the computational complexity of the self-attention mechanism. After the feature map s5 passes through the encoder, it will be upsampled layer by layer and fused with s3 and s4 to obtain the final three feature maps s`3, s`4, and s`5;
[0065] The decoupled feature prediction module contains multiple parallel decoders. Its main function is to perform object detection based on the feature sequence of the encoder. The decoder uses a query mechanism to generate a series of query vectors, which interact with the features output by the encoder through the cross-attention layer to extract the features corresponding to the target. The decoupled feature prediction module is used to decouple the feature information in the three-layer feature maps s3, s4, and s5 respectively. Each feature map has a dedicated decoder to decode the features and map the three-layer feature maps to detection boxes of large, medium, and small sizes. Different feature maps predict detection boxes of corresponding sizes. Then, through the classification and regression module, the class information and location information of the detection box are obtained.
[0066] Furthermore, the semi-supervised multi-view feature alignment network realizes knowledge transfer through the feature map alignment of the teacher network and the student network. Specifically:
[0067] The i-th layer feature map s of the student network i and the i-th layer t of the teacher network i and the upsampled feature map t of the (i + 1)-th layer i+1 calculate the consistency loss;
[0068] The semi-feature alignment loss function of a certain layer of the feature map is:
[0069] l i = l(s i , t i ) + l(s i , up(t i+1 ))
[0070] The total loss function is the sum of the losses of each layer:
[0071] L = ∑(loss(s i , t i ) + loss(s i , up(t i+1 )))
[0072] Among them, l represents the l1 loss function, and up represents bilinear interpolation upsampling;
[0073] Specifically, for a certain layer of feature s of the student network i , the knowledge it needs to learn not only comes from t i , but also from t i+1 . That is, when calculating the semi-supervised consistency loss between the student network and the teacher network feature maps, the loss function is calculated not only with the feature maps of the same layer, but also with the features of the upsampled feature maps of the next layer. Finally, the semi-feature alignment loss function of a certain layer of the feature map is l i, multi - perspective feature alignment is adopted because the features of the teacher network in the latter layer are the high - level semantic information of the previous layer, containing more fine - grained features, which is beneficial to guiding the learning of the student network.
[0074] Furthermore, the test results of the algorithm are evaluated by detecting the classification score. Under the same noise conditions, its detection accuracy is higher than that of YOLOv5 and the original RTDETR model;
[0075] Specifically, as shown in the following figure, the student model RTDETR in the trained semi - supervised network is used as the inference model, and the classification score of the detected object is used as the evaluation index. Some test results are shown in Figure (a), Figure (b), and Figure (c);
[0076] Figure (a) shows the detection results of the YOLOv5 model for lung nodule pictures in lung CT images. Figure (b) shows the detection results of the RTDETR model for lung nodule pictures in CT images. It can be found that there are differences in the detection performance of the two models for lung nodule pictures in lung CT images under different noises. The worse the weather, the worse the detection performance;
[0077] Figure (c) shows the detection model RTDETR obtained by training with the semi - supervised network. It can be found that under the same type of noise, the detection performance of RTDETR is higher than that of the other two comparison models, verifying the effectiveness of the detection of the lung CT image dataset and improving the detection upper limit of the RTDETR detection model.
[0078]
[0079] Table 1
[0080] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A semi-supervised multi-view feature alignment algorithm for CT detection of pulmonary nodules, characterized in that, It includes the following steps: Step S1: Collect and preprocess the open-source pulmonary CT image dataset, perform noise addition processing at three levels of low, medium, and high according to the noise intensity, divide it into labeled and unlabeled parts, and divide it into a training set, a validation set, and a test set according to the deep learning paradigm; Step S2: Construct a pulmonary nodule detection network composed of a semi-supervised multi-view feature alignment network, a convolutional-based teacher network, and a Transformer-based student network; Step S3: Use the training set of labeled data to train the teacher network, guide the student network to train with labeled and unlabeled data through the teacher network, and optimize the weight parameters using gradient descent to obtain the target detection model; Step S4: Adjust the hyperparameters of the model through the validation set; Step S5: Use the optimized model to detect pulmonary nodules in the test set and the pulmonary CT images to be detected.
2. A semi-supervised multi-view feature alignment-based lung nodule CT detection algorithm according to claim 1, characterized in that, The preprocessing in step S1 includes: Perform normalization processing on all images to unify the input size; The division ratio of labeled data to unlabeled data is 2:8, and the division ratio of the training set, the validation set, and the test set is 6:2:
2.
3. A semi-supervised multi-view feature alignment-based CT lung nodule detection algorithm according to claim 1, characterized in that The convolutional-based teacher network in step S2 includes: A data augmentation module for performing rotation, flipping, and cropping operations on the input image; A feature extraction module that divides the input image into four pieces and reduces the computational amount through channel splicing; A feature fusion module that fuses multi-scale features through upsampling and downsampling; A classification and regression module that predicts the category and location information of the detection box.
4. A semi-supervised multi-view feature alignment-based lung nodule CT detection algorithm according to claim 1, characterized in that The Transformer-based student network in step S2 includes: A feature extraction module composed of ResNet34 that outputs three-layer feature maps; An encoder that extracts global features based on the last layer of feature maps and performs upsampling fusion; A decoupled feature prediction module that contains parallel decoders corresponding to the prediction of large, medium, and small detection boxes respectively.
5. A semi-supervised multi-view feature alignment lung nodule CT detection algorithm according to claim 1, characterized in that The semi-supervised multi-view feature alignment network in step 2 realizes knowledge transfer through the alignment of the feature maps of the teacher network and the student network. Specifically: The i-th layer feature map s of the student network i with the i-th layer t of the teacher network i and the upsampled feature map t of the (i + 1)-th layer i+1 calculate the consistency loss; The semi-feature alignment loss function of a certain layer of feature map is: l i = l(s i , t i ) + l(s i , up(t i+1 )) The total loss function is the sum of the losses of each layer: L = ∑(loss(s i , t i )) + loss(s i , up(t i+1 ))) Among them, l represents the l1 loss function, and up represents bilinear interpolation upsampling.
6. The semi-supervised multi-view feature alignment-based lung nodule CT detection algorithm according to claim 3, wherein, The operations of the feature fusion module include: Gradually upsample the high-level features to the low-level resolution to transfer semantic information; Gradually downsample the low-level features to the high-level resolution to fuse detailed information.
7. A semi-supervised multi-view feature alignment-based lung nodule CT detection algorithm according to claim 4, characterized in that The parallel decoder of the decoupled feature prediction module interacts with the encoder features through a cross-attention mechanism to generate query vectors and predict detection boxes of different sizes respectively. And the final results are output through the classification and regression modules.
8. A semi-supervised multi-view feature alignment lung nodule CT detection algorithm according to claim 1, characterized in that The training of the teacher network and the student network adopts cross-model category alignment, and multi-view feature interaction is performed between the three-layer feature maps t3, t4, t5 of the teacher network and the three-layer feature maps s3, s4, s5 of the student network.
9. A semi-supervised multi-view feature alignment lung nodule CT detection algorithm according to claim 1, wherein The classification module multiplies the category probability by the position accuracy of the detection box, and the regression module assigns weights according to the size of the detection box to enhance the prediction ability of targets at different scales.
10. A semi-supervised multi-view feature alignment-based lung nodule CT detection algorithm according to claim 1, characterized in that, The test results of the algorithm are evaluated by the detection classification score. Under the same noise conditions, its detection accuracy is higher than that of YOLOv5 and the original RTDETR model.