Lightweight pulmonary nodule detection method and system based on attention mechanism and structure optimization, and storage medium

The RS-YOLO model with RFA-C2f, Optimized Neck, and LS-Detect modules addresses the challenges of lung nodule detection by enhancing feature extraction and context fusion, improving accuracy and reducing computational complexity for lung nodule detection across varied CT images.

CN120318145APending Publication Date: 2025-07-15SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510234466.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing lung nodule detection algorithms have difficulty in processing complex CT images, especially the detection accuracy of small targets and irregular shapes of lung nodules is not high, and lacks stability and robustness, making it difficult to deal with CT image data of different qualities and equipment.

Method used

Using a lightweight lung nodule detection method based on attention mechanism and structural optimization, the RS-YOLO model, including RFA-C2f, Optimized Neck and LS-Detect modules, uses multi-branch convolution and context information fusion to optimize feature extraction and detection heads to achieve decoupling of training and inference, and reduce parameter redundancy through shared convolution.

Benefits of technology

It significantly improves the accuracy and efficiency of pulmonary nodules detection, reduces the missed detection rate, improves the generalization ability of the model on different data sets, reduces the amount of calculation, and maintains the detection quality and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318145A_ABST
    Figure CN120318145A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight pulmonary nodule detection method and system based on an attention mechanism and structure optimization, and a storage medium, and belongs to the technical field of medical detection. According to the method, an RFA-C2f module is designed, a re-parameterization technology and an attention mechanism are fused, decoupling of training and reasoning is achieved, the ability of a backbone network in the aspect of space and channel feature extraction is remarkably enhanced, and therefore the detection precision is improved; according to the method, a feature fusion structure Optimized Neck module focusing on up-sampling is constructed, a high-resolution up-sampling layer is introduced, a sampling path is simplified, feature redundancy is effectively removed, feature expression granularity is expanded, and therefore the pulmonary nodule omission ratio is reduced; a Lite Shared Desection module is provided, parameter redundancy is reduced through a shared convolution and group normalization technology, the consistency of feature aggregation is improved, and the model precision is kept while the calculated amount is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical detection, and particularly relates to a lightweight pulmonary nodule detection method, system and storage medium based on attention mechanism and structural optimization. Background Art

[0002] Currently, the detection of pulmonary nodules mainly relies on medical imaging technologies such as CT and PET. By improving the detection rate and diagnostic accuracy of pulmonary nodules, lung cancer can be detected earlier and corresponding treatment measures can be taken. However, the detection process of pulmonary nodules faces huge challenges, mainly due to the complexity and variability of pulmonary nodules in CT images, ranging from tiny nodules of a few millimeters to masses of several centimeters. This wide range of sizes requires automatic detection algorithms to have high sensitivity and specificity to avoid misdiagnosis and missed detection. In addition, the shapes of pulmonary nodules are extremely irregular, and may present various forms such as spherical, oval, flat, etc., and may even be accompanied by complex features such as calcification and cavities. These irregular shapes and complex internal features further increase the difficulty of detection. Therefore, it is particularly important to develop an efficient and accurate automatic detection algorithm for pulmonary nodules. This not only requires the algorithm to accurately identify pulmonary nodules of various sizes and shapes, but also needs to have high stability and robustness to cope with CT image data of different qualities and different devices.

[0003] In recent years, the development of deep learning and computer vision technologies has greatly promoted the research progress in the field of pulmonary nodule detection. In 2022, SA Agne proposed a two-stage pulmonary nodule detection framework based on enhanced U-Net and convolutional LSTM networks, combining two-dimensional slice detection and three-dimensional feature extraction, which improved the spatio-temporal consistency and accuracy of detection. In 2023, H Mkindu proposed a pulmonary nodule diagnosis model based on Bayesian optimization and Vision Transformer, enhancing the feature processing and propagation efficiency through a sliding window mechanism, and using Bayesian optimization to adjust hyperparameters to improve the detection performance of chest CT images. In the same year, Z Ji proposed the YOLOv5-CASP model, introducing the CBAM attention mechanism, ASPP module and CoT optimization, enhancing the detection ability of small targets while reducing the model complexity. In 2024, Z UrRehman designed a convolutional neural network (CNN) based on dual attention mechanism, enhancing the feature expression by combining channel and spatial attention, and fully integrating spatial information using global average pooling. These technologies have significantly improved the performance of the model in small target detection, complex background processing and feature extraction through advanced feature extraction and object detection methods.

[0004] Currently, the research on lung nodule target detection mainly focuses on two types of methods: those based on CNN and those based on Transformer. CNN is good at capturing local features, but has limitations in processing global context information. On the other hand, Transformer uses the attention mechanism to significantly enhance the global feature modeling ability, but lacks inductive bias in the case of insufficient data, thus ignoring local details. Considering the special requirements of the lung nodule detection task, an ideal model should be able to accurately process complex boundary features and effectively identify small target regions, which requires the model to capture local details while also being able to understand global context information. Summary of the Invention

[0005] The purpose of the present invention is to provide a lightweight lung nodule detection method, system and storage medium based on the attention mechanism and structural optimization, which can accurately identify lung nodules of various sizes and shapes, and also has high stability and robustness, and can handle CT image data of different qualities and different devices.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] The first purpose of the present invention is to provide a lightweight lung nodule detection method based on the attention mechanism and structural optimization, including the following specific steps:

[0008] S1. Obtain a lung image dataset and preprocess the lung image dataset;

[0009] S2. Perform lung parenchyma segmentation on the original lung CT image dataset generated in step S1, optimize the original lung CT image dataset, retain the region of interest, and construct a dataset for training the RS-YOLO model;

[0010] S3. Construct an RS-YOLO model, where the RS-YOLO model includes RFA-C2f, Optimized Neck, and LS-Detect. The RFA-C2f module is connected in series by two sub-modules, namely the token mixer and the channel mixer, and uses multi-branch convolution and context information fusion. The Optimized Neck module is composed of an upsampling and a C2f module. The detection head of the RS-YOLO model is composed of a shared decoupling module LS-Detect;

[0011] S4. Use the RS-YOLO model obtained in step S3 to extract the image features of the training set of the original lung CT image dataset, obtain a trained CT image lung nodule detection model, and test it on the test set;

[0012] S5. Use the trained RS-YOLO model to process the lung CT images to be detected and obtain the detection results.

[0013] Further, in step S1, the lung image dataset is LUNA16. The lung nodule information of the lung CT images is obtained from the csv file provided by the LUNA16 dataset to generate the original lung CT image dataset.

[0014] Further, in step S2, lung parenchyma segmentation is performed on the original lung CT image dataset generated in step S1, specifically including: using the lung parenchyma mask data provided by the LUNA16 dataset, generating a CT image containing only the lung parenchyma through pixel-by-pixel multiplication operation. The size of the obtained lung nodule image is 512×512 pixels, and the side length range of the gold standard border corresponding to the nodule is 5 to 41 pixels.

[0015] Further, the dataset for training the RS-YOLO model is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2, and all slices of the same patient belong to the same subset.

[0016] Further, the token mixer adopts the reparameterization technique. The specific process is to fuse a 3×3 convolutional layer and a BN layer into a 3×3 convolutional layer. Then, the 7×7 convolutional layer and the BN layer are converted into a 7×7 convolutional layer in the same way. Finally, the weights of the two branches are added to form a 7×7 convolutional layer with bias.

[0017] Further, the Optimized Neck module replaces the traditional upsampling and downsampling dual paths in the Neck with a single-branch upsampling path, and also introduces the high-resolution feature map 160×160 to participate in the fusion.

[0018] Further, the LS-Detect module first unifies the number of channels of the input feature map through three 1×1 convolutional layers, and then defines two consecutive input-output channels as the shared convolutional module of the hidden layer. Among them, the number of channels of the hidden layer controls the depth of feature extraction, and the number of detection layers is dynamically adjusted according to the number of input feature maps; the normalization operation in all shared convolutional modules adopts group normalization.

[0019] The second object of the present invention is to provide a lung nodule detection system. The system is based on the above method and at least includes the following components built in the system:

[0020] A data acquisition module configured to obtain a lung image dataset and preprocess the lung image dataset;

[0021] The dataset construction module is configured to perform lung parenchyma segmentation on the original dataset of lung CT images obtained by the data acquisition module, optimize the original dataset of lung CT images, and retain the regions of interest;

[0022] The processing module is configured to construct an RS-YOLO model, and use the RS-YOLO model to extract the image features of the training set of the original dataset of lung CT images, obtain a trained lung nodule detection model for CT images, and perform tests on the test set; the RS-YOLO model includes RFA-C2f, Optimized Neck, and LS-Detect. The RFA-C2f module is connected in series by two sub-modules, namely the token mixer and the channel mixer, and uses multi-branch convolution and context information fusion. The Optimized Neck module is composed of an upsampling module and a C2f module; the detection head of the RS-YOLO model is composed of a shared decoupling module LS-Detect;

[0023] The detection module is configured to use the trained RS-YOLO model to process the lung CT image to be detected, and obtain a detection result; the detection result includes whether there are lung nodules in the lung CT scan image to be detected, as well as the calibration position and region size of the lung nodules.

[0024] The third object of the present invention is to provide a computer-readable storage medium, which stores a program that can be executed by one or more processors to implement the above-mentioned lightweight lung nodule detection method.

[0025] The fourth object of the present invention is to provide an electronic device, including a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes the instructions of the above-mentioned lightweight lung nodule detection method.

[0026] The English character explanations in the present invention are as follows:

[0027] RFA-C2f: Reparameterized Feature Attention C2f; Optimized Neck: Optimized Feature Fusion Structure; LiteShared Detection: Lightweight Shared Detection Head; AP: Average Precision; Recall: Recall Rate; Parameters: Number of Parameters; ConvModule: Convolution Module; SPPF: Spatial Pyramid Pooling Fast; token mixer: Token Mixer; channelmixer: Channel Mixer; BN: Batch Normalization; Shared Conv: Shared Convolution Module; Reg Conv: Regression Convolution; ClsConv: Classification Convolution; Scale: Scale Factor; GN: Group Normalization; EMA: Efficient Multi-Scale Attention Mechanism; mask: Mask; batchsize: Batch Size; epoch: Epoch; Precision: Precision; AFROC: Area Under the Free-Response Receiver Operating Characteristic Curve; FROC: Free-Response Receiver Operating Characteristic Curve; TPR: True Positive Rate; AP 50 : It represents the average precision calculated in the object detection task when the intersection over union (IoU) between the predicted box and the ground truth box is greater than or equal to 0.5; M: Million.

[0028] Compared with the prior art, the beneficial effects brought by the technical solution provided by the present invention are:

[0029] (1) The present invention provides a lightweight pulmonary nodule detection method based on the attention mechanism and structural optimization to improve the accuracy and efficiency of pulmonary nodule detection. First, the Reparameterized Feature Attention C2f (RFA-C2f) module is designed. This module combines the reparameterization technology and the attention mechanism, decouples training and inference, significantly enhances the backbone network's ability in spatial and channel feature extraction, and thus improves the detection accuracy. Then, an Optimized Neck feature fusion structure focusing on upsampling is constructed. By introducing a high-resolution upsampling layer and simplifying the sampling path, it effectively removes feature redundancy and expands the feature expression granularity, thereby reducing the missed detection rate of pulmonary nodules. Finally, LiteShared Detection (LS-Detect) is proposed. By sharing convolution and group normalization technologies, it reduces parameter redundancy and improves the consistency of feature aggregation, maintaining the model accuracy while effectively reducing the computational amount. Experiments on the LUNA16 dataset show that RS-YOLO improves from 83.7% to 85.31% in AP compared with the baseline model, the Recall increases by 5.12% (from 74.07% to 79.19%), and the Parameters decrease from 3.01M to 1.19M, highlighting its superior detection performance and computational efficiency. In addition, RS-YOLO shows strong generalization ability on different datasets through transfer learning.

[0030] (2) The lightweight pulmonary nodule detection method provided by the present invention realizes the accurate detection of pulmonary nodule regions, improves the inspection efficiency, ensures the detection quality, and improves the stability and efficiency of assisting doctors in disease diagnosis. Brief Description of the Drawings

[0031] Figure 1 It is a schematic diagram of the RS-YOLO model structure provided by the present invention;

[0032] Figure 2 It is a schematic diagram of the RFA-C2f structure provided by the present invention;

[0033] Figure 3 It is a diagram of the equivalent conversion process of reparameterization provided by the present invention;

[0034] Figure 4 It is a schematic diagram of the EMA module;

[0035] Figure 5 It is a comparison diagram of the Neck structure before and after optimization;

[0036] Figure 6 It is a schematic diagram of the Lite Shared Detection module;

[0037] Figure 7 It is a diagram for lung parenchyma extraction. In the diagram, (a) is the original lung CT image, (b) is the mask image corresponding to the lung parenchyma region, and (c) is the lung parenchyma image obtained by multiplying (a) and (b) pixel by pixel;

[0038] Figure 8 It is a diagram of the detection results of solitary and adherent vascular lung nodules by different models;

[0039] Figure 9 It is a diagram of the detection results of ground-glass lung nodules by different models;

[0040] Figure 10 It is a diagram of the detection results of adherent pleural lung nodules by different models;

[0041] Figure 11 It is a comparative diagram of heatmap analysis before and after the optimization of the feature fusion structure;

[0042] Figure 12 It is a schematic diagram of the structure of a lung nodule detection system provided by the present invention. Detailed implementation manners

[0043] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention in combination with specific embodiments and the accompanying drawings. For those not specified in the embodiments regarding specific technologies or conditions, they shall be carried out according to the technologies or conditions described in the literature in this field or according to the product specifications.

[0044] Embodiment 1

[0045] The present invention provides a lightweight lung nodule detection method based on the attention mechanism and structure optimization, which specifically includes the following steps:

[0046] 1. Data collection

[0047] The medical image dataset used in the present invention is LUNA16, which originated from the 2016 LUng Nodule Analysis competition and contains 888 low-dose lung CT images stored in the mhd format. These images were selected from a larger dataset LIDC-IDRI, which has a total of 1018 CT scans. During the selection process, CT data with a slice thickness greater than 3 mm, inconsistent slice spacing, or missing slice information were excluded, and finally 888 CTs were retained.

[0048] LUNA16 provides 1186 lung nodules annotated by at least three experts. According to the provided csv file, this article extracted the slices where the center points of each lung nodule are located, and combined with a small number of multi-nodule slices, and finally obtained 1174 CT images containing 1272 lung nodules.

[0049] 2. Dataset Construction

[0050] To further focus on the lung parenchyma region, the official lung parenchyma mask data is used to generate CT images containing only the lung parenchyma through pixel-by-pixel multiplication operations (as shown in Figure 7 ). In the figure, (a) is the original lung CT image, (b) is the mask image corresponding to the lung parenchyma region, and (c) is the lung parenchyma image obtained by pixel-by-pixel multiplication of (a) and (b). The size of the obtained lung nodule image is 512×512 pixels, and the side length range of the gold standard border corresponding to the nodule is 5 to 41 pixels, reflecting the small target characteristics of the lung nodule detection task. To ensure data independence and scientifically evaluate the model performance, this paper divides the dataset into a training set, a validation set, and a test set according to a ratio of 6:2:2, and ensures that all slices of the same patient belong to the same subset. The specific distribution is shown in Table 1.

[0051] Table 1 The number of images and lung nodules in the dataset.

[0052]

[0053] 3. Construction of the RS-YOLO Model

[0054] The RS-YOLO provided by the present invention is an efficient and lightweight lung nodule detection network proposed based on YOLOv8-n. Its design mainly focuses on three core modules: RFA-C2f, Optimized Neck, and LS-Detect, which are optimized in the backbone network, feature fusion, and detection head respectively, achieving a balance between detection performance and model complexity. The specific structure is shown in Figure 1 .

[0055] The RFA-C2f module constructs the backbone network as the main body. This module effectively optimizes the model parameters and significantly enhances the feature expression ability for complex details. In addition, the ConvModule is responsible for feature extraction and downsampling, and the SPPF module effectively captures global information by expanding the receptive field. The backbone network finally generates multi-resolution feature maps {C2, C3, C4, C5}, providing a rich semantic basis for subsequent feature fusion and object detection.

[0056] The Optimized Neck consists of an upsampling and a C2f module, adopting a single-branch structure focused on upsampling. This design effectively alleviates the problems of small nodule information loss and missed detection caused by excessive downsampling in the traditional Neck structure through deep fusion of feature maps with different resolutions.

[0057] Finally, the detection head of RS-YOLO consists of a lightweight shared decoupling module LS-Detect. This module selects the three feature layers {P3, P4, P5} with the highest resolution as inputs, and uses shared convolution and group normalization to share the parameters of multiple detection heads, which not only optimizes the classification and regression results, but also significantly reduces the computational complexity.

[0058] To better elaborate on the RS-YOLO model structure provided by the present invention, the following details the three modules as follows:

[0059] 3.1. Reparameterized Feature Attention C2f Module

[0060] In the pulmonary nodule detection task, pulmonary nodules usually have the characteristics of small size, blurred boundaries and irregular shapes, which make the detection of small nodules and nodules with weak boundaries particularly difficult. Due to the limitation of local feature extraction ability of traditional convolutional neural networks, they often cannot effectively process these fine-grained features, resulting in low detection accuracy. To solve this problem, the present invention designs the RFA-C2f module (as shown in Figure 2 ), which significantly enhances the receptive field of the model and effectively aggregates multi-scale information through a streamlined residual structure and an innovative RFA strategy.

[0061] As the core part of the RFA-C2f module, RFA is connected in series by two sub-modules, namely token mixer and channel mixer, and uses multi-branch convolution and context information fusion to significantly improve the ability to extract fine-grained features in complex medical images. In the token mixer part of RFA, by combining 3×3 and 7×7 depthwise separable convolutions, rich spatial features are effectively extracted under different receptive fields. Different from the traditional residual structure, the token mixer cancels the residual connection to simplify the module structure and reduce the computational complexity. After BN normalization, the outputs of the two branches are weighted and summed, and the SiLU activation function is used to effectively alleviate the problem of gradient disappearance, thereby improving the stability and non-linear expression ability of the model. The SiLU activation function allows the input to pass through negative values, so as to better capture the subtle changes of the input and avoid excessive loss of information. The token mixer uses the reparameterization technique to decouple training and inference, reducing the model complexity in the inference stage. Specifically, first, a 3×3 convolutional layer and a BN layer are fused into a 3×3 convolutional layer (as shown in (a) of Figure 3 ). Then, the 7×7 convolutional layer and the BN layer are converted into a 7×7 convolutional layer in the same way (as shown in (b) of Figure 3 ). Finally, the weights of the two branches are added together to form a 7×7 convolutional layer with bias (as shown in Figure 3As shown in (c), the equivalent conversion from a multi-branch structure to a single-path structure is achieved.

[0062] The channel mixer part of the RFA module further enhances the context information fusion of multi-scale features. By introducing the EMA mechanism, the model can adaptively weight features at different scales, especially effectively focusing on key information when dealing with complex nodules, improving the model's recognition ability. The EMA mechanism (as Figure 4 shown) enhances the feature representation through grouped interactions, retains the key information between channels, and optimizes the encoding of global spatial information. Specifically, this attention mechanism divides the input tensor into g groups by channel, performs pooling, convolution, normalization, and attention weighting operations on each group separately, and finally fuses and outputs the features, thereby enhancing the overall perception ability of small targets and difficult-to-detect nodules. After the EMA module, two 1×1 pointwise convolutions are then performed, using the GELU activation function to smooth the features, and the residual structure reduces the possible information loss in the deep network, thereby enhancing the perception ability of global spatial information.

[0063] 3.2, Optimized Neck Module

[0064] In the overall framework of RS-YOLO, Optimized Neck, as the core feature fusion module, deeply fuses feature maps with different resolutions by focusing on the single-branch path of upsampling. This module is stacked by the C2f module and the upsampling layer (as Figure 1 shown), where the C2f module optimizes the feature expression ability, and the upsampling layer ensures the alignment and fusion of feature maps at different scales.

[0065] Compared with the traditional Neck structure (as Figure 5 shown in (a)), Optimized Neck (as Figure 5 shown in (b)) effectively alleviates the problem of feature redundancy by optimizing the information transmission path, improving the information fusion efficiency in pulmonary nodule detection. First, the single-branch upsampling path replaces the up-and-down sampling dual paths in the traditional Neck, more efficiently fusing the information of feature maps with different resolutions. Second, introducing the high-resolution feature map C2 (160×160) to participate in the fusion provides rich spatial information for subsequent feature processing, effectively alleviating the loss of small target information caused by multiple downsamplings, thereby reducing the risk of missed detection of small nodules. It is found by comparison that the output resolution after optimization is doubled, effectively enhancing the model's sensitivity to small targets. Overall, Optimized Neck improves the expression ability of feature maps, reduces redundant information, and significantly enhances the ability to capture global spatial information by optimizing the feature fusion strategy. This design provides high-quality feature inputs for the subsequent detection head.

[0066] 3.3, Lite Shared Detection Module

[0067] Traditional detection heads lack consistency when aggregating features at different levels, resulting in parameter redundancy. Especially in detection tasks with small targets or blurred boundaries, this redundancy may increase the risk of false detection and thus lead to performance fluctuations. To address the above problems, the present invention proposes a lightweight shared-parameter detection head LS-Detect (as Figure 6 shown), which makes the feature representations between different detection layers more consistent by sharing convolutional weights, enhances the model's focusing ability on real targets, and reduces misjudgment of non-target regions.

[0068] LS-Detect first unifies the number of channels of the input feature map through three 1×1 convolutions, and then defines two consecutive input-output channels as the shared convolutional module Shared Conv of the hidden layer. Among them, the number of channels of the hidden layer controls the depth of feature extraction, and the number of detection layers is dynamically adjusted according to the number of input feature maps. This flexible design enables the model to adapt to different input conditions and further improves its generalization ability. Through the shared convolutional operation, the three feature maps no longer perform redundant convolutional calculations repeatedly, thereby reducing parameter redundancy and improving computational efficiency. After two Shared Conv operations, they are re-divided into three parallel branches, and each branch independently executes Reg Conv and Cls Conv. In addition, a Scale module is introduced after Reg Conv to learn adjustable scale factors to adaptively adjust the feature amplitude. To maintain light weight without reducing the model's accuracy, Group Normalization (GN) is used for the normalization operation in all convolutional blocks. GN can effectively improve the performance of the detection head in localization and classification tasks. Therefore, LS-Detect uses GN, effectively compensating for the possible deficiency in feature extraction ability caused by the lightweight design.

[0069] 4. Model Evaluation

[0070] Training was performed on an NVIDIA GeForce 4080RTX GPU with 64GB of memory using the PyTorch framework. The batch size used was 32, and the step decay learning rate schedule started from 0.01. The training spanned a total of 300 epochs.

[0071] The performance of the model is comprehensively evaluated in this paper by the following metrics: Average Precision (AP), Precision, Recall, the Area under the Free-response Receiver Operating Characteristic curve (AFROC), Params, and GigaFloating-point Operations (GFLOPs). The definitions of Precision and Recall are shown in Equation 1 and Equation 2, where TP, FP, and FN represent the numbers of true positives, false positives, and false negatives, respectively.

[0072]

[0073] The P-R curve constructed by Precision and Recall, and its area is AP (see Equation 3 in detail, where P is Precision and R is Recall), which comprehensively reflects the balanced performance of the model between precision and recall.

[0074]

[0075] The Free-response Receiver Operating Characteristic curve (FROC) is used to evaluate the recall ability of the model under different numbers of false positives. This curve takes False Positives Per Image (FPPI) as the horizontal axis and True Positive Rate (TPR) as the vertical axis. FPPI represents the average number of false detections per image (see Equation 4 in detail), where N represents the total number of images, and the definition of TPR is the same as Recall. In addition, to reduce the impact of excessive false positives on clinical significance, this paper limits the number of false positives per image to no more than 2. The larger the area of the FROC curve, the stronger the object detection ability of the model under the premise of controlling the number of false detections.

[0076]

[0077] Params represents the total number of parameters of all layers of the model, and its size directly affects the storage requirements of the model, with the unit of M. GFLOPs represents the total floating-point operation volume when the model processes a single image, and it is a commonly used indicator to measure the computational complexity of the model.

[0078] 5. Comprehensive evaluation of the detection performance of the RS-YOLO model

[0079] To comprehensively evaluate the detection performance of the RS-YOLO model, the present invention systematically compared it with classical models SSD, RetinaNet, Faster-RCNN, YOLO series models of the same scale, and RT-DETR (the results are shown in Table 2). The experimental results show that RS-YOLO has achieved significant improvements in key performance indicators such as Recall and AFROC, and optimized the model efficiency through parameter compression. Although the computational complexity has increased slightly, the performance gain is significant, and it is overall superior to the existing comparison models. In terms of the Recall indicator, RS-YOLO has improved by 5.02% compared to the baseline model, and a high recall rate is particularly crucial in medical image detection, reflecting that the model is more sensitive to potential lesion areas. The AFROC value is 1.820, which is superior to all comparison models under low false positive detection conditions, indicating its excellent performance in balancing the recall rate and false alarm rate. At the same time, in terms of the AP 50 and Precision indicators, RS-YOLO has also achieved a steady improvement, proving that it can improve the recall rate without sacrificing detection accuracy, reflecting an effective balance between the overall detection ability and positioning accuracy. In terms of the number of parameters, RS-YOLO compressed the number of parameters from 3.01M to 1.19M through structural optimization, significantly reducing the storage requirements of the model and improving the deployment efficiency. Although the GFLOPs have increased slightly compared to the baseline model, combined with the significant improvements in Recall and AFROC, this computational cost is acceptable.

[0080] To statistically demonstrate the advantages of RS-YOLO, according to the calculation method of formula (4), record the TPR when FPPI is 0.75, 1.00, 1.25, and 1.50. Selecting these four intermediate points can not only avoid the influence of random fluctuations at the edge points on the results, but also more fairly compare the detection performance of the models within a reasonable range, thus more representatively demonstrating the improvement effect. Table 3 shows that RS-YOLO performs best overall. For example, when FPPI is 0.75, which means that 3 false positives are allowed in every 4 images, the TPR of RS-YOLO reaches a high level of 92.50. The p-value in the table is obtained based on T-test statistical analysis. When the p-value is less than 0.05, it indicates that there are significant differences between the two groups of data. The results show that except for YOLOv9 and GELAN, the p-values of other methods and RS-YOLO are all less than 0.05, indicating its significant advantages in target detection ability, sensitivity, and false alarm control, and strong practicality and robustness in practical applications.

[0081] Table 2. Results of the comparison between RS-YOLO and other detection models.

[0082]

[0083] Table 3. Performance of TPR in comparative experiments under different FPPIs.

[0084]

[0085] From the experimental results, RS-YOLO shows high accuracy and stability in detecting different types of pulmonary nodules (isolated type, ground-glass type, and adherent type). Figure 8 、 Figure 9 and Figure 10 respectively show the detection results of each model for the three types of pulmonary nodules, indicating that RS-YOLO shows strong robustness and generalization ability in the pulmonary nodule detection task.

[0086] 1) Isolated pulmonary nodules: Due to the similarity in morphology between blood vessels, bronchi, and lung texture in the pulmonary parenchyma region and pulmonary nodules, the detection difficulty of isolated pulmonary nodules increases significantly. These similar features easily lead to missed or misdetected cases in traditional models. However, RS-YOLO can more accurately capture the boundary contours and positions of isolated nodules, effectively identify their significant features, and thus improve accuracy.

[0087] 2) Ground-glass nodules: The boundaries of pulmonary nodules are blurred and have fewer features, so traditional models are prone to missed detection. However, the experimental results show that RS-YOLO can accurately locate and detect such nodules, demonstrating strong feature extraction and recognition capabilities, significantly superior to other models.

[0088] 3) Adherent pulmonary nodules: They are often closely connected to surrounding blood vessels or tissues, increasing the complexity of detection. RS-YOLO can effectively separate nodules from surrounding tissues, accurately locate adherent nodules, and the detection results have a high consistency with the gold standard, further verifying its adaptability to complex backgrounds.

[0089] Overall, RS-YOLO performs excellently, especially under complex backgrounds and weak boundary conditions, effectively reducing false positives and missed detections, further verifying the robustness and adaptability of the model.

[0090] Example 2

[0091] To verify the contributions of RFA-C2f, Optimized Neck, and LS-Detect to the overall model performance, ablation experiments were also conducted.

[0092] By analyzing the impact of different module combinations on performance, the effectiveness and synergy of each module were clarified, as shown in Table 4. RF, O, and S in the table represent RFA-C2f, Optimized Neck, and LS-Detect, respectively. The results show that after adding the RFA-C2f module alone, both the AP and Precision of the model are improved, effectively compensating for the deficiency of the feature fusion structure in feature extraction ability and significantly improving the model detection accuracy. After using Optimized Neck, the recall rate is increased by 2.33% compared with the baseline model, and the AFROC value is increased from 1.783 to 1.793, proving that the optimized model captures more targets while reducing the false alarm rate, so it is more reliable in practical applications. In addition, the single-branch structure focusing on upsampling reduces the number of model parameters by 0.59M, which helps to remove redundant information and reduces the storage requirements of the model. Although the AP and Precision decrease, the improvement of the optimized model in Recall and AFROC fully proves the enhanced target capture ability of the model. The value of GFLOPs reflects the improvement of the optimized model in computational complexity, which may lead to an increase in inference time, but also means that the model has stronger computational power when processing higher-resolution features. When using LS-Detect alone, the AFROC value of the model is increased from 1.783 to 1.795, greatly reducing the false detection rate of the model. The introduction of the shared parameter strategy aims to achieve lightweight while maintaining the model accuracy. Params and GFLOPs in Table 4 show that compared with before the improvement, the number of model parameters and computational complexity have changed significantly. In the single-class detection task, since both the Cls classification branch and the Box regression branch are related to the class, the effect of the coupled head with shared parameters optimizes the pulmonary nodule detection effect. Table 4 shows that when the three improved modules act simultaneously, AP 50 , Precision, and Recall are increased by 1.61%, 0.82%, and 5.12% respectively, and the AFROC increases from 1.783 to 1.820, indicating that more true positive examples are correctly identified. The Params are reduced by 1.82M without significantly increasing the GFLOPs. The synergy of the three modules further improves the detection performance and achieves the best results in most evaluation metrics.

[0093] To further numerically illustrate the detection differences of each model, the TPRs were calculated when the FPPI was 0.75, 1.00, 1.25, and 1.50 respectively. RS-YOLO showed the best overall performance, indicating that the model performs stably under different numbers of false positives and will not sharply change its sensitivity due to an increase or decrease in false positives. This performance stability is particularly important for the practical application of pulmonary nodule detection, especially when the number of false positives is limited, and it can still maintain a high detection performance. In addition, the p-value of RS-YOLO is less than 0.05, indicating that the improved module has a significant advantage over the baseline model. The specific results are shown in Table 5.

[0094] Table 4. Influence of each module on the detection performance of pulmonary nodules.

[0095]

[0096] Note: RF represents RFA-C2f; O represents Optimized Neck; S represents LS-Detect.

[0097] Table 5. Performance of TPR in ablation experiments under different FPPIs.

[0098]

[0099] Example 3

[0100] Thermodynamic map analysis before and after optimization of the feature fusion structure.

[0101] A CT slice image containing three ground-glass pulmonary nodules was selected, and thermodynamic maps of the output layers P3, P4, and P5 of the feature fusion structure were generated using the HiResCAM method. The annotation box score threshold was set to 0.2. The specific results are shown in Figure 11 . Each row in the figure corresponds to a different feature layer. The baseline model could only detect one pulmonary nodule, and only the P3 layer with the highest resolution extracted effective information. In contrast, the optimized model could detect all pulmonary nodules, and all three feature layers could effectively focus on the pulmonary nodule region. The thermodynamic map shows that the dark regions are more precisely concentrated in the target area, while the light regions provide the necessary global information. This result indicates that the optimized model improves the comprehensiveness of feature expression while focusing on key targets.

[0102] Example 4

[0103] Examine the generalization ability of the model.

[0104] The LIDC-IDRI dataset is introduced, where the training set, validation set, and test set contain 2,869, 679, and 556 lung parenchyma region images respectively. Comparative experiments between the mainstream models and RS-YOLO are conducted using the divided dataset, as shown in Table 6 specifically. To accelerate the training speed and enhance the adaptability to the new dataset, transfer learning is used to freeze-train the backbone network, significantly reducing the training cost. The results show that the AP 50 has increased by 2.80 percentage points, indicating that the detection ability of the model has been enhanced. In particular, Recall has significantly increased by 3.06 percentage points, indicating that the model can better capture more targets. In terms of the AFROC metric, the performance of RS-YOLO is also better than that of the baseline model, demonstrating its ability to effectively improve the recall while controlling the false positive rate. Generally speaking, except for a slight decrease in Precision and a slight increase in GFLOPs, other metrics have been significantly improved, verifying the superior performance and wide adaptability of the model on the new dataset.

[0105] Table 6. Comparative experiments of LIDC-IDRI on different models.

[0106]

[0107] Example 5

[0108] Based on the design of Example 1, this example discloses a lung nodule detection system, as Figure 12 shown. This system is based on the above lightweight lung nodule detection method and at least includes the following components built in the system:

[0109] A data acquisition module, configured to obtain a lung image dataset and preprocess the lung image dataset;

[0110] A dataset construction module, configured to perform lung parenchyma segmentation on the original lung CT image dataset obtained by the data acquisition module, optimize the original lung CT image dataset, and retain the regions of interest;

[0111] A processing module, configured to build an RS-YOLO model, extract the image features of the training set of the original lung CT image dataset using the RS-YOLO model to obtain a trained CT image lung nodule detection model, and conduct tests on the test set; among them, the RS-YOLO model includes RFA-C2f, Optimized Neck, and LS-Detect. The RFA-C2f module is composed of two sub-modules, namely token mixer and channel mixer, in series, using multi-branch convolution and context information fusion. The Optimized Neck module consists of an upsampling and a C2f module; the detection head of the RS-YOLO model is composed of a shared decoupling module LS-Detect;

[0112] The detection model construction module is configured to process the lung CT image to be detected by using the trained RS-YOLO model, and obtain the detection result; the detection result includes whether there are pulmonary nodules in the lung CT scan image to be detected, as well as the calibrated position and regional size of the pulmonary nodules.

[0113] Example 6

[0114] Based on the designs of Example 1 and Example 5, this example discloses a computer-readable storage medium storing a program that can be executed by one or more processors to implement the above-mentioned lightweight pulmonary nodule detection method.

[0115] Example 7

[0116] Based on the designs of Example 1 and Example 5, this example discloses an electronic device, including a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions of the above-mentioned lightweight pulmonary nodule detection method to implement the above-mentioned lightweight pulmonary nodule detection method. The method includes:

[0117] Obtain a lung image dataset and preprocess the lung image dataset;

[0118] Perform lung parenchyma segmentation on the generated original lung CT image dataset, optimize the original lung CT image dataset, retain the region of interest, and construct a dataset for training the RS-YOLO model;

[0119] Construct the RS-YOLO model, where the RS-YOLO model includes RFA-C2f, Optimized Neck, and LS-Detect. The RFA-C2f module is connected in series by two sub-modules, namely the token mixer and the channel mixer, and uses multi-branch convolution and context information fusion. The Optimized Neck module is composed of an upsampling module and a C2f module. The detection head of the RS-YOLO model is composed of the shared decoupling module LS-Detect;

[0120] Extract the image features of the training set of the original lung CT image dataset by using the obtained RS-YOLO model to obtain a trained CT image pulmonary nodule detection model, and perform tests on the test set;

[0121] Process the lung CT image to be detected by using the trained RS-YOLO model to obtain the detection result.

[0122] References:

[0123] [1] Liu, W., et al. Ssd: Single shot multibox detector. in Computer Vision - ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11 - 14, 2016, Proceedings, Part I 14. 2016. Springer;

[0124] [2] Ross, T.-Y. and G. Dollár. Focal loss for dense object detection. in proceedings of the IEEE conference on computer vision and pattern recognition. 2017;

[0125] [3] Ren, S., et al., Faster R - CNN: Towards real - time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence, 2016. 39(6): p. 1137 - 1149;

[0126] [4] Jocher, G. a. C., A. and Qiu, J. Ultralytics yolov8, version 8.0.0, https: / / github.com / ultralytics / ultralytics.2023 ;

[0127] [5] Wang, C.-Y., I.-H. Yeh, and H.-Y. Mark Liao. Yolov9: Learning what you want to learn using programmable gradient information. in European Conference on Computer Vision. 2025. Springer;

[0128] [6]Wang, A., et al., Yolov10: Real-time end-to-end object detection. arXiv preprint arXiv:2405.14458, 2024;

[0129] [7] Zhao, Y., et al. Detrs beat yolos on real-time object detection. in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024.;

[0130] In the case of no conflict, the above-mentioned embodiments and features in the embodiments in this article may be combined with each other.

[0131] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A lightweight pulmonary nodule detection method based on attention mechanism and structure optimization, characterized in that It includes the following specific steps: S1. Obtain a lung image dataset and preprocess the lung image dataset; S2. Perform lung parenchyma segmentation on the original lung CT image dataset generated in step S1, optimize the original lung CT image dataset, retain the region of interest, and construct a dataset for training the RS-YOLO model; S3. Construct an RS-YOLO model, where the RS-YOLO model includes RFA-C2f, Optimized Neck, and LS-Detect. The RFA-C2f module is connected in series by two sub-modules, namely the token mixer and the channel mixer, and utilizes multi-branch convolution and context information fusion. The Optimized Neck module is composed of an upsampling and a C2f module. The detection head of the RS-YOLO model is composed of a shared decoupling module LS-Detect; S4. Use the RS-YOLO model obtained in step S3 to extract the image features of the training set of the original lung CT image dataset, obtain a trained CT image lung nodule detection model, and test it on the test set; S5. Use the trained RS-YOLO model to process the lung CT image to be detected and obtain the detection result.

2. The detection method according to claim 1, wherein In step S1, the lung image dataset is LUNA16. The lung nodule information of the lung CT image is obtained from the csv file provided by the LUNA16 dataset to generate the original lung CT image dataset.

3. The detection method according to claim 2, wherein In step S2, performing lung parenchyma segmentation on the original lung CT image dataset generated in step S1 specifically includes: using the lung parenchyma mask data provided by the LUNA16 dataset, generating a CT image containing only the lung parenchyma through a pixel-by-pixel multiplication operation. The size of the obtained lung nodule image is 512×512 pixels, and the side length range of the gold standard border corresponding to the nodule is 5 to 41 pixels.

4. The detection method according to claim 3, wherein The dataset for training the RS-YOLO model is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2, and all slices of the same patient belong to the same subset.

5. The detection method according to claim 1, wherein The token mixer adopts a reparameterization technique. The specific process is to fuse a 3×3 convolutional layer and a BN layer into a 3×3 convolutional layer. Then, the 7×7 convolutional layer and the BN layer are converted into a 7×7 convolutional layer in the same way. Finally, the weights of the two branches are added to form a 7×7 convolutional layer with a bias.

6. The detection method according to claim 1, wherein The Optimized Neck module replaces the traditional up-and-down sampling dual path in the Neck with a single-branch upsampling path and also introduces a high-resolution feature map of 160×160 to participate in the fusion.

7. The detection method according to claim 1, wherein The LS-Detect module first unifies the number of channels of the input feature map through three 1×1 convolutions, and then defines two consecutive input-output channels as a shared convolutional module of the hidden layer. Among them, the number of channels of the hidden layer controls the depth of feature extraction, and the number of detection layers is dynamically adjusted according to the number of input feature maps; the normalization operation in all shared convolutional modules uses group normalization.

8. A pulmonary nodule detection system, characterized in that, The system is based on the method described in any one of claims 1-7, and at least includes the following components built in the system: A data acquisition module, configured to acquire a lung image dataset and preprocess the lung image dataset; A dataset construction module, configured to perform lung parenchyma segmentation on the original lung CT image dataset acquired by the data acquisition module, optimize the original lung CT image dataset, and retain the region of interest; A processing module, configured to build an RS-YOLO model, extract image features of the training set of the original lung CT image dataset using the RS-YOLO model to obtain a trained CT image lung nodule detection model, and perform tests on the test set; the RS-YOLO model includes RFA-C2f, Optimized Neck, and LS-Detect. The RFA-C2f module is connected in series by two sub-modules, namely the token mixer and the channel mixer, and uses multi-branch convolution and context information fusion. The Optimized Neck module is composed of an upsampling module and a C2f module; the detection head of the RS-YOLO model is composed of a shared decoupling module LS-Detect; A detection module, configured to process the lung CT image to be detected using the trained RS-YOLO model to obtain a detection result; the detection result includes whether there are lung nodules in the lung CT scan image to be detected, as well as the calibration position and region size of the lung nodules.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that can be executed by one or more processors to implement the lightweight lung nodule detection method described in any one of claims 1-7.

10. An electronic device, characterized in that, It includes a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions of the lightweight lung nodule detection method described in any one of claims 1-7.

Citation Information

Cited By

  • Defect detection method and device based on improved YOLO model

    CN121169795A