Lightweight pulmonary nodule detection method based on multi-dimensional collaborative attention mechanism

By introducing the MS-YOLO11 detection method with a multi-dimensional synergistic attention mechanism in the YOLO11 model, the problems of low recognition accuracy of small targets and poor adaptability in complex backgrounds in lung nodules detection are solved, and high-precision and stable lung nodules detection are achieved, which is suitable for real-time auxiliary screening of low-computing medical terminals.

CN120339256APending Publication Date: 2025-07-18KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510504321.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the detection of lung nodules, the existing technology has problems such as low recognition accuracy of small targets, complex model structure, and poor adaptability to complex backgrounds. In particular, the detection of lung nodules in CT images is difficult to meet the clinical application standards.

Method used

MS-YOLO11, an improved deep learning detection model, was built, and a multi-dimensional collaborative attention module (MCA) and a collaborative multi-attention Transformer module (SMAT) were introduced to enhance feature understanding capabilities in channel, height and width dimensions, and improve local texture capture capabilities through pixel, channel and spatial attention, which was suitable for unsegmented original CT image detection.

Benefits of technology

It significantly improves the detection accuracy of small targets and the stability of the model, achieves 83.66% detection accuracy, is suitable for low-computing medical terminals and real-time clinical auxiliary screening, and has lightweight deployment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339256A_ABST
    Figure CN120339256A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing, in particular to a lightweight pulmonary nodule detection method based on a multi-dimensional collaborative attention mechanism. According to the method, a multi-dimensional collaborative algorithm named as MS-YOLO11 is provided, based on an improved YOLO11 framework, a multi-dimensional collaborative attention mechanism and a multiple attention transformation module are fused, key feature expression is dynamically enhanced through channel, height and width dimensions, and the local and global feature interaction capability is improved in combination with pixel, channel and space attention. According to the MS-YOLO11 model, 83.66% of detection precision is obtained in an unsegmented pulmonary parenchyma CT image while only 9.35 MB model volume is maintained, the small target detection capability is effectively improved, the background interference robustness is enhanced, and the MS-YOLO11 model is suitable for pulmonary nodule auxiliary diagnosis scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and medical image processing, and particularly to a lightweight pulmonary nodule detection method based on a multi-dimensional collaborative attention mechanism. Background Art

[0002] Pulmonary nodules are important imaging markers in the early screening of lung cancer. They have problems such as small target size, complex morphology, scattered positions, and blurred boundaries in CT images. At the same time, they are accompanied by complex background interferences such as blood vessels and trachea, which are extremely likely to lead to misdiagnosis and missed diagnosis. Traditional detection algorithms such as Faster R-CNN and SSD perform well in large target detection, but have problems such as low recall rate and poor accuracy in the detection of micro-nodules. Although the YOLO series has good real-time performance, it is still difficult to meet the clinical application standards when faced with characteristics such as small size, low contrast, and complex background of pulmonary nodules.

[0003] In addition, existing attention mechanisms such as the SE module and the CBAM module only enhance features from a single dimension of channels or space, and cannot achieve collaborative modeling of multi-dimensional features. Although the Transformer structure has the ability to model global dependencies, it is insufficient in modeling local texture details. Some methods rely on pulmonary parenchyma segmentation to enhance feature focus, but it is easy to cause omission of edge nodule information and affect the integrity of detection. Therefore, there is an urgent need for a pulmonary nodule detection method with multi-dimensional modeling ability, adaptable to complex backgrounds, and capable of taking into account both global and local features. Summary of the Invention

[0004] The purpose of the present invention is to provide a lightweight pulmonary nodule detection method based on a multi-dimensional collaborative attention mechanism to solve the technical problems such as low recognition accuracy of small targets, complex model structure, and poor adaptability to complex backgrounds existing in the prior art.

[0005] To achieve the above purpose, the present invention proposes a multi-dimensional synergistic (MultidimensionalSynergistic) detection method and constructs an improved deep learning detection model MS-YOLO11. The MS-YOLO11 model introduces two key structural modules on the basis of the YOLO11 architecture: a multi-dimensional collaborative attention module (MCA) and a collaborative multi-attention Transformer module (SMAT).

[0006] The MCA module establishes an attention weight mechanism for the input feature map in three dimensions of channels, height, and width respectively, dynamically enhances important features, and thus improves the model's ability to understand spatial position and channel semantic information.

[0007] The SMAT module integrates pixel attention, channel attention, and spatial attention, and further enhances the ability to capture local textures by introducing an enhanced multi-layer perceptron module (E-MLP), adapting to the small and blurred edge features of early lung nodules.

[0008] Preferably, the MCA module is applied to the shallow feature map part of the YOLO11 detection head to improve the positioning accuracy of small targets; the SMAT module is embedded in the deep feature extraction layer of the YOLO11 backbone network to enhance the global feature perception ability.

[0009] Preferably, the model uses the original CT image without segmentation as the data source, and realizes full-image detection without destroying the nodule context structure, improving the detection stability and generalization ability.

[0010] Preferably, the model is trained and validated on the LUNA16 public dataset, and reaches 83.66%, 37.88%, and 50.74% respectively in the three indicators of mAP50, APsmall, and ARsmall, significantly superior to existing mainstream methods such as Faster R-CNN, SSD, YOLOv8, and MSDet.

[0011] The model of the present invention has a volume of only 9.35MB, has the ability of lightweight deployment, and is suitable for low-computing-power medical terminals and real-time clinical auxiliary screening scenarios. Description of the Drawings

[0012] Figure 1 It is a schematic diagram of the MS-YOLO11 detection model structure of the present invention, showing the embedding positions and cooperative effects of the MCA and SMAT modules in the YOLO11 architecture;

[0013] Figure 2 It is a schematic diagram of the multi-dimensional collaborative attention module (MCA) of the present invention, showing the structural path of calculating attention weights for the input feature map in three dimensions: channel, height, and width;

[0014] Figure 3 It is a schematic diagram of the collaborative multi-attention Transformer module (SMAT) of the present invention, including a collaborative multi-attention module (SMA) and an enhanced multi-layer perceptron module (E-MLP);

[0015] Figure 4 It is a trend curve graph of mAP50 changing with Epoch during the training process of the present invention, comparing the detection accuracies of different module combinations;

[0016] Figure 5 It is a curve graph of the Loss value changing with the number of training rounds of the model of the present invention under different experimental settings;

[0017] Figure 6 This is a comparison chart of the detection visualization results of the present invention, the true annotation information, and the YOLO11 model under different data segmentation conditions;

[0018] Figure 7 This is a comparison chart of the detection results of the present invention, the true annotation information, and the YOLO11 model in the lung nodule detection task, with model attention visualization based on heat maps.

[0019] Figure 8 This is a chart of the training parameter settings used in the experiment of the present invention;

[0020] Figure 9 This is a comparison chart of the performance of different model combinations in the ablation experiment of the present invention;

[0021] Figure 10 This is a comparison chart of the comparison results of the present invention and other mainstream detection models under different metrics. Detailed implementation manners

[0022] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0023] Example 1: Model structure construction

[0024] As Figure 1 shown, the MS-YOLO11 proposed by the present invention optimizes the structure on the basis of YOLO11.

[0025] Preferably, the 8th layer C3 module is replaced with an SMAT module to form a global modeling area in the deep feature map (P5 / 32);

[0026] Preferably, MCA modules are respectively embedded in the detection head C3 modules of the 13th, 16th, 19th, and 22nd layers, and cooperate with the FPN structure to achieve attention enhancement at three scales of P3, P4, and P5.

[0027] Example 2: MCA module design

[0028] As Figure 2 shown, the MCA module extracts statistical description information in the channel, width, and height dimensions respectively, and uses 1D convolution to generate attention weights. Among them, the channel branch uses joint modeling of average pooling and standard deviation pooling; the spatial branch realizes directional modeling through tensor dimension transformation, and finally, after the three-branch feature fusion, point-by-point multiplication weighting processing is performed on the original feature map.

[0029] Example 3: SMAT module design

[0030] As Figure 3As shown, the SMAT consists of an SMA module and an E-MLP module. Among them, the SMA module collaboratively obtains multi-dimensional significant regions through channel attention, pixel attention, and spatial attention; the E-MLP structure uses linear projection and 3×3 convolution fusion to enhance the local feature expression ability and strengthen edge detection.

[0031] Example 4: Training and Evaluation Strategies

[0032] Preferably, the LUNA16 publicly available CT image dataset is used to construct two datasets, namely the Undivided dataset and the Segmentation dataset, each divided into a training set and a validation set in a ratio of 9:1. The SGD optimizer is used, with an initial learning rate of 0.01, a momentum coefficient of 0.937, a weight decay of 0.0005, and a total of 700 training epochs. As Figure 8 shown are the training parameter settings used in the experiments of the present invention.

[0033] Preferably, mAP50, APsmall, and ARsmall in the COCO evaluation criteria are used as evaluation indicators. Experiments show that the model of the present invention performs better on the undivided data.

[0034] Example 5: Comparative Analysis

[0035] As Figure 4 shown, the model of the present invention shows a faster convergence speed and higher final accuracy during the training process; as Figure 5 shown, the model of the present invention maintains the lowest Loss value and converges quickly during the training process; as Figure 6 shown, compared with the original YOLO11 and the true annotation information, the MS-YOLO11 method of the present invention detects clearer boundaries of pulmonary nodules and higher confidence on the undivided images; as Figure 7 shown, the heatmap shows the detection performance of the true annotation, the YOLO11 model, and the MS-YOLO11 method of the present invention on multiple samples. The heatmap of the YOLO11 model has a more scattered response to pulmonary nodules, with more noise and misjudgments, limited feature extraction ability, and is easily affected by the background. However, through multi-scale feature fusion and improvement strategies, the MS-YOLO11 method of the present invention has a more concentrated heatmap response, can accurately focus on pulmonary nodules, suppress the response of irrelevant regions, improve the detection ability of small nodules, and reduce the false alarm rate. As Figure 9 shown is the performance comparison of different model combinations in the ablation experiment of the present invention; as Figure 10 shown are the comparison results of the present invention with other mainstream detection models under different metrics.

[0036] In summary, the lightweight pulmonary nodule detection method based on the multi-dimensional collaborative attention mechanism proposed by the present invention combines local and global feature modeling techniques, can effectively improve the detection accuracy, stability and deployment efficiency, and has broad clinical application prospects.

Claims

1. A lightweight pulmonary nodule detection method based on a multi-dimensional collaborative attention mechanism, characterized in that: This method constructs a deep learning model called MS-YOLO11 for detecting pulmonary nodule targets in lung CT images. The model integrates a multi-dimensional collaborative attention mechanism and a multiple attention transformation module based on the YOLO11 structure, including a backbone network, a feature fusion structure, and a detection head. Among them, a collaborative multi-attention transformation module is integrated into the backbone network to enhance the global modeling ability of deep features, and a multi-dimensional collaborative attention module is embedded in the detection head to enhance the small target feature expression at different scales and improve the localization accuracy and robustness of the model for pulmonary nodules.

2. The method according to claim 1, wherein: The collaborative multi-attention transformation module consists of a collaborative multiple attention mechanism and an enhanced multi-layer perceptron unit. The collaborative multiple attention mechanism fuses pixel attention, channel attention, and spatial attention, and enhances the input feature map through residual connection and normalization operations to highlight the edge structure, density change, and key texture of pulmonary nodules. The enhanced multi-layer perceptron unit strengthens the local area detail capture ability based on convolution and activation combination operations, and optimizes the recognition performance of the small nodule edge area.

3. The method according to claim 1, wherein: The multi-dimensional collaborative attention module independently models the attention weights in three dimensions: channel, height, and width, and performs weighted modulation on the input feature map. The module dynamically generates a feature enhancement matrix by statistically analyzing the activation responses in different dimensions, and realizes multi-directional feature enhancement of suspected lesion areas in lung images while keeping the model calculation cost low.

4. The method according to claim 3, wherein: The multi-dimensional collaborative attention module is inserted at multiple positions in the detection head, and acts on feature maps of different resolution levels respectively to improve the small target detection ability of the model on shallow and deep features. Among them, the shallow enhancement focuses on the expression of pulmonary nodules in the edge area, and the deep layer improves the target recognition performance in complex backgrounds.

5. The method according to claim 1, wherein: In the backbone network, the high-level module of the original YOLO11 model is replaced with a collaborative multi-attention transformation structure to enhance the global context feature extraction ability of pulmonary nodules. This structure establishes long-range dependence relationships between lung structures through a multi-layer attention mechanism, so as to more accurately identify pulmonary nodules of various shapes and different sizes.

6. The method according to claim 1, characterized in that: This method uses the original chest CT image without pulmonary parenchyma segmentation as input data, avoids the truncation problem of nodule boundaries caused by traditional segmentation operations, maintains the integrity of lung context information, and effectively improves the small nodule detection rate and boundary localization accuracy.

7. The method according to claim 1, characterized in that: The model is an end-to-end detection framework, which uses a multi-scale feature pyramid structure that combines top-down and bottom-up fusion for feature integration, and outputs the bounding box coordinates, target category, and confidence score of nodules, suitable for automated pulmonary nodule screening and clinical auxiliary diagnosis.

8. The method according to claim 1, wherein: The overall weight of the MS-YOLO11 model does not exceed 10MB, adapts to edge computing devices and lightweight deployment environments, and has the characteristics of low resource consumption, high accuracy, and high real-time performance in pulmonary nodule detection, suitable for pulmonary health screening tasks in primary hospitals or mobile terminal environments.

9. The method according to claim 1, wherein: The model was trained and tested on the unsegmented images of the LUNA16 public dataset, achieving an average detection accuracy of 83.66%. It outperformed existing models such as the YOLO series, Faster R-CNN, SSD, and MSDet in terms of the small object detection metrics AP and AR, and has significant engineering practicality and clinical promotion value.

Citation Information

Cited By

  • Deep learning-based endoscopic anatomical structure auxiliary detection and identification system and method

    CN121685496A