Method and system for detecting miniature focus of pulmonary tuberculosis based on multi-scale convolution
By improving the YOLOv10 network model, using the C2f-Faster-EMA module and the multi-scale convolution module, combined with the WIOUv3 loss function, the problems of misdiagnosis and misdiagnosis in the detection of micro-lesions of pulmonary tuberculosis were solved, significantly improving the average accuracy and recall of the detection, and assisting doctors to improve diagnosis accuracy.
Patent Information
- Application Number
- CN202510578992.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has problems of misdiagnosis and misdiagnosis when detecting micro lesions of pulmonary tuberculosis, especially in early micro lesions detection, the accuracy of doctors and neural networks assisted diagnosis is insufficient, resulting in poor clinical diagnosis effect.
The YOLOv10 network model is improved based on multi-scale convolution. By using the C2f-Faster-EMA module and the multi-scale convolution module in the backbone network, the feature extraction and multi-scale feature fusion capabilities are improved. Combined with the WIOUv3 loss function, the average accuracy and recall rate of target detection are optimized.
The improved network model significantly improves the average accuracy and recall rate of target detection of micro lesions, assists doctors in accurately determining the location of lesions in CT images, improves diagnosis accuracy and efficiency, and reduces the work burden of medical staff.
Smart Images

Figure CN120107251A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and relates to a method and system for detecting micro-lesions of pulmonary tuberculosis based on multi-scale convolution. Background Art
[0002] Since there are more micro lesions in the early stage of tuberculosis, micro lesions are easily missed by doctors through CT scans because they are judged based on the size, shape, posture, density and other factors of the nodules. For example, cavities and tree-bud signs are mostly micro pulmonary nodules, and the tree-bud sign is also the least obvious micro pulmonary nodules. The effect of neural networks directly applied to medical images is not satisfactory, with a high risk of missed diagnosis and misdiagnosis, and poor clinical diagnosis accuracy. Due to the insufficient accuracy of auxiliary diagnosis, a large number of doctors are required to jointly diagnose patients with micro lesions. Summary of the invention
[0003] The purpose of the present invention is to provide a method and system for detecting micro-lesions of pulmonary tuberculosis based on multi-scale convolution, and to improve the network model framework of YOLOv10. The improved network model improves the average accuracy of micro-lesion target detection and has a more effective recall rate.
[0004] The technical solution to achieve the purpose of the present invention is: A method for detecting micro-lesions of pulmonary tuberculosis based on multi-scale convolution comprises the following steps: S01: Acquire lung CT image data; S02: Feature extraction is performed through the convolution layer of the neural network detection model of lung nodule lesions. After the extracted feature map is input into the Faster and efficient multi-scale attention fusion modules, the outputs of different levels in the feature pyramid in the backbone network are input into the corresponding Adown module, Conv module and UpSample module in the multi-scale convolution module. Multiple feature information is obtained through multi-scale convolution, and finally a prediction box is generated to obtain the detection result.
[0005] In the preferred technical solution, the neural network detection model of pulmonary nodule lesions described in step S02 is improved on the basic framework of YOLO v10, and the C2f module of the backbone network and the neck network in the YOLOv10 model is replaced with a C2f-Faster-EMA module, and the Bottleneck in the C2f module is replaced with a Faster module to obtain a C2f-Faster-EMA module. After the last effective feature extraction layer of the backbone network, feature information is obtained from pooling windows of different sizes through a spatial pyramid pooling module, and a PSA module is added at the end of the backbone network. The features are divided by the PSA module and self-attention is applied to some features. In the backbone network, feature maps of multiple scales are output for use in the multi-scale convolution module of the neck network.
[0006] In the preferred technical solution, the C2f-Faster-EMA module includes a first Conv layer, a Split layer, a first Faster module, a second Faster module, a Concat layer and a second Conv layer connected in sequence, and the Split layer output feature map and the feature maps of the first Faster module and the second Faster module are input into the Concat layer together.
[0007] In the preferred technical solution, each Faster module has a PConv layer, followed by two PWConv layers or Conv layers, and an EMA attention mechanism module is added to form an inverted residual block, in which the middle layer has more channels and a shortcut connection is placed to reuse the input function, and its output is added element by element with the PConv layer output; The EMA attention mechanism module calculates the attention score by the dot product of the query vector Q and the key vector K to obtain the attention weight matrix A. The efficient multi-scale attention mechanism smoothing mathematical expression is applied to the attention weight matrix A as follows: ;
[0008] ;
[0009] in, is the dimension of the key vector K, represents the smoothing coefficient; The output is calculated by smoothing the attention weights, and the smoothed attention weights are weighted summed over the value vector V to obtain the final output representation.
[0010] In the preferred technical solution, the input end of the multi-scale convolution module is divided into three branches, which are respectively connected to the ADown module, the Conv module and the upsample module, and then fused together through a Concat, and then enter a parallel group deep separable convolution. Multiple separable volumes can exist in the parallel group deep separable convolution at the same time. Multiple separable convolutions are added into a feature map, which is added to the parallel group deep separable convolution after passing through the Conv module and the BAM attention mechanism module.
[0011] In the preferred technical solution, the parallel group depth-separable convolution is expressed as: ;
[0012] in, for The convolution extracts local feature information. To customize the size of the convolution kernel, Indicates that the feature map value is a real number is a three-dimensional tensor whose elements are real numbers and whose dimensions are the number of channels ,high ,width , In the backbone network Layer input feature map, The second scale feature of the branch, is the nth parallel branch convolution, is the number of layers for extracting features from the backbone network, For the mth DSConv, middle There are 4 different depth-separable convolutions, is the size of the convolution kernel, , , Represent the number of channels, height and width respectively; The relationship between each channel is expressed as follows: ;
[0013] in, Represents the output features, and 1×1 convolution is the channel fusion mechanism.
[0014] In the preferred technical solution, the parallel group depth separable convolution output point multiplication in the multi-scale convolution module is added to the output of the point multiplication. For a given input feature map , BAM attention mechanism infers a 3D attention map , the channel attention branch is calculated on two independent attention branches and the spatial attention branch , and finally calculate the attention map : ;
[0015] In the formula, is the input feature map, The output is an attention-weighted feature map. represents element-by-element multiplication, is a sigmoid function, and the feature maps of the spatial attention branch and the channel attention branch are real numbers. Its elements are real numbers and its dimension is the number of channels. ,high ,width ; In the spatial attention branch, 1×1 convolution is used to reduce the dimension. The dimensionality reduction formula is: ;
[0016] In the formula, represents the convolution operation, represents the feature map, Represents a batch normalization operation; The channel attention branch performs global average pooling on the feature map F and generates a channel vector , encode the global information of each channel, add a hidden layer of multi-layer perceptron to obtain cross-channel attention, the mathematical expression is: ;
[0017] In the formula, is a multi-layer perceptron, is the average pooling layer, To reduce the dimension, we transform the input from Dimension compression to dimension, To adjust the feature distribution after dimensionality reduction, To restore the dimension from Restore to dimension, It is the benchmark value that affects the final channel weight.
[0018] In the preferred technical solution, the neural network detection model of pulmonary nodule lesions uses the WIOUv3 loss function as the regression loss of the bounding box, includes a dynamic non-monotonic mechanism and gradient gain allocation, and constructs distance attention by distance metric to obtain WIoUv1 with a two-layer attention mechanism. WIoUv1 introduces a dynamic weight mechanism based on the traditional IoU, and automatically adjusts the loss weight according to the scale of the target box or the regression difficulty, so that the model pays more attention to difficult samples or small targets. The mathematical expression is: ;
[0019] In the formula, Used to measure the matching degree between the predicted box and the real box. is the overlap between the real box and the predicted box, and is the width and height of the minimum prediction box, and are the variances of the width and height differences between the predicted box and the true box; A non-monotonic focusing coefficient is constructed using the outlier degree β and applied to WIoU v1 to obtain WIoU v3 with dynamic non-monotonic FM. FM assigns larger gradient gains to these anchor boxes. The dynamic non-monotonic FM is used to determine the gradient gain allocation strategy. The mathematical expression is: , , ,in is the gradient gain, and is a hyperparameter that controls the shape of the gradient gain curve. To adjust the sharpness of the gain, To determine the peak position of the gain, is an exponential move with kinetic energy m.
[0020] The present invention also discloses a pulmonary tuberculosis micro-lesion detection system based on multi-scale convolution, comprising: An image data acquisition module, which acquires lung CT image data; The pulmonary nodule lesion detection module extracts features through the convolution layer of the neural network detection model of pulmonary nodule lesions. After the extracted feature map is input into the Faster and efficient multi-scale attention fusion modules, the outputs of different levels in the feature pyramid in the backbone network are input into the corresponding Adown module, Conv module and UpSample module in the multi-scale convolution module. Multiple feature information is obtained through multi-scale convolution, and finally a prediction box is generated to obtain the detection result.
[0021] The present invention further discloses a computer storage medium on which a computer program is stored. When the computer program is executed, the above-mentioned method for detecting micro-lesions of pulmonary tuberculosis based on multi-scale convolution is implemented.
[0022] Compared with the prior art, the present invention has the following significant advantages: The present invention improves on the network model framework of YOLOv10, and the improvement mainly includes the use of C2f-Faster-EMA module in the backbone network to improve the performance of feature extraction of micro-tuberculosis lesions. A multi-scale convolution module (MSCM) is added to the feature fusion network to process features of different scales, capture context feature information across multiple scales, obtain multiple feature information, and increase the ability of information interaction. The overall network model uses the WIoUv3 loss function to not only reflect the degree of overlap between the predicted frame and the real frame, but also reflects that large and small objects do not affect each other under cover. The improved network model improves the average accuracy of micro-lesion target detection and a more effective recall rate. The network model assists doctors in diagnosing micro-lesions and accurately determines the location information of lesions in CT images for the subsequent basic diagnosis and treatment, while also reducing the workload and energy of medical staff. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Flow chart of the method for detecting micro-lesions of pulmonary tuberculosis based on multi-scale convolution in this embodiment; Figure 2 This is a schematic diagram of the overall structure of the improved YOLOv10 algorithm network in this embodiment; Figure 3 This is the overall structure diagram of C2f-Faster-EMA; Figure 4 It is the overall structure diagram of the multi-scale convolution module (MSCM); Figure 5 Comparison data between the improved YOLOv10 and YOLOv10 training results; FIG6 is a comparison diagram of CT prediction results in this embodiment, wherein Figure 6a is the real annotation box in the test dataset; Figure 6b Predict box for YOLOv10s; Figure 6c This is the improved YOLOv10s prediction box. DETAILED DESCRIPTION
[0024] The principle of the present invention is: first, according to the collection of lung CT image data of multiple patients, after preprocessing and doctor annotation, it is used to build a database; based on the network architecture of the improved YOLOv10 algorithm, a neural network optimal optimization model for detecting lung nodule lesions is built, and the preprocessed image is input into the network in the network structure, and after a series of convolutional layers for feature extraction, the feature map is input into the C2f-Faster-EMA module, and the outputs P3, P4, and P5 of different levels in the feature pyramid in the backbone network are input into the corresponding ADown, Conv, and UpSample modules in the multi-scale convolution module, and finally a prediction frame is generated. For the obtained prediction frame, the confidence loss and positioning loss are calculated respectively, and finally, the training is stopped when the set number of iterations is reached or the model can no longer be optimized.
[0025] Embodiment 1: like Figure 1 As shown, a method for detecting micro-lesions of pulmonary tuberculosis based on multi-scale convolution comprises the following steps: S01: Acquire lung CT image data; S02: Feature extraction is performed through the convolution layer of the neural network detection model of lung nodule lesions. After the extracted feature map is input into the Faster and efficient multi-scale attention fusion modules, the outputs of different levels in the feature pyramid in the backbone network are input into the corresponding Adown module, Conv module and UpSample module in the multi-scale convolution module. Multiple feature information is obtained through multi-scale convolution, and finally a prediction box is generated to obtain the detection result.
[0026] In a preferred embodiment, the neural network detection model of pulmonary nodule lesions described in step S02 is improved on the basic framework of YOLO v10, and the C2f module of the backbone network and the neck network in the YOLOv10 model is replaced with a C2f-Faster-EMA module, and the Bottleneck in the C2f module is replaced with a Faster module to obtain a C2f-Faster-EMA module. After the last effective feature extraction layer of the backbone network, feature information is obtained from pooling windows of different sizes through a spatial pyramid pooling module, and a PSA module is added at the end of the backbone network. The features are divided by the PSA module and self-attention is applied to some features. In the backbone network, feature maps of multiple scales are output for use in the multi-scale convolution module of the neck network.
[0027] In a preferred embodiment, the C2f-Faster-EMA module includes a first Conv layer, a Split layer, a first Faster module, a second Faster module, a Concat layer and a second Conv layer connected sequentially, and the Split layer output feature map and the feature maps of the first Faster module and the second Faster module are input into the Concat layer together.
[0028] In a preferred embodiment, each Faster module has a PConv layer, followed by two PWConv layers or Conv layers, and an EMA attention mechanism module is added to form an inverted residual block, in which the middle layer has more channels and a shortcut connection is placed to reuse the input function, and its output is added element by element with the PConv layer output; The EMA attention mechanism module calculates the attention score by the dot product of the query vector Q and the key vector K to obtain the attention weight matrix A. The efficient multi-scale attention mechanism smoothing mathematical expression is applied to the attention weight matrix A as follows: ;
[0029] ;
[0030] in, is the dimension of the key vector K, represents the smoothing coefficient; The output is calculated by smoothing the attention weights, and the smoothed attention weights are weighted summed over the value vector V to obtain the final output representation.
[0031] In a preferred embodiment, the input end of the multi-scale convolution module is divided into three branches, which are respectively connected to the ADown module, the Conv module and the upsample module, and then fused together through a Concat, and then enter a parallel group deep separable convolution. Multiple separable volumes can exist in the parallel group deep separable convolution at the same time. Multiple separable convolutions are added into a feature map, which is added to the parallel group deep separable convolution after passing through the Conv module and the BAM attention mechanism module.
[0032] In a preferred embodiment, the parallel group depthwise separable convolution is expressed as: ;
[0033] in, for The convolution extracts local feature information. To customize the size of the convolution kernel, Indicates that the feature map value is a real number is a three-dimensional tensor whose elements are real numbers and whose dimensions are the number of channels ,high ,width , In the backbone network Layer input feature map, The second scale feature of the branch, is the nth parallel branch convolution, is the number of layers for extracting features from the backbone network, For the mth DSConv, middle There are 4 different depth-separable convolutions, is the size of the convolution kernel, , , Represent the number of channels, height and width respectively; The relationship between each channel is expressed as follows: ;
[0034] in, Represents the output features, and 1×1 convolution is the channel fusion mechanism.
[0035] In a preferred embodiment, the parallel group depth-separable convolution output point multiplication in the multi-scale convolution module is added to the output of the point multiplication. For a given input feature map , BAM attention mechanism infers a 3D attention map , the channel attention branch is calculated on two independent attention branches and the spatial attention branch , and finally calculate the attention map : ;
[0036] In the formula, is the input feature map, The output is an attention-weighted feature map. represents element-by-element multiplication, is a sigmoid function, and the feature maps of the spatial attention branch and the channel attention branch are real numbers. Its elements are real numbers and its dimension is the number of channels. ,high ,width , is the channel attention, for spatial attention; In the spatial attention branch, 1×1 convolution is used to reduce the dimension. The dimensionality reduction formula is: ;
[0037] In the formula, represents the convolution operation, represents the feature map, Represents a batch normalization operation; The channel attention branch performs global average pooling on the feature map F and generates a channel vector , encode the global information of each channel, add a hidden layer of multi-layer perceptron to obtain cross-channel attention, the mathematical expression is: ;
[0038] In the formula, is a multi-layer perceptron, is the average pooling layer, To reduce the dimension, we transform the input from Dimension compression to dimension, To adjust the feature distribution after dimensionality reduction, To restore the dimension from Restore to dimension, It is the benchmark value that affects the final channel weight.
[0039] In a preferred embodiment, the neural network detection model of pulmonary nodule lesions uses the WIOUv3 loss function as the regression loss of the bounding box, includes a dynamic non-monotonic mechanism and gradient gain allocation, and constructs distance attention by distance metric to obtain WIoUv1 with a two-layer attention mechanism. WIoUv1 introduces a dynamic weight mechanism based on the traditional IoU, and automatically adjusts the loss weight according to the scale of the target box or the regression difficulty, so that the model pays more attention to difficult samples or small targets. The mathematical expression is: ;
[0040] In the formula, The basic pre-selected anchor boxes will be significantly optimized , used to measure the matching degree between the predicted box and the real box, is the overlap between the real box and the predicted box, Generally, the IoU is transformed to make it non-negative and the closer it is to the true value, the smaller the loss is, and the prediction accuracy of the bounding box in the target detection model is optimized. and is the width and height of the minimum prediction box, and are the variances of the width and height differences between the predicted box and the true box; A non-monotonic focusing coefficient is constructed using the outlier degree β and applied to WIoUv1 to obtain WIoUv3 with dynamic non-monotonic FM. FM assigns larger gradient gains to these anchor boxes. The dynamic non-monotonic FM is used to determine the gradient gain allocation strategy. The mathematical expression is: , , ,in is the gradient gain, and is a hyperparameter that controls the shape of the gradient gain curve. To adjust the sharpness of the gain, To determine the peak position of the gain, is an exponential move with kinetic energy m.
[0041] In another embodiment, a computer storage medium stores a computer program, and when the computer program is executed, the multi-scale convolution-based pulmonary tuberculosis micro-lesion detection method is implemented. The above detection method is adopted and will not be described in detail here.
[0042] In another embodiment, a pulmonary tuberculosis micro-lesion detection system based on multi-scale convolution includes: An image data acquisition module, which acquires lung CT image data; The pulmonary nodule lesion detection module extracts features through the convolution layer of the neural network detection model of pulmonary nodule lesions. After the extracted feature map is input into the Faster and efficient multi-scale attention fusion modules, the outputs of different levels in the feature pyramid in the backbone network are input into the corresponding Adown module, Conv module and UpSample module in the multi-scale convolution module. Multiple feature information is obtained through multi-scale convolution, and finally a prediction box is generated to obtain the detection result.
[0043] Specifically, the workflow of the pulmonary tuberculosis micro-lesion detection system based on multi-scale convolution is described as follows by taking a preferred embodiment as an example: S01: Collect lung CT image data from multiple patients and create a training data set in the database after preprocessing and doctor annotation; S02: The backbone network optimizes feature extraction, and the feature information will be used for subsequent target detection tasks; S03: Multi-scale convolution module (MSCM), which can obtain multiple feature information through multi-scale convolution and can also capture the details and global information of the image at the same time; S04: Iterate the training data set according to the improved YOLOv10 network model, and finally train the best network model.
[0044] Step S01 collects lung CT image data of multiple patients, and uses them for database construction after preprocessing and doctor annotation, including: S11: Use data enhancement on the collected lung CT images to increase the number of images, and then combine with doctors to annotate the data.
[0045] S21: performing data enhancement processing on the annotated micro-lesions, wherein the data enhancement method comprises: translation, flipping, contrast transformation and random forms of noise disturbance.
[0046] The backbone network uses C2f-Faster-EMA to improve the performance of feature extraction of micro-tuberculosis lesions. A multi-scale convolution module is added to the feature fusion network to process features of different scales, capture contextual feature information across multiple scales, obtain multiple feature information, and increase the ability of information interaction.
[0047] See also Figure 2 As shown in the figure, based on the network architecture of the improved YOLOv10 algorithm, a neural network optimal optimization model for detecting lung nodule lesions is built. The improved algorithm is based on the basic framework of YOLO. The C2f-Faster-EMA module is used for feature extraction in the improved backbone network. After the last effective feature extraction layer, a spatial pyramid pooling (SPPF) module is added to obtain feature information from pooling windows of different sizes. The PSA module unique to YOLOv10 is added at the end of the backbone network to divide features through PSA and apply self-attention to some features. Combined with effective self-attention, the computational complexity and memory usage are reduced, while the global representation learning is enhanced. In the backbone network, feature maps of three scales of 64×64, 32×32 and 16×16 are output for the subsequent feature fusion network.
[0048] See also Figure 3 The one shown is based on the improved Y0L0v10 model, and the improvement lies in that in the backbone network and the neck network, the C2f module in the Y0L0v10 model is replaced by the C2f-faster-EMA module, wherein the C2f-Faster-EMA module is connected in order of the first Conv layer, the Split layer, the first Faster module, the second Faster module, the Concat layer and the second Conv layer, and the output feature map of the Split layer and the feature maps of the first Faster module and the second Faster module are input into the Concat layer together.
[0049] The Bottleneck in C2f is replaced by the Faster module. Each Faster module has a PConv layer, followed by two PWConv (or Conv 1×1) layers, and an EMA attention mechanism. Together, they form an inverted residual block, where the middle layer has more channels and a shortcut connection is placed to reuse the input function. The EMA attention mechanism is a weighted average attention mechanism. The more recent the data, the greater the weight, and the weight of historical data decays exponentially. The formula is as follows: ;
[0050] Represents the input value at the current moment (such as attention weight or gradient). Represents the exponential moving average at the current moment. Represents the smoothing coefficient (0<α<10<α<1), which controls the weight distribution between the current value and the historical value. The larger it is, the more the model focuses on the current value; The smaller it is, the more the model relies on historical values.
[0051] The EMA attention mechanism calculates the attention score by the dot product of Query and Key to obtain the attention weight matrix A. The mathematical expression of applying the efficient multi-scale attention mechanism to the attention weight matrix A is as follows: ;
[0052] ;
[0053] The output is calculated by smoothing the attention weights, and the smoothed attention weights are weighted summed on the Value to get the final output representation.
[0054] See also Figure 3 As shown in the figure, the backbone network feature extraction uses the C2f-Faster-EMA module. FasterBlock can not only extract spatial features more effectively, but also enhance the network's sensitivity to the central area, thereby improving the ability to capture local details. Efficient multi-scale attention can effectively aggregate cross-dimensional interactive information and achieve refined processing of high-level feature maps.
[0055] See also Figure 4As shown in the figure, the multi-scale convolution module (MSCM) includes three branches at the input end, which are connected to an ADown, a Conv and an upsample respectively, and then fused together through a Concat to enter a parallel group depthwise separable convolution. Multiple separable volumes (such as 3×3, 5×5, 7×7 and no convolution) can exist in the parallel group depthwise separable convolution at the same time. Multiple separable convolutions are added into a feature map, which is then added to the parallel group depthwise separable convolution through a Conv and BAM attention mechanism.
[0056] The multi-scale convolutional module (MSCM) is an Inception-style module; at the input end, the output of the backbone network is analyzed using ADown, Conv, and upsample respectively; and then through the parallel group depthwise separable convolution (DSConv), it can be expressed by mathematical expression: ;
[0057] in for The convolution extracts local feature information. For the mth DSConv is set as and .
[0058] The mathematical expression of the relationship between each channel is: ;
[0059] in Represents the output features, and 1×1 convolution is the channel fusion mechanism.
[0060] DSConv is a lightweight structure in the multi-scale convolution module. It is mainly divided into two steps: separation convolution and point-by-point convolution. They perform addition operations. The mathematical expressions comparing the parameters and calculations of DSConv and Conv are as follows: ;
[0061] in Expressed as the size of the convolution kernel, the size of the output feature map is , the input channel is , the output channel is .
[0062] The parallel DSConv outputs in the multi-scale convolutional module are multiplied, and a two-layer BAM attention mechanism is added to the output of the dot product to reduce the effect of denoising low-level features. Multi-layer BAM can gradually focus on smaller targets. BAM for a given input feature map , BAM infers a 3D attention map , the channel attention branch is calculated on two independent attention branches and the spatial attention branch Finally, calculate the attention map It can be expressed mathematically as: ;
[0063] In the formula, represents element-by-element multiplication, is the sigmoid function, and both independent attentions are .
[0064] In the spatial attention branch, 1×1 convolution is used to reduce the dimension to integrate and compress the feature maps across channel dimensions. The mathematical expression of dimensionality reduction is: ;
[0065] In the formula, Represents the convolution operation represents the feature map, Represents a batch normalization operation.
[0066] The channel attention branch aggregates the feature map F of each channel, globally averages the feature map F, and generates a channel vector . Encode the global information of each channel. To obtain cross-channel attention, a multi-layer perceptron (MLP) with a hidden layer is added, and the activation size is set to , where r is the reduction ratio. The mathematical expression is: ;
[0067] In the formula, , , , .
[0068] The WIOUv3 loss function used as the regression loss of the bounding box in the overall neural network framework contains a dynamic non-monotonic mechanism and designs a reasonable gradient gain distribution. This strategy reduces the large gradients or harmful gradients that appear in extreme samples. The distance metric constructs the distance attention and obtains the WIoUv1 with a two-layer attention mechanism. The mathematical expression is: ;
[0069] In the formula, Used to measure the matching degree between the predicted box and the real box. and is the width and height of the minimum prediction box, and are the variances of the width and height differences between the predicted box and the true box, respectively.
[0070] A non-monotonic focusing coefficient is constructed using β and applied to WIoU v1 to obtain WIoU v3 with dynamic non-monotonic FM. WIoU v3 achieves superior performance by using a judicious gradient gain allocation strategy with dynamic non-monotonic FM, as shown in the mathematical expression: , , is the gradient gain, represents the degree of outlier, , It is an exponential movement with kinetic energy m and the ability to self-regulate.
[0071] See also Figure 5 As shown in the figure, a comparative experiment was conducted using the experimental data of the improved YOLOv10 and the experimental data of the existing YOLOv10. In the experiment, the mAP% of YOLOv10 on the test set was 86.7%, and the improved YOLOv10 showed significant performance on the test set, with mAP% reaching 89.9%. Compared with the original YOLOv10 network, the improved YOLO has significantly improved overall performance, taking into account the recall rate and average detection accuracy of lesions, and also significantly improved the loss function in the training and verification pre-selection boxes.
[0072] The improved YOLOv10 has been improved in different scenarios and its generalization ability has been improved under the effect of multi-scale convolution modules and C2f-Faster-EMA. Figure 6a-6c Comparison chart of CT prediction results.
[0073] The C2f-Faster-EMA module improves the feature extraction performance of micro-tuberculosis lesions. A multi-scale convolution module is added to the feature fusion network to process features of different scales, capture contextual feature information across multiple scales, obtain multiple feature information, and increase the ability of information interaction. The overall network model uses the WIoUv3 loss function to not only reflect the degree of overlap between the predicted box and the true box, but also reflects that large and small objects do not affect each other under occlusion. This method improves the average accuracy of micro-lesion target detection and has a more effective recall rate. The network model assists doctors in diagnosing micro-lesions and accurately determines the location information of lesions in CT images to prepare for subsequent basic diagnosis and treatment.
[0074] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principles of the present invention shall be equivalent replacement modes and shall be included in the protection scope of the present invention.
Claims
1. A method for detecting micro-lesions of pulmonary tuberculosis based on multi-scale convolution, characterized in that: The following steps are involved: S01: Acquire lung CT image data; S02: Feature extraction is performed through the convolution layer of the neural network detection model of lung nodule lesions. After the extracted feature map is input into the Faster and efficient multi-scale attention fusion modules, the outputs of different levels in the feature pyramid in the backbone network are input into the corresponding Adown module, Conv module and UpSample module in the multi-scale convolution module. Multiple feature information is obtained through multi-scale convolution, and finally a prediction box is generated to obtain the detection result.
2. The method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution according to claim 1, characterized in that: The neural network detection model for pulmonary nodule lesions described in step S02 is improved on the basic framework of YOLO v10. The C2f module of the backbone network and the neck network in the YOLOv10 model is replaced with a C2f-Faster-EMA module, and the Bottleneck in the C2f module is replaced with a Faster module to obtain a C2f-Faster-EMA module. After the last effective feature extraction layer of the backbone network, feature information is obtained from pooling windows of different sizes through a spatial pyramid pooling module. A PSA module is added at the end of the backbone network. Features are divided through the PSA module and self-attention is applied to some features. Feature maps of multiple scales are output in the backbone network for use in a multi-scale convolution module of the neck network.
3. The method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution according to claim 2, characterized in that: The C2f-Faster-EMA module includes a first Conv layer, a Split layer, a first Faster module, a second Faster module, a Concat layer and a second Conv layer connected in sequence, and the Split layer output feature map and the feature maps of the first Faster module and the second Faster module are input into the Concat layer together.
4. The method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution according to claim 3, characterized in that: Each Faster module has a PConv layer followed by two PWConv layers or Conv layers, plus an EMA attention mechanism module, forming an inverted residual block, where the middle layer has more channels and a shortcut connection is placed to reuse the input features, and its output is added element-wise with the PConv layer output; The EMA attention mechanism module calculates the attention score by the dot product of the query vector Q and the key vector K to obtain the attention weight matrix A. The efficient multi-scale attention mechanism smoothing mathematical expression is applied to the attention weight matrix A as follows: , , in, is the dimension of the key vector K, represents the smoothing coefficient; The output is calculated by smoothing the attention weights, and the smoothed attention weights are weighted summed over the value vector V to obtain the final output representation.
5. The method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution according to claim 1, characterized in that: The input end of the multi-scale convolution module is divided into three branches, which are respectively connected to the ADown module, the Conv module and the upsample module. After being fused together through a Concat, they enter a parallel group deep separable convolution. Multiple separable volumes can exist in the parallel group deep separable convolution at the same time. Multiple separable convolutions are added into a feature map, which is added to the parallel group deep separable convolution after passing through the Conv module and the BAM attention mechanism module.
6. The method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution according to claim 5, characterized in that: The parallel group depthwise separable convolution is expressed as: , in, for The convolution extracts local feature information. To customize the size of the convolution kernel, Indicates that the feature map value is a real number is a three-dimensional tensor whose elements are real numbers and whose dimensions are the number of channels ,high ,width , In the backbone network Layer input feature map, The second scale feature of the parallel branch convolution, is the number of layers for extracting features from the backbone network, For the mth DSConv, middle For 4 different depth-separable convolutions, is the size of the convolution kernel; The relationship between each channel is expressed as follows: , in, Represents the output features, and 1×1 convolution is the channel fusion mechanism.
7. The method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution according to claim 5, characterized in that: The parallel group depth-separable convolution output point multiplication in the multi-scale convolution module is added to the output of the point multiplication. For a given input feature map , BAM attention mechanism infers a 3D attention map , the channel attention branch is calculated on two independent attention branches and the spatial attention branch , and finally calculate the attention map : , In the formula, is the input feature map, The output is an attention-weighted feature map. represents element-by-element multiplication, is a sigmoid function, and the feature maps of the spatial attention branch and the channel attention branch are real numbers. , whose elements are real numbers and whose dimension is the number of channels ,high ,width ; In the spatial attention branch, 1×1 convolution is used to reduce the dimension. The dimensionality reduction formula is: , In the formula, represents the convolution operation, represents the feature map, Represents a batch normalization operation; The channel attention branch performs global average pooling on the feature map F and generates a channel vector , encode the global information of each channel, add a hidden layer of multi-layer perceptron to obtain cross-channel attention, the mathematical expression is: , In the formula, is a multi-layer perceptron, is the average pooling layer, To reduce the dimension, we transform the input from Dimension compression to dimension, To adjust the feature distribution after dimensionality reduction, To restore the dimension from Restore to dimension, It is the benchmark value that affects the final channel weight.
8. The method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution according to claim 1, characterized in that: The neural network detection model of pulmonary nodule lesions uses the WIOUv3 loss function as the regression loss of the bounding box, includes a dynamic non-monotonic mechanism and gradient gain allocation, and constructs distance attention by distance metric to obtain the WIoUv1 loss function with a two-layer attention mechanism. The mathematical expression is: , In the formula, Used to measure the matching degree between the predicted box and the real box. is the overlap between the real box and the predicted box, and is the width and height of the minimum prediction box, and are the variances of the width and height differences between the predicted box and the true box; A non-monotonic focusing coefficient is constructed using the outlier degree β and applied to WIoUv1 to obtain WIoUv3 with dynamic non-monotonic allocation of gradient gain. The dynamic non-monotonic allocation of gradient gain is used to determine the gradient gain allocation strategy. The mathematical expression is: , , ,in is the gradient gain, and is a hyperparameter, is an exponential move with kinetic energy m.
9. A pulmonary tuberculosis micro-lesion detection system based on multi-scale convolution, characterized in that: include: An image data acquisition module, which acquires lung CT image data; The pulmonary nodule lesion detection module extracts features through the convolution layer of the neural network detection model of pulmonary nodule lesions. After the extracted feature map is input into the Faster and efficient multi-scale attention fusion modules, the outputs of different levels in the feature pyramid in the backbone network are input into the corresponding Adown module, Conv module and UpSample module in the multi-scale convolution module. Multiple feature information is obtained through multi-scale convolution, and finally a prediction box is generated to obtain the detection result.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the method for detecting pulmonary tuberculosis micro-lesions based on multi-scale convolution as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Power field operation specification detection method based on improved YOLOv8 algorithm
CN118736307A
Methods of assessing lung disease in chest x-rays
US20220180514A1
Cited By
Multi-scale adaptive lesion detection method based on breast ultrasound
CN120374631A
A multi-scale adaptive lesion detection method based on breast ultrasound
CN120374631B
Pulmonary tuberculosis focus detection method and system based on shared feature map, and storage medium
CN121414755A
Tuberculosis lesion detection method and system based on shared feature map and storage medium
CN121414755B
Improved YOLOv8-based low-altitude citrus pest and disease damage identification method, equipment and medium
CN121545151A