Mobile pulmonary nodule detection and multi-class classification method based on deep learning

By applying deep learning-based mobile lung nodule detection and multi-level classification methods in medical image processing, combining morphological processing and multi-head self-attention mechanism, the problems of complex, time-consuming and difficult to process non-standardized images in the prior art are solved, and efficient and accurate lung nodule detection and multi-level classification are achieved.

CN119273695BActive Publication Date: 2025-05-09EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411816040.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-09
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

The prior art is complex and time-consuming in medical image processing, is susceptible to noise, difficult to process non-standardized images, and traditional models are difficult to detect lung nodules of different sizes at the same time.

Method used

A mobile lung nodule detection and multi-level classification method based on deep learning is proposed, and lung parenchymal segmentation is performed through morphological processing, combined with improved lightweight network design and multi-head self-attention mechanism to realize automatic size standardization and multi-scale feature extraction.

Benefits of technology

It realizes efficient detection of non-standardized medical images on the mobile terminal, improves detection accuracy and efficiency, meets doctors' needs for richer diagnostic content, and provides clearer and easier-to-understand diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119273695B_ABST
    Figure CN119273695B_ABST
Patent Text Reader

Abstract

The present invention proposes a mobile terminal lung nodule detection and multi-level classification method based on deep learning, the method comprising: converting an image into a grayscale image, obtaining a segmentation threshold based on the grayscale image, and obtaining a binary image through the segmentation threshold; obtaining an expanded mask image through the binary image; obtaining a mask based on the binary image, and obtaining a lung parenchyma image processed by the lung parenchyma segmentation part based on the mask and the expanded mask image; based on the lung parenchyma image processed by the lung parenchyma segmentation part, obtaining an enhanced feature map processed by a multi-head self-attention mechanism, a connected output feature map, and a feature map after a fast spatial pyramid pooling operation; the three feature maps are input into an efficient detection head to obtain results and predict and track the target. The present invention takes multi-scale feature extraction and fusion as the core for lung nodules, and improves detection accuracy and efficiency by fully integrating and processing the features of targets of different sizes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and in particular to a mobile terminal lung nodule detection and multi-level classification method based on deep learning. Background Art

[0002] The traditional machine learning methods based on CAD technology that are commonly used in medical image processing are not only computationally complex and time-consuming, but also easily affected by noise; traditional models can only process standard CT images, and the non-standardized images provided by patients make it difficult to identify lung areas and lesions, affecting the final diagnostic accuracy; as deep learning models become more and more complex, their training difficulty is greatly increased, which also places high demands on computing resources, and the model is difficult to implement; the sizes of lung nodules of different malignant degrees vary by dozens of times, and traditional diagnostic networks are difficult to simultaneously meet the detection of lung nodules of different sizes. Summary of the invention

[0003] In view of the above situation, the main purpose of the present invention is to propose a mobile lung nodule detection and multi-level classification method based on deep learning to solve the above technical problems.

[0004] The present invention proposes a mobile terminal lung nodule detection and multi-level classification method based on deep learning, which comprises the following steps:

[0005] Step 1: Convert the RGB image to a grayscale image, and then perform data normalization on the grayscale image to obtain a standardized grayscale image; select the middle area from the standardized grayscale image, calculate the mean of the middle area, divide the pixel values ​​in the mean of the middle area into two categories by the K-means clustering algorithm (K-maeans clustering algorithm), and obtain the segmentation threshold based on the center points of the two categories of pixel values; convert the standardized grayscale image into a binary image based on the segmentation threshold;

[0006] Step 2: performing an erosion operation on the binary image to obtain an eroded binary image; performing an expansion operation on the eroded binary image to obtain an expanded binary image;

[0007] Use the connected region marking function (medasure.label) to mark the connected regions in the dilated binary image to obtain the marked regions, and then use the unique value function (np.unique) to capture all the marked regions to obtain the captured regions;

[0008] Traverse all captured areas, filter out qualified areas according to the bounding box size and position, and add the labels corresponding to the qualified areas to the list; traverse all labels in the list, use the mask image of all zeros to iterate to generate the lung mask image, then convert the lung mask image to uint8 type and perform the expansion operation again to obtain the expanded lung mask image;

[0009] Step 3, assign and mark the pixels in the binary image to obtain the marked pixels; calculate and obtain the pixel sum of each connected area in the binary image, and take the maximum value based on the pixel sum of all connected areas to obtain the pixel sum of the largest connected area; filter the pixel sum of the largest connected area to obtain a connected area index list that is greater than half of the pixel sum of the largest connected area; generate a mask of the largest connected area based on the marked pixels and the connected area index list;

[0010] A lung parenchyma image processed by lung parenchyma segmentation is obtained based on the expanded lung mask image and the mask of the maximum connected area;

[0011] Collect the lung parenchyma images processed by the lung parenchyma segmentation part and the original image labels to generate a training data set, and use the training data set to train the model;

[0012] Step 4, standardizing the lung parenchyma image after the lung parenchyma segmentation process to obtain a standardized image;

[0013] The standardized image is input into the improved residual shift module (iRMB module), and after residual connection and deep convolution operations, an enhanced feature map is obtained; the enhanced feature map is processed by the multi-head self-attention mechanism to obtain an enhanced feature map processed by the multi-head self-attention mechanism;

[0014] The standardized image is processed using the group convolution module (GSConv module) and the multi-scale group convolution pooling module (MsGSCP module), and the processing results are skip-connected through the efficient bottleneck module to obtain the connected output feature map;

[0015] The standardized image is processed using the fast spatial pyramid pooling module (SPPF module) to obtain a feature map after the fast spatial pyramid pooling operation;

[0016] Step 5: Input the enhanced feature map processed by the multi-head self-attention mechanism, the concatenated output feature map, and the feature map after the fast spatial pyramid pooling operation into the efficient detection head, and obtain the bounding box coordinates through multi-scale convolution and DFL mechanism;

[0017] Objects are predicted and tracked via bounding box coordinates.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] 1. The present invention constructs a lung parenchyma segmentation module based on morphological processing in the network, and performs automatic size standardization when inputting into the detection network, which can more effectively detect the non-standard data provided by patients and break the time and space boundaries of medical treatment;

[0020] 2. The present invention reduces the amount of calculation and improves the detection effect through improved lightweight network design and optimized loss function, ensuring that the lightweight network can smoothly implement functions on the mobile terminal;

[0021] 3. The present invention adopts a five-level classification standard for the malignancy of pulmonary nodules based on the LU-RADS classification method, which enhances interpretability and meets the needs of doctors for richer and more interpretable diagnostic content, such as lesion localization and automatic generation of diagnostic reports, and meets the needs of patients for clearer and easier-to-understand diagnostic results;

[0022] 4. The present invention focuses on multi-scale feature extraction and fusion of lung nodules, and improves detection accuracy and efficiency by fully integrating and processing the features of targets of different sizes.

[0023] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flow chart of the mobile terminal lung nodule detection and multi-level classification method based on deep learning proposed by the present invention;

[0025] Figure 2 A schematic diagram of a module for lung parenchyma segmentation and lung nodule detection and classification of a mobile terminal lung nodule detection and multi-level classification method based on deep learning proposed by the present invention;

[0026] Figure 3 This is a flow chart of the morphological segmentation module of the deep learning-based mobile lung nodule detection and multi-level classification method proposed in the present invention;

[0027] Figure 4 A data set production flow chart for the deep learning-based mobile lung nodule detection and multi-level classification method proposed in the present invention;

[0028] Figure 5 This is a schematic diagram of the network structure of the lung nodule classification detection module of the mobile terminal lung nodule detection and multi-level classification method based on deep learning proposed by the present invention;

[0029] Figure 6 This is the iRMB module structure diagram of the mobile terminal lung nodule detection and multi-level classification method based on deep learning proposed by the present invention;

[0030] Figure 7 A schematic diagram of the recall rate-confidence curve of the mobile terminal lung nodule detection and multi-level classification method based on deep learning proposed in the present invention;

[0031] Figure 8Schematic diagram of the recall-confidence curve of YOLOv10;

[0032] Fig. 9 A schematic diagram of the precision-recall curve of the deep learning-based mobile lung nodule detection and multi-level classification method proposed in the present invention;

[0033] Fig.10 Schematic diagram of the precision-recall curve of YOLOv10. DETAILED DESCRIPTION

[0034] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0035] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0036] See also Figure 1 , Figure 2 and Figure 3 The embodiment of the present invention proposes a mobile terminal lung nodule detection and multi-level classification method based on deep learning, which includes the following steps:

[0037] Step 1: Convert the RGB image to a grayscale image, and then perform data normalization on the grayscale image to obtain a standardized grayscale image; select the middle area from the standardized grayscale image, calculate the mean of the middle area, divide the pixel values ​​in the mean of the middle area into two categories by the K-means clustering algorithm, and obtain the segmentation threshold based on the center points of the two categories of pixel values; convert the standardized grayscale image into a binary image based on the segmentation threshold;

[0038] In step 1, the RGB image is converted into a grayscale image, and then the grayscale image is subjected to data normalization to obtain a standardized grayscale image; the middle area is selected from the standardized grayscale image, the mean of the middle area is obtained by calculation, and the pixel values ​​in the mean of the middle area are divided into two categories by the K-means clustering algorithm, and the segmentation threshold is obtained based on the center points of the two categories of pixel values; the standardized grayscale image is converted into a binary image based on the segmentation threshold. The specific steps are as follows:

[0039] The relationship between the conversion of RGB images to grayscale images is:

[0040] ;

[0041] in, Represents the grayscale value after conversion, Represents the intensity value of the blue channel, Represents the intensity value of the red channel, Represents the intensity value of the green channel;

[0042] The grayscale image is subjected to data standardization, and the corresponding relationship in the process of obtaining the standardized grayscale image is:

[0043] ;

[0044] in, represents the mean value of the grayscale image, represents the standard deviation of the grayscale image, represents a grayscale image, Represents the standardized grayscale image;

[0045] The middle area is selected from the standardized grayscale image, the mean of the middle area is obtained by calculation, and the pixel values ​​in the mean of the middle area are divided into two categories by the K-means clustering algorithm. The segmentation threshold is obtained based on the center points of the two categories of pixel values. The relationship between the corresponding process is:

[0046] ;

[0047] in, and Represents the number of two types of pixel values, and Represent two types of pixel values, represents the segmentation threshold, and Represent two cluster centers respectively;

[0048] Based on the segmentation threshold, the standardized grayscale image is converted into a binary image. The corresponding relationship is:

[0049] ;

[0050] in, Represents a binary image; Represents processing by a function, where this function returns one of two values ​​based on a given condition. If the condition is true, it returns the first value (1.0), otherwise it returns the second value (0.0).

[0051] Furthermore, the purpose of converting RGB images to grayscale images is designed based on the different sensitivities of the human eye to different colors. The human eye is most sensitive to green light and least sensitive to blue light, so the green channel has the largest weight and the blue channel has the smallest weight.

[0052] Step 2: performing an erosion operation on the binary image to obtain an eroded binary image; performing an expansion operation on the eroded binary image to obtain an expanded binary image;

[0053] The connected regions in the dilated binary image are marked using the connected region marking function to obtain the marked regions, and then all the marked regions are captured using the unique value finding function to obtain the captured regions;

[0054] Traverse all captured areas, filter out qualified areas according to the bounding box size and position, and add the labels corresponding to the qualified areas to the list; traverse all labels in the list, use the mask image of all zeros to iterate to generate the lung mask image, then convert the lung mask image to uint8 type and perform the expansion operation again to obtain the expanded lung mask image;

[0055] In step 2, the binary image is corroded to obtain a corroded binary image; the corroded binary image is expanded to obtain an expanded binary image. The specific steps are as follows:

[0056] Perform an erosion operation on the binary image to obtain the eroded binary image:

[0057] ;

[0058] in, Represents a structural element, represents a translated version of the structural element, represents the corrosion operation, Represents the pixel position in a binary image;

[0059] The binary image after corrosion is expanded to obtain the expanded binary image. The relationship between the corresponding process is:

[0060] ;

[0061] in, represents the expansion operation, represents the empty set, Represents the pixel position in the binary image after corrosion;

[0062] Traverse all captured areas, filter out qualified areas according to the bounding box size and position, and add the labels corresponding to the qualified areas to the list; traverse all labels in the list, use the mask image of all zeros to iterate to generate the lung mask image, then convert the lung mask image to uint8 type and expand it again to obtain the expanded lung mask image. The specific steps are as follows:

[0063] Traverse all captured areas, filter out qualified areas according to the size and position of the bounding box, and add the labels corresponding to the qualified areas to the list. The relationship in the corresponding process is:

[0064] ;

[0065] in, represents the bounding box of the region, Indicates the filter conditions. Represents the coordinate values ​​of the bounding box, Represents the bounding box in the region attribute;

[0066] Traverse all the labels in the list, and use the mask image of all zeros to iteratively generate the lung mask image. The corresponding relationship in the process is:

[0067] ;

[0068] in, represents the lung mask image, Indicates the label, Indicates the filtered tags.

[0069] Step 3, assign and mark the pixels in the binary image to obtain the marked pixels; calculate and obtain the pixel sum of each connected area in the binary image, and take the maximum value based on the pixel sum of all connected areas to obtain the pixel sum of the largest connected area; filter the pixel sum of the largest connected area to obtain a connected area index list that is greater than half of the pixel sum of the largest connected area; generate a mask of the largest connected area based on the marked pixels and the connected area index list;

[0070] A lung parenchyma image processed by lung parenchyma segmentation is obtained based on the expanded lung mask image and the mask of the maximum connected area;

[0071] Collect the lung parenchyma images processed by the lung parenchyma segmentation part and the original image labels to generate a training data set, and use the training data set to train the model;

[0072] Please refer to Figure 4In step 3, the pixels in the binary image are allocated and marked to obtain the marked pixels; the pixel sum of each connected area in the binary image is calculated and obtained, and the maximum value is taken based on the pixel sum of all connected areas to obtain the pixel sum of the largest connected area; the pixel sum of the largest connected area is screened to obtain a connected area index list that is greater than half of the pixel sum of the largest connected area; based on the marked pixels and the connected area index list, a mask of the largest connected area is generated. The specific steps are as follows:

[0073] The pixel points in the binary image are allocated and marked to obtain the marked pixel points. The corresponding relationship in the process is:

[0074] ;

[0075] in, Indicates the midpoint of the marked image The label value of represents the labeled image, Indicates the current label value;

[0076] Calculate and obtain the pixel sum of each connected area in the binary image. The corresponding process has the following relationship:

[0077] ;

[0078] in, Indicates The sum of pixels in the connected region is Indicates A set of pixels in a connected region, Represents the midpoint of a binary image The pixel value of

[0079] Based on the pixel sum of all connected areas, the maximum value is taken to obtain the pixel sum of the maximum connected area. The corresponding relationship is:

[0080] ;

[0081] in, represents the total number of connected regions, represents the pixel sum of the largest connected region, Indicates The pixel values ​​of the connected regions;

[0082] The pixel sum of the maximum connected area is screened to obtain the connected area index list that is greater than half of the pixel sum of the maximum connected area. The relationship between the corresponding process is:

[0083] ;

[0084] in, represents the pixel sum of the connected region, A list of connected region indices that represent pixels that are larger than half of the maximum connected region;

[0085] Based on the marked pixels and the connected area index list, the mask of the maximum connected area is generated. The relationship between the corresponding process is:

[0086] ;

[0087] in, Represents the midpoint of the maximum connected region The value of .

[0088] Step 4, standardizing the lung parenchyma image after the lung parenchyma segmentation process to obtain a standardized image;

[0089] The standardized image is input into the improved residual shift module, and after residual connection and deep convolution operations, an enhanced feature map is obtained; the enhanced feature map is processed by the multi-head self-attention mechanism to obtain an enhanced feature map processed by the multi-head self-attention mechanism;

[0090] The standardized image is processed using the group convolution module and the multi-scale group convolution pooling module, and the processing results are skip-connected through the efficient bottleneck module to obtain the connected output feature map;

[0091] The standardized image is processed using the fast spatial pyramid pooling module to obtain a feature map after the fast spatial pyramid pooling operation;

[0092] Please refer to Figure 5 and Figure 6 In step 4, the lung parenchyma image after the lung parenchyma segmentation is standardized to obtain a standardized image. The relationship between the corresponding process is:

[0093] ;

[0094] in, represents the normalized image, represents the lung parenchyma image after the lung parenchyma segmentation part is processed. represents the standard deviation of the lung parenchyma image after the lung parenchyma segmentation part is processed. represents the mean value of the lung parenchyma image after being processed by the lung parenchyma segmentation part;

[0095] The standardized image is input into the improved residual shift module, and after residual connection and deep convolution operations, an enhanced feature map is obtained. The enhanced feature map is processed by the multi-head self-attention mechanism to obtain an enhanced feature map processed by the multi-head self-attention mechanism, where the expression of the residual connection is:

[0096] ;

[0097] in, represents the convolution operation, represents the normalized image, represents the weight of the convolutional sum, represents the output feature map;

[0098] The depth convolution operation includes depth-wise separable convolution operation and point-wise convolution operation. The expression of depth-wise separable convolution is:

[0099] ;

[0100] in, represents the channels of the normalized image, represents the channel of the convolution kernel, Represents the elements of the output feature map, Represents the row index, Represents the column index, represents the column index of the convolution kernel, represents the row index of the convolution kernel, represents the height of the convolution kernel, Indicates the width of the convolution kernel;

[0101] The expression of point-by-point convolution is:

[0102] ;

[0103] in, Represents the output feature map The pixel value at represents the channel index of the normalized image, Represents the total number of channels of the normalized image, Indicates the image after normalization The pixel value at Indicates that the convolution kernel of size 1×1 is in the channel The weight on represents the row index of the output feature map, Represents the column index of the output feature map;

[0104] The expression of multi-head self-attention is:

[0105] ;

[0106] in, represents the query matrix, represents the key matrix, represents the value matrix, represents the transpose operation, represents the scaling factor;

[0107] The standardized image is processed using the group convolution module and the multi-scale group convolution pooling module, and the processing results are skipped through the efficient bottleneck module to obtain the connected output feature map. The group convolution module includes the group convolution operation and the channel shuffle operation. The relationship corresponding to the group convolution operation is:

[0108] ;

[0109] in, Indicates the number of groups, Indicates The input feature map of the group, represents the convolution kernel, and Both represent the relative displacement offset of the convolution kernel;

[0110] The specific operations for channel shuffle are as follows:

[0111] Use the reshape operation to rearrange the channels of the standardized image;

[0112] Perform shuffle operation on the rearranged channels to disrupt the order of each group of channels;

[0113] The image after the shuffle operation is restored and output through reshape. The relationship between the corresponding process is:

[0114] ;

[0115] in, Represents a shuffle operation;

[0116] The multi-scale group convolution pooling module includes a lightweight convolution operation and a SiLU activation function, where the relationship between the lightweight convolution operation is:

[0117] ;

[0118] in, represents the concatenated feature map, All represent standardized images;

[0119] The expression of SiLU activation function is:

[0120] ;

[0121] in, Represents the feature map after the convolution operation, Represents the Sigmoid function;

[0122] The working steps of the efficient bottleneck module are as follows:

[0123] Reduce the number of channels of the input feature map through a 1×1 convolution layer;

[0124] The output feature map of the channel input is then reduced through a 3×3 convolution layer to maintain the integrity of the spatial features;

[0125] Then restore the number of channels to the original value through a 1×1 convolution layer;

[0126] Finally, the result is output through skip connection processing. The relationship between the corresponding process is:

[0127] ;

[0128] in, Represents the feature map of the network layer output, represents the output feature map after residual connection, Represents the input feature map;

[0129] The working steps of the fast spatial pyramid pooling module are as follows:

[0130] Perform a maximum pooling operation on the standardized feature map to generate multiple feature maps of different sizes;

[0131] The feature maps of different scales are concatenated with the standardized feature map to obtain a feature map containing multi-scale information;

[0132] The feature map containing multi-scale information is processed by standard convolution operation to obtain the feature map after fast spatial pyramid pooling operation. The relationship between the corresponding process is:

[0133] ;

[0134] in, Represents the feature map after fast spatial pyramid pooling operation, , and They represent the maximum pooling operations of 5×5, 9×9, and 13×13 respectively.

[0135] Furthermore, the EMO module proposed in the present invention includes a residual movement module and a multi-head self-attention processing mechanism.

[0136] Furthermore, this step can capture feature information at different levels by using dilated convolutions with different dilation rates (e.g., 1×1, 3×3, and 5×5), so that the network has good detection capabilities for lung nodules of various sizes;

[0137] In this step, multiple convolution results are spliced ​​together to form a multi-dimensional feature map for use by subsequent modules;

[0138] This step also introduces the Squeeze-and-Excitation mechanism in the iRMB module. Through the Squeeze-and-Excitation mechanism, the weights of feature channels can be adaptively recalibrated to emphasize important features.

[0139] In this step, during the point-by-point convolution process, linear combinations between channels are achieved through 1×1 convolution.

[0140] Step 5: Input the enhanced feature map processed by the multi-head self-attention mechanism, the concatenated output feature map, and the feature map after the fast spatial pyramid pooling operation into the efficient detection head, and obtain the bounding box coordinates through multi-scale convolution and DFL mechanism;

[0141] Predict and track objects through bounding box coordinates;

[0142] In step 5, the enhanced feature map processed by the multi-head self-attention mechanism, the concatenated output feature map, and the feature map after the fast spatial pyramid pooling operation are input into the efficient detection head, and the bounding box coordinates are obtained through multi-scale convolution and DFL mechanism. The calculation relationship of multi-scale convolution is:

[0143] ;

[0144] in, Indicates A collection of convolution kernels, Represents the total number of sets of different convolution kernels, and They represent the offset under different expansion rates respectively;

[0145] The bounding box coordinates are obtained through the DFL mechanism, and the corresponding relationship is:

[0146] ;

[0147] in, represents the adjusted bounding box coordinates, represents the prior box, represents the scaling factor, Represents the bounding box adjustment operation, Represents the predicted values ​​of the bounding box parameters.

[0148] Furthermore, in the classification layer, the error between the model output and the true label is measured by the cross entropy loss. The expression of the cross entropy loss function is:

[0149] ;

[0150] in, represents the cross entropy loss function, represents the true label, represents the class probability predicted by the model, represents the total number of categories, Represents the category index.

[0151] Please refer to Figure 7 and Figure 8 Furthermore, the performance of the present invention on the recall-confidence curve is better than that of YOLOv10, showing that it can maintain a high recall rate at various confidence levels, which means that no matter how the user's confidence requirements for the detection results are, this method can provide more real positive samples. In addition, it can still maintain a high recall rate in the high confidence interval, indicating high robustness to outliers and noise, as well as stable working ability in complex environments. This flexibility allows users to adjust the confidence threshold according to actual needs to adapt to different application scenarios.

[0152] Please refer to Fig. 9 and Fig.10 Furthermore, the performance of the present invention on the precision-recall curve is better than that of YOLOv10, showing that while maintaining a high recall rate, the precision is also significantly improved, achieving a good balance between precision and recall. This balance is crucial to reducing false positives and improving user satisfaction, especially in situations where high accuracy is required. In addition, the overall trend of the present method on the PR curve is closer to the upper right corner, indicating its better generalization ability and ability to maintain high precision at different recall levels.

[0153] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0154] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0155] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A mobile lung nodule detection and multi-level classification method based on deep learning, characterized in that: The method comprises the following steps: Step 1: Convert the RGB image to a grayscale image, and then perform data normalization on the grayscale image to obtain a standardized grayscale image; select the middle area from the standardized grayscale image, calculate the mean of the middle area, divide the pixel values ​​in the mean of the middle area into two categories by the K-means clustering algorithm, and obtain the segmentation threshold based on the center points of the two categories of pixel values; convert the standardized grayscale image into a binary image based on the segmentation threshold; Step 2: performing an erosion operation on the binary image to obtain an eroded binary image; performing an expansion operation on the eroded binary image to obtain an expanded binary image; The connected regions in the dilated binary image are marked using the connected region marking function to obtain the marked regions, and then all the marked regions are captured using the unique value finding function to obtain the captured regions; Traverse all captured areas, filter out qualified areas based on bounding box size and position, and add labels corresponding to qualified areas to the list; Traverse all the labels in the list, use the mask image of all zeros to iteratively generate the lung mask image, then convert the lung mask image to uint8 type and expand it again to obtain the expanded lung mask image; Step 3: Allocate and mark the pixels in the binary image to obtain marked pixels; Calculate and obtain the pixel sum of each connected area in the binary image, take the maximum value based on the pixel sum of all connected areas, and obtain the pixel sum of the maximum connected area; The pixel sum of the maximum connected area is screened to obtain a connected area index list that is greater than half of the pixel sum of the maximum connected area; Generate the mask of the largest connected region based on the marked pixels and the connected region index list; A lung parenchyma image processed by lung parenchyma segmentation is obtained based on the expanded lung mask image and the mask of the maximum connected area; Collect the lung parenchyma images processed by the lung parenchyma segmentation part and the original image labels to generate a training data set, and use the training data set to train the model; Step 4, standardizing the lung parenchyma image after the lung parenchyma segmentation process to obtain a standardized image; The standardized image is input into the improved residual shift module, and after residual connection and deep convolution operations, an enhanced feature map is obtained; the enhanced feature map is processed by the multi-head self-attention mechanism to obtain an enhanced feature map processed by the multi-head self-attention mechanism; The standardized image is processed using the group convolution module and the multi-scale group convolution pooling module, and the processing results are skip-connected through the efficient bottleneck module to obtain the connected output feature map; The standardized image is processed using the fast spatial pyramid pooling module to obtain a feature map after the fast spatial pyramid pooling operation; Step 5: Input the enhanced feature map processed by the multi-head self-attention mechanism, the concatenated output feature map, and the feature map after the fast spatial pyramid pooling operation into the efficient detection head, and obtain the bounding box coordinates through multi-scale convolution and DFL mechanism; Objects are predicted and tracked via bounding box coordinates.

2. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 1, characterized in that: In the step 1, the RGB image is converted into a grayscale image, and then the grayscale image is subjected to data normalization processing to obtain a standardized grayscale image; the middle area is selected from the standardized grayscale image, the middle area mean is obtained by calculation, the pixel values ​​in the middle area mean are divided into two categories by the K-means clustering algorithm, and the segmentation threshold is obtained based on the center points of the two categories of pixel values; the standardized grayscale image is converted into a binary image based on the segmentation threshold, and the specific steps are as follows: The relationship between the conversion of RGB images to grayscale images is: ; in, Represents the grayscale value after conversion, Represents the intensity value of the blue channel, Represents the intensity value of the red channel, Represents the intensity value of the green channel; The grayscale image is subjected to data standardization, and the corresponding relationship in the process of obtaining the standardized grayscale image is: ; in, represents the mean value of the grayscale image, represents the standard deviation of the grayscale image, represents a grayscale image, Represents the standardized grayscale image; The middle area is selected from the standardized grayscale image, the mean of the middle area is obtained by calculation, and the pixel values ​​in the mean of the middle area are divided into two categories by the K-means clustering algorithm. The segmentation threshold is obtained based on the center points of the two categories of pixel values. The relationship between the corresponding process is: ; in, and Represents the number of two types of pixel values, and Represent two types of pixel values, represents the segmentation threshold, and Represent two cluster centers respectively; Based on the segmentation threshold, the standardized grayscale image is converted into a binary image. The corresponding relationship is: ; in, represents a binary image, Indicates that it is processed by a function. Indicates the filtered tags.

3. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 2, characterized in that: In step 2, the binary image is corroded to obtain a corroded binary image; the corroded binary image is expanded to obtain an expanded binary image. The specific steps are as follows: Perform an erosion operation on the binary image to obtain the eroded binary image: ; in, Represents a structural element, represents a translated version of the structural element, represents the corrosion operation, Represents the pixel position in a binary image; The binary image after corrosion is expanded to obtain the expanded binary image. The relationship between the corresponding process is: ; in, represents the expansion operation, represents the empty set, Represents the pixel position in the binary image after erosion.

4. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 3, characterized in that: In step 2, all captured areas are traversed, and qualified areas are screened out according to the size and position of the bounding box, and the labels corresponding to the qualified areas are added to the list; all labels in the list are traversed, and a lung mask image is generated by iteration using an all-zero mask image, and then the lung mask image is converted to a uint8 type and expanded again to obtain an expanded lung mask image. The specific steps are as follows: Traverse all captured areas, filter out qualified areas according to the size and position of the bounding box, and add the labels corresponding to the qualified areas to the list. The relationship in the corresponding process is: ; in, represents the bounding box of the region, Indicates the filter conditions. Represents the coordinate values ​​of the bounding box, Represents the bounding box in the region attribute; Traverse all the labels in the list, and use the mask image of all zeros to iteratively generate the lung mask image. The corresponding relationship in the process is: ; in, represents the lung mask image, Indicates the label, Indicates the filtered tags.

5. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 4, characterized in that: In the step 3, the pixel points in the binary image are allocated and marked to obtain the marked pixel points; the pixel sum of each connected area in the binary image is calculated and obtained, and the maximum value is taken based on the pixel sum of all connected areas to obtain the pixel sum of the maximum connected area; The pixel sum of the maximum connected area is screened to obtain a connected area index list that is greater than half of the pixel sum of the maximum connected area; based on the marked pixels and the connected area index list, a mask of the maximum connected area is generated. The specific steps are as follows: The pixel points in the binary image are allocated and marked to obtain the marked pixel points. The corresponding relationship in the process is: ; in, Indicates the midpoint of the marked image The label value of represents the labeled image, Indicates the current label value; Calculate and obtain the pixel sum of each connected area in the binary image. The corresponding process has the following relationship: ; in, Indicates The sum of pixels in the connected region is Indicates A set of pixels in a connected region, Represents the midpoint of a binary image The pixel value of Based on the pixel sum of all connected areas, the maximum value is taken to obtain the pixel sum of the maximum connected area. The corresponding relationship is: ; in, represents the total number of connected regions, represents the pixel sum of the largest connected region, Indicates The pixel values ​​of the connected regions; The pixel sum of the maximum connected area is screened to obtain the connected area index list that is greater than half of the pixel sum of the maximum connected area. The relationship between the corresponding process is: ; in, represents the pixel sum of the connected region, A list of connected region indices that represent pixels that are larger than half of the maximum connected region; Based on the marked pixels and the connected area index list, the mask of the maximum connected area is generated. The relationship between the corresponding process is: ; in, Represents the midpoint of the maximum connected region The value of .

6. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 5, characterized in that: In step 4, the lung parenchyma image after the lung parenchyma segmentation part is subjected to standardization processing to obtain a standardized image. The relationship between the corresponding process is: ; in, represents the normalized image, represents the lung parenchyma image after the lung parenchyma segmentation part is processed. represents the standard deviation of the lung parenchyma image after the lung parenchyma segmentation part is processed. Represents the mean value of the lung parenchyma image after processing the lung parenchyma segmentation part.

7. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 6, characterized in that: In step 4, the standardized image is input into the improved residual shift module, and an enhanced feature map is obtained after residual connection and deep convolution operations; the enhanced feature map is processed by the multi-head self-attention mechanism to obtain an enhanced feature map processed by the multi-head self-attention mechanism, wherein the expression of the residual connection is: ; in, represents the convolution operation, represents the normalized image, represents the weight of the convolutional sum, represents the output feature map; The depth convolution operation includes depth-wise separable convolution operation and point-wise convolution operation. The expression of depth-wise separable convolution is: ; in, represents the channels of the normalized image, represents the channel of the convolution kernel, Represents the elements of the output feature map, Represents the row index, Represents the column index, represents the column index of the convolution kernel, represents the row index of the convolution kernel, represents the height of the convolution kernel, Indicates the width of the convolution kernel; The expression of point-by-point convolution is: ; in, Represents the output feature map The pixel value at represents the channel index of the normalized image, Represents the total number of channels of the normalized image, Indicates the image after normalization The pixel value at Indicates that the convolution kernel of size 1×1 is in the channel The weight on represents the row index of the output feature map, Represents the column index of the output feature map; The expression of multi-head self-attention is: ; in, represents the query matrix, represents the key matrix, represents the value matrix, represents the transpose operation, Represents the scaling factor.

8. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 7, characterized in that: In step 4, the standardized image is processed using a group convolution module and a multi-scale group convolution pooling module, and the processing result is jump-connected through an efficient bottleneck module to obtain a connected output feature map, wherein the group convolution module includes a group convolution operation and a channel shuffle operation, and the relationship corresponding to the group convolution operation is: ; in, Indicates the number of groups, Indicates The input feature map of the group, represents the convolution kernel, and Both represent the relative displacement offset of the convolution kernel; The following relationship is included in the process of channel shuffle operation: ; in, Represents a shuffle operation; The multi-scale group convolution pooling module includes a lightweight convolution operation and a SiLU activation function, where the relationship between the lightweight convolution operation is: ; in, represents the concatenated feature map, All represent standardized images; The expression of SiLU activation function is: ; in, Represents the feature map after the convolution operation, Represents the Sigmoid function; In the efficient bottleneck module, the expression of the skip connection is: ; in, Represents the feature map output by the network layer, represents the output feature map after residual connection, Represents the input feature map.

9. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 8, characterized in that: In step 4, the standardized image is processed using a fast spatial pyramid pooling module to obtain a feature map after a fast spatial pyramid pooling operation. The relationship between the corresponding process is: ; in, Represents the feature map after fast spatial pyramid pooling operation, , and They represent the maximum pooling operations of 5×5, 9×9, and 13×13 respectively.

10. The method for mobile terminal lung nodule detection and multi-level classification based on deep learning according to claim 9, characterized in that: In step 5, the enhanced feature map processed by the multi-head self-attention mechanism, the connected output feature map, and the feature map after the fast spatial pyramid pooling operation are input into the efficient detection head, and the bounding box coordinates are obtained through multi-scale convolution and DFL mechanism, wherein the calculation relationship of the multi-scale convolution is: ; in, Indicates A collection of convolution kernels, Represents the total number of sets of different convolution kernels, and They represent the offset under different expansion rates respectively; The bounding box coordinates are obtained through the DFL mechanism, and the corresponding relationship is: ; in, represents the adjusted bounding box coordinates, represents the prior box, represents the scaling factor, Represents the bounding box adjustment operation, Represents the predicted values ​​of the bounding box parameters.

Citation Information

Patent Citations

  • Pulmonary nodule benign and malignant identification model training method, application method and system

    CN116468103A

  • Pulmonary nodule detection method based on progressive multi-scale network model

    CN118823547A