Lung focus identification method and system based on image deep learning
Through adaptive contrast enhancement and noise suppression, multi-scale feature matching and image registration of elastic deformation models, combined with attention mechanism and cascade structure, the image complexity and individual differences in lung lesion recognition are solved, and lesion recognition with high accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510511658.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing lung lesions identification methods rely on the experience of doctors, pose a risk of misdiagnosis and misdiagnosis, and the image quality is complex, which affects the accuracy of identification, making it difficult to effectively deal with problems such as clarity, respiratory movement and individual differences in lung images.
Adaptive contrast enhancement and noise suppression technology are used to combine multi-scale feature matching and elastic deformation models for image registration, and a convolutional neural network with multi-scale receptive fields is built, attention mechanism is introduced, classification and regression network of cascade structure is built, morphological filtering and connectivity domain analysis are applied, uncertainty evaluation is carried out, and artifacts and noise are removed.
It improves the accuracy and reliability of lung lesions identification, reduces missed diagnosis and misdiagnosis, provides more effective diagnostic auxiliary tools, and ensures the reliability and accuracy of identification results.
Smart Images

Figure CN120388005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly refers to a method and system for identifying lung lesions based on image deep learning. Background Art
[0002] There are various types of lung diseases. Lung cancer, as one of the cancers with relatively high incidence and mortality rates globally, early diagnosis is crucial for improving the survival rate of patients. Other lung diseases such as pneumonia and tuberculosis also seriously affect people's health. Currently, the identification of lung lesions mainly relies on medical imaging examinations such as X-rays, CTs, MRIs, etc. However, traditional image recognition methods have many drawbacks.
[0003] On the one hand, manual film reading highly depends on doctors' professional experience and subjective judgment. During the long process of reading films, doctors are prone to a decline in attention due to fatigue, thus increasing the risk of missed diagnosis and misdiagnosis. The differences in diagnostic criteria and experience among different doctors also make it difficult to ensure the consistency of diagnostic results. On the other hand, there are complex interference factors in lung images themselves. The morphological changes of the lungs caused by respiratory movement and the natural differences in lung structures among different individuals will all affect the image quality and increase the difficulty of lesion identification. Moreover, there are often noises and artifacts in the images, which may be misjudged as lesions, further reducing the accuracy of identification.
[0004] With the development of computer technology and artificial intelligence, methods based on image deep learning have gradually been applied to the field of lung lesion identification. However, when dealing with lung images, existing deep learning methods face many challenges, such as the clarity of lung images, the influence of respiratory movement and individual differences on images, the accurate extraction of lesion features, and the uncertainty assessment of identification results.
[0005] Therefore, a method and system for identifying lung lesions based on image deep learning have become an urgent problem to be solved by people. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and system for identifying lung lesions based on image deep learning, so as to improve the accuracy and reliability of lung lesion identification, reduce missed diagnosis and misdiagnosis, and provide a more effective diagnostic auxiliary tool for doctors.
[0007] To solve the above technical problem, the technical solution provided by the present invention is: A method for identifying lung lesions based on image deep learning, including the following steps,
[0008] S1. Enhance and register the input lung images to construct an annotated data set;
[0009] S2. Construct a convolutional neural network model that can capture local details and global structural information in lung images and has a multi-scale receptive field. Introduce an attention mechanism to make the model focus on the lesion area, and construct a cascaded classification and regression network to identify the lesion location and estimate its size in a coarse-to-fine manner.
[0010] S3. Apply morphological filtering and connected component analysis to further purify the lesion area and remove artifacts and noise.
[0011] S4. Conduct an uncertainty assessment on the recognition results to obtain the lesion recognition results.
[0012] Furthermore, in step S1, enhance the clarity of lung images through adaptive contrast enhancement and noise suppression; achieve the registration of lung images through multi-scale feature matching and elastic deformation models to eliminate the influence of respiratory motion and individual differences on the recognition results; use a semi-automatic annotation tool to construct a lung lesion annotation dataset.
[0013] Furthermore, the method of adaptive contrast enhancement is as follows: Use the adaptive histogram equalization method to divide the lung image into several non-overlapping small blocks, perform histogram equalization within each small block respectively, and then eliminate the boundary effect between small blocks through bilinear interpolation.
[0014] Furthermore, the method of noise suppression is as follows: For each pixel in the lung image, search for pixel blocks similar to this pixel in its neighborhood, assign weights to each similar pixel block according to the similarity between pixel blocks, and then perform weighted average on the pixel values of these similar pixel blocks to obtain the filtered value of this pixel. Let the neighborhood of pixel p in image I be N p , for pixel q in the neighborhood, its similarity weight w(p,q) with p is calculated as follows:
[0015]
[0016] where V(p) and V(q) are pixel blocks centered on p and q respectively, h is a parameter controlling the similarity attenuation, is the normalization factor. The value I ’ (p) of the filtered pixel p is:
[0017] The method to achieve the registration of lung images is as follows:
[0018] First, use the Scale-Invariant Feature Transform (SIFT) algorithm to extract feature points of lung images at different scales. For two images I1 and I2, calculate their SIFT feature descriptors respectively, and then find the corresponding feature point pairs in the two images through a feature matching algorithm (such as nearest neighbor matching). Then, use the Thin Plate Spline (TPS) elastic deformation model to register the images. This model calculates the deformation field of the images by minimizing the deformation energy between control points. Let the control point set be Its corresponding points in the two images are The TPS model obtains the deformation field by solving the following energy minimization problem:
[0019]
[0020] where f is the deformation function and λ is the parameter controlling smoothness.
[0021] Furthermore, the convolutional neural network model with multi-scale receptive fields includes:
[0022] Multiple convolutional layers, each convolutional layer is configured with convolutional kernels of different sizes to capture feature information at different scales; for the input feature map X, the output feature map Y of the convolutional layer is calculated as follows:
[0023]
[0024] where M and N are the sizes of the convolutional kernels, C is the number of channels of the input feature map, W is the convolutional kernel weight, and b is the bias term.
[0025] Pooling layers, used to reduce the resolution of the feature map and reduce the computational amount; for the input feature map X, the output feature map Y of the pooling layer is calculated as follows:
[0026] Feature fusion layers, which splice the feature maps output by different convolutional layers in the channel dimension to generate feature representations; let the feature maps output by different convolutional layers be X1, X2, …, X n , and the output feature map Y of the feature fusion layer is: Y = concat(X1, X2, …, X n ).
[0027] Furthermore, the method for obtaining the enhanced feature map after introducing the attention mechanism is as follows:
[0028] For the input feature map X, obtain the channel descriptor z through global average pooling:
[0029]
[0030] where H and W are the height and width of the feature map, and c is the number of channels.
[0031] The channel attention weight s is obtained through two fully connected layers and an activation function:
[0032] s = σ(W2ReLU(W1z));
[0033] where W1 and W2 are the weight matrices of the fully connected layers, and σ is the Sigmoid activation function.
[0034] The enhanced feature map Y is obtained by multiplying the channel attention weight with the input feature map:
[0035] Y i,j,c = s c X i,j,c .
[0036] Furthermore, the method for constructing the classification and regression network with a cascade structure is as follows:
[0037] The first-level network: A convolutional neural network is established. The input is the multi-scale feature extraction network and the feature map enhanced by the attention mechanism, and a heat map containing the probability of the lesion location is output; it is trained through the cross-entropy loss function:
[0038]
[0039] where y i is the true label, is the predicted probability, and N is the number of samples.
[0040] The second-level network: Secondary convolution and fully connected operations are performed on the regions with relatively high probabilities in the heat map output by the first-level network, and the precise location coordinates (x, y) and size estimation values (w, h) of the lesion are output; it is trained through the mean squared error loss function:
[0041]
[0042] where (x i , y i , w i , h i ) are the true location and size, is the predicted value.
[0043] Furthermore, the morphological operations include:
[0044] Using opening and closing operations to remove small objects and fill small holes in the image at the lesion location; among them, the specific content of the opening operation is as follows: First, perform an erosion operation on the binary image of the lesion area, and then perform a dilation operation. The erosion operation is:
[0045]
[0046] Among them, A is the input image, B is the structuring element, and (B) z is the result of translating the structuring element B by z; the dilation operation is as follows:
[0047]
[0048] The specific content of the closing operation is as follows: first perform the dilation operation, and then perform the erosion operation.
[0049] Use the labeled connected component algorithm (such as 4-connected or 8-connected) to process the images after the opening and closing operations, and label the independent lesion regions; according to the characteristics of the lesion regions such as area and perimeter, remove the regions with relatively too small area or relatively too long perimeter, and consider that these regions may be artifacts or noise.
[0050] Adopt the cubic spline interpolation method to smooth the lesion boundary; let the discrete points on the lesion boundary be Fit these points through the cubic spline interpolation function S(x), so that S(x) is a cubic polynomial on each sub-interval [x i , x i:1 , and satisfies certain continuity conditions.
[0051] Furthermore, in step S4, perform an uncertainty assessment on the recognition result, and the method for obtaining the lesion recognition result is as follows:
[0052] Adopt the Monte Carlo Dropout method for uncertainty assessment. During the model training process, introduce Dropout layers in the fully connected layer and the convolutional layer. In the test phase, perform multiple forward propagations on each input lung image sample. Each time during forward propagation, still randomly discard neurons according to the probability set during training. This is like predicting the sample from different "perspectives", and the result of each prediction will vary because of the different neurons being discarded.
[0053] For each sample, multiple prediction results will be obtained after multiple forward propagations. Evaluate the uncertainty by calculating the variance of these prediction results. Let the multiple prediction results of the i-th sample be where m is the number of forward propagations, then the uncertainty U i is calculated as follows:
[0054]
[0055] Among them, is the mean of the prediction results. The larger the variance, the greater the fluctuation of the prediction result and the higher the uncertainty; conversely, the smaller the variance, the lower the uncertainty.
[0056] Set an uncertainty threshold, and divide the recognition results into high-confidence and low-confidence results according to this threshold; for high-confidence results, directly output them as the final lesion recognition results; for low-confidence results, they are judged manually by professionals or the imaging data is re-acquired for recognition;
[0057] Comprehensively consider the results of uncertainty evaluation, and output the final lesion recognition results including lesion location, type, size, and confidence information.
[0058] The present invention also provides a lung lesion recognition system based on image deep learning, including a preprocessing module, a deep learning model, and a postprocessing module;
[0059] The preprocessing module is used to enhance, register, and construct an annotation dataset for the input lung images;
[0060] The deep learning model module, including a multi-scale feature extraction network, an attention mechanism enhancement, and a cascaded classification and regression network, is used to train a deep learning model to recognize lung lesions;
[0061] The postprocessing module is used to perform morphological operations and uncertainty evaluation on the recognition results of the deep learning model, and give the recognition results.
[0062] The advantages of the present invention compared with the prior art are as follows:
[0063] The present invention effectively improves the clarity of lung images and reduces noise interference through adaptive contrast enhancement and noise suppression, providing a better data basis for subsequent lesion recognition. The image registration achieved through multi-scale feature matching and elastic deformation models can eliminate the influence of respiratory motion and individual differences on the recognition results and improve the accuracy of recognition.
[0064] The convolutional neural network model with multi-scale receptive fields constructed by the present invention can comprehensively capture local details and global structure information in lung images. Combining the attention mechanism, the model can more accurately focus on the lesion area and improve the accuracy of lesion recognition. The cascaded classification and regression network first roughly and then precisely identifies the lesion location and estimates its size, further improving the recognition accuracy.
[0065] The present invention applies morphological filtering and connected component analysis to effectively remove artifacts and noise in the recognition results, purify the lesion area, make the final recognition results more reliable, and reduce misjudgment.
[0066] The present invention uses the Monte Carlo Dropout method to evaluate the uncertainty of the recognition results, providing confidence information of the recognition results for doctors. For results with low confidence, measures such as manual judgment or re-acquisition of data for recognition can be taken to improve the reliability of diagnosis and avoid misdiagnosis and missed diagnosis caused by model uncertainty. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a flowchart of a method for identifying lung lesions based on image deep learning according to the present invention.
[0068] Figure 2 is a model architecture diagram of a convolutional neural network model with multi-scale receptive fields.
[0069] Figure 3 is a system block diagram of a system for identifying lung lesions based on image deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0071] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present invention, its application, or its use.
[0072] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered as part of the specification.
[0073] In all examples shown and discussed herein, any specific value should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.
[0074] The following further details the method and system for identifying lung lesions based on image deep learning according to the present invention with reference to the accompanying drawings.
[0075] Combined with the attached Figures 1-3 , the present invention is introduced in detail.
[0076] A method for identifying lung lesions based on image deep learning includes the following steps:
[0077] S1. Image preprocessing and dataset construction
[0078] Image Enhancement: Improve the clarity of lung images through adaptive contrast enhancement and noise suppression. Adaptive contrast enhancement uses adaptive histogram equalization. The lung image is divided into non-overlapping small blocks, and histogram equalization is performed separately within each small block to enhance local contrast. Then, the bilinear interpolation algorithm is used to eliminate the boundary effect between small blocks, making the whole image transition naturally. For noise suppression, for each pixel in the lung image, similar pixel blocks are searched in its neighborhood. Weights are assigned to each similar pixel block according to the similarity between pixel blocks, and then the pixel values of these similar pixel blocks are weighted and averaged to obtain the filtered value of this pixel. Specifically, let the neighborhood of pixel p in image I be N p , for pixel q in the neighborhood, its similarity weight w(p,q) with p is calculated as follows:
[0079]
[0080] where V(p) and V(q) are pixel blocks centered on p and q respectively, and h is a parameter controlling the similarity attenuation, is the normalization factor. The value I ’ (p) of the filtered pixel p is:
[0081] Image Registration: Use multi-scale feature matching and elastic deformation model to achieve lung image registration. First, the Scale-Invariant Feature Transform (SIFT) algorithm is used to extract feature points of the lung image at different scales. For two images to be registered, I1 and I2, their SIFT feature descriptors are calculated respectively, and then corresponding feature point pairs in the two images are found through a feature matching algorithm (such as nearest neighbor matching). Then, the Thin Plate Spline (TPS) elastic deformation model is used to register the images. This model calculates the deformation field of the image by minimizing the deformation energy between control points. Let the control point set be and its corresponding points in the two images be The TPS model obtains the deformation field by solving the following energy minimization problem:
[0082]
[0083] where f is the deformation function and λ is a parameter controlling smoothness. Through image registration, the influence of respiratory motion and individual differences on the recognition result is effectively eliminated.
[0084] Dataset Construction: Build a labeled dataset of lung lesions with the help of a semi-automatic annotation tool. Professional doctors, assisted by the semi-automatic annotation tool, label the lesions on the preprocessed lung images, and the tool automatically records the annotation information, providing high-quality data support for the training of subsequent deep learning models.
[0085] S2. Construction of Deep Learning Model
[0086] Multi-scale Feature Extraction Network: Construct a convolutional neural network model with multi-scale receptive fields. This model contains multiple convolutional layers, and each convolutional layer is configured with convolutional kernels of different sizes to capture feature information at different scales. For the input feature map X, the output feature map Y of the convolutional layer is calculated as follows:
[0087]
[0088] where M and N are the sizes of the convolutional kernels, C is the number of channels of the input feature map, W is the convolutional kernel weight, and b is the bias term.
[0089] In addition, the model is also equipped with a pooling layer to reduce the resolution of the feature map and the computational amount. For the input feature map X, the output feature map Y of the pooling layer is calculated as follows:
[0090] Finally, through the feature fusion layer, the feature maps output by different convolutional layers are concatenated in the channel dimension to generate a comprehensive feature representation. Let the output feature maps of different convolutional layers be X1, X2, …, X n , and the output feature map Y of the feature fusion layer is: Y = concat(X1, X2, …, X n ).
[0091] Attention Mechanism Enhancement: Introduce an attention mechanism into the model to enable the model to focus on the lesion area. For the input feature map X, obtain the channel descriptor z through global average pooling:
[0092]
[0093] where H and W are the height and width of the feature map, and c is the number of channels.
[0094] Then, obtain the channel attention weight s through two fully connected layers and activation functions:
[0095] s = σ(W2ReLU(W1z));
[0096] where W1 and W2 are the weight matrices of the fully connected layers, and σ is the Sigmoid activation function.
[0097] Finally, multiply the channel attention weight by the input feature map to obtain the enhanced feature map Y, and the formula is Y i,j,c = s c X i,j,c .
[0098] Cascaded Classification and Regression Network: Construct a cascaded classification and regression network to identify the lesion location and estimate its size in a coarse-to-fine manner. The first-level network is a convolutional neural network. The input is the feature map enhanced by the multi-scale feature extraction network and the attention mechanism. It outputs a heat map containing the probability of the lesion location and is trained using the cross-entropy loss function:
[0099]
[0100] where y i is the ground truth label, is the predicted probability, and N is the number of samples.
[0101] The second-level network performs secondary convolution and fully connected operations on the regions with relatively high probabilities in the heat map output by the first-level network, and outputs the precise location coordinates (x, y) and size estimate (w, h) of the lesion, and is trained using the mean squared error loss function:
[0102]
[0103] where (x i , y i , w i , h i ) are the ground truth location and size, is the predicted value.
[0104] S3. Post-processing and Refinement
[0105] Morphological Filtering: Apply morphological filtering to further refine the lesion area and remove artifacts and noise. Use opening and closing operations to process the image at the lesion location. The opening operation first performs an erosion operation on the binary image of the lesion area and then a dilation operation. The erosion operation is where A is the input image, B is the structuring element, and (B) z is the result of translating the structuring element B by z; the dilation operation is: The closing operation first performs a dilation operation and then an erosion operation. Through the opening and closing operations, small objects in the image are removed and small holes are filled.
[0106] Connected Component Analysis: Use the labeled connected component algorithm (such as 4-connected or 8-connected) to process the image after the opening and closing operations, and label the independent lesion areas. According to the characteristics such as the area and perimeter of the lesion area, remove the areas with relatively small area or relatively long perimeter, which are probably artifacts or noise.
[0107] Boundary Smoothing: Let the discrete points on the lesion boundary be Fit these points using the cubic spline interpolation function S(x) so that S(x) is in each sub-interval [x i,x i:1 It is a cubic polynomial above and satisfies certain continuity conditions, thus making the lesion boundary clearer and more accurate.
[0108] S4. Uncertainty Assessment and Result Output
[0109] Uncertainty Assessment: The Monte Carlo Dropout method is used to assess the uncertainty of the recognition results. During the model training process, Dropout layers are introduced in the fully connected layer and the convolutional layer. In the test phase, multiple forward propagations are performed for each input lung image sample. Each time during forward propagation, neurons are randomly discarded according to the probability set during training, and the sample is predicted from different "perspectives". The prediction results vary each time due to different discarded neurons. For each sample, multiple prediction results are obtained after multiple forward propagations. The uncertainty is evaluated by calculating the variance of these prediction results. Let the multiple prediction results of the i-th sample be where m is the number of forward propagations, then the uncertainty U of this sample i is calculated as follows:
[0110]
[0111] where, is the mean of the prediction results. The larger the variance, the greater the fluctuation of the prediction results and the higher the uncertainty; conversely, the smaller the variance, the lower the uncertainty.
[0112] Result Output: Set an uncertainty threshold, and divide the recognition results into high-confidence and low-confidence results according to this threshold. For high-confidence results, they are directly output as the final lesion recognition results, including lesion location, type, size, and confidence information. For low-confidence results, they are manually judged by professionals or the image data is re-collected for recognition to ensure the reliability of the diagnosis results.
[0113] This application also provides a lung lesion recognition system based on image deep learning for implementing the above method. This system includes a preprocessing module, a deep learning model, and a postprocessing module.
[0114] Preprocessing Module: Responsible for enhancing, registering, and constructing the annotation dataset for the input lung images, and providing high-quality data for the subsequent deep learning model.
[0115] Deep Learning Model Module: Includes a multi-scale feature extraction network, an attention mechanism enhancement, and a cascaded classification and regression network, and realizes the accurate recognition of lung lesions by training the deep learning model.
[0116] Post - processing module: Perform morphological operations and uncertainty assessment on the recognition results of the deep - learning model, and finally give accurate lesion recognition results.
[0117] The specific implementation process of a lung lesion recognition method and system based on image deep learning of the present invention is as follows:
[0118] I. Image pre - processing and dataset construction
[0119] (I) Image enhancement
[0120] 1. Adaptive contrast enhancement:
[0121] Both images I1 and I2 are grayscale images with a size of 512×512 pixels. Each image is divided into non - overlapping small blocks of size 32×32. In this way, each image is divided into (512÷32)×(512÷32) = 256 small blocks.
[0122] Perform adaptive histogram equalization on each small block. For example, for a certain small block in image I1, calculate its grayscale histogram, and then re - distribute the grayscale values according to the rules of histogram equalization to enhance the contrast within the small block.
[0123] Use bilinear interpolation to eliminate the boundary effect between small blocks. Assume that there is a pixel P at the boundary of two adjacent small blocks, its coordinates in the left small block are (x1, y1), and its coordinates in the right small block are (x2, y2). According to the bilinear interpolation formula, the final grayscale value of this pixel is jointly determined by its positions in the two small blocks and the grayscale values of surrounding pixels. In this way, the transition at the small - block splicing is natural, and the visual effect of the whole image is more continuous.
[0124] 2. Noise suppression:
[0125] Set the pixel neighborhood size to 7×7, and the parameter h for controlling the similarity attenuation is 15.
[0126] Taking the pixel p with coordinates (100, 100) in image I2 as an example, its neighborhood N p is a 7×7 area centered on this pixel. Search for pixel blocks similar to pixel p within this neighborhood. For pixel q within the neighborhood, assume its coordinates are (102, 101), and calculate its similarity weight w(p, q) with pixel p.
[0127] First, determine the pixel blocks V (p) and V (q) , where the pixel - block size is set to 3×3. Calculate ||V(p)-V(q)|| 2 , that is, the sum of the squares of the differences between the pixel values at the corresponding positions of the two pixel blocks. Assume that after calculation ||V(p)-V(q)||2 = 20。
[0128] Calculate the normalization factor Z (p) , for the neighborhood N p calculate w(p, q) for all pixels q within it and sum them up to obtain Z (p) . Assume Z (p) = 50。
[0129] Then
[0130] According to the formula perform a weighted sum of the weights and corresponding pixel values of all pixels within the neighborhood to obtain the gray value of pixel p after filtering. Assume that after calculating the weights and pixel values of other pixels within the neighborhood, finally I′(100, 100) = 120 (the original pixel value is assumed to be 110), and the noise suppression process for this pixel is completed. Perform such operations on all pixels in the image in sequence to complete the noise suppression of the entire image.
[0131] (2) Image registration
[0132] 1. Feature point extraction and matching:
[0133] Use the SIFT algorithm to extract feature points from images I1 and I2 at different scales. By constructing a difference-of-Gaussians pyramid, detect extreme points in each scale space. A total of 500 feature points are detected in image I1 and 480 feature points are detected in image I2.
[0134] Calculate the 128-dimensional SIFT feature descriptor for each feature point. For example, for a certain feature point P1 in image I1, its feature descriptor is a 128-dimensional vector [0.1, 0.3, 0.5,...].
[0135] Adopt the nearest neighbor matching algorithm. According to the Euclidean distance between the feature descriptors, find the corresponding feature point pairs in the two images. After matching, 300 pairs of corresponding feature points are obtained.
[0136] 2. Thin plate spline (TPS) elastic deformation model registration:
[0137] Let the control point set be where n = 300, that is, the corresponding feature point pairs found above.
[0138] Set the parameter λ for controlling smoothness to 0.05.
[0139] By solving the energy minimization problem calculate the deformation field of image I2 relative to image I1. After complex mathematical calculations, obtain the deformation function f. Deform image I2 through this deformation function to align it with image I1 in space and complete image registration.
[0140] (III) Dataset Construction
[0141] Using a semi-automatic annotation tool, professional doctors annotate the lesions on the preprocessed images. Suppose there are 100 preprocessed lung images. Doctors outline the lesions in each image in the annotation tool. For example, for one of the images, the doctor marks 3 lesion areas, and the tool automatically records the annotation information such as the positions and shapes of these lesions, finally constructing a lung lesion annotation dataset containing 100 images and their corresponding lesion annotation information.
[0142] II. Construction of Deep Learning Model
[0143] (I) Multi-scale Feature Extraction Network
[0144] 1. Convolutional Layer Settings:
[0145] The constructed convolutional neural network model contains 4 convolutional layers. The first convolutional layer has a kernel size of 3×3, the second convolutional layer has a kernel size of 5×5, the third convolutional layer has a kernel size of 7×7, and the fourth convolutional layer has a kernel size of 9×9. The size of the input feature map X is 128×128×1 (single-channel image data after preprocessing).
[0146] For the first convolutional layer, assume that the value at a certain position (i, j, k) in the input feature map X is X(i, j, k) = 0.5, the convolutional kernel weights W(m, n, l, k) (where m, n range from 0 - 2, l = 0, k = 0), and the bias term b k = 0.1. According to the formula (here M = 3, N = 3, C = 1), calculate the value at the corresponding position of the output feature map Y. Suppose after calculation, Y(10, 10, 0) = 0.8.
[0147] Similarly, calculate the other convolutional layers in turn, and each convolutional layer performs corresponding operations according to its kernel size and the input feature map to capture feature information at different scales.
[0148] 2. Pooling Layer Operations:
[0149] The pooling layer uses max pooling. For the input feature map X, the output feature map Y of the pooling layer is calculated as follows:
[0150] For example, for the 2×2 region at the position (20, 20, 0) in the input feature map X After max pooling, the value at the position (10, 10, 0) in the output feature map Y is 0.7. In this way, the resolution of the feature map is reduced and the computational amount is decreased.
[0151] 3. Feature Fusion Layer Concatenation:
[0152] Let the feature maps output by the four convolutional layers be X1, X2, X3, and X4 respectively. The feature fusion layer concatenates them in the channel dimension. Suppose the size of X1 is 64×64×32, the size of X2 is 64×64×64, the size of X3 is 64×64×128, and the size of X4 is 64×64×256. After concatenation, the size of the output feature map Y of the feature fusion layer is 64×64×(32 + 64 + 128 + 256) = 64×64×480, generating a comprehensive feature representation.
[0153] (II) Attention Mechanism Enhancement
[0154] 1. Channel Descriptor Calculation:
[0155] For the input feature map X (with a size of 64×64×480), the channel descriptor z is obtained through global average pooling. Taking the channel c = 10 as an example, (where H = 64, W = 64). Suppose after calculation, z 10 = 0.3.
[0156] 2. Channel Attention Weight Calculation:
[0157] The channel attention weight s is obtained through two fully connected layers and activation functions. The first fully connected layer has 256 neurons, and the weight matrix W1 has a size of 480×256. The second fully connected layer has 480 neurons, and the weight matrix W2 has a size of 256×480. Suppose after calculation through the ReLU and Sigmoid activation functions, for the channel c = 10, s 10 = 0.8. Here s = σ(W2ReLU(W1z)), where σ is the Sigmoid function.
[0158] 3. Generation of Enhanced Feature Map:
[0159] The channel attention weight is multiplied by the input feature map to obtain the enhanced feature map Y. For the value X(15, 15, 10) = 0.6 at the position (15, 15, 10) in the feature map X, the value Y(15, 15, 10) of the enhanced feature map Y at this position is Y(15, 15, 10) = s 10 ×X(15, 15, 10) = 0.8×0.6 = 0.48. In this way, the model can focus on the features related to the lesion area.
[0160] (III) Cascade Classification and Regression Network
[0161] 1. Training of the First-Level Network:
[0162] The first-level network is a convolutional neural network. The input is the multi-scale feature extraction network and the feature map enhanced by the attention mechanism, and the output is a heat map containing the probability of the lesion location. Suppose there are 50 training samples (selected from the dataset constructed above), and the number of samples N = 50.
[0163] For one of the samples, the true label y i represents the location of the lesion in the image (assumed to be a binary matrix of 0-1, where 1 indicates the presence of a lesion and 0 indicates the absence of a lesion), and the predicted probability is the probability value at the corresponding location in the heat map output by the network.
[0164] It is trained through the cross-entropy loss function Suppose after one training iteration, L cls = 0.2. By continuously adjusting the network parameters, the loss function is gradually reduced to improve the network's initial prediction ability for the lesion location.
[0165] 2. Training of the second-level network:
[0166] Perform secondary convolution and fully connected operations on the regions with relatively high probabilities (for example, probabilities greater than 0.6) in the heat map output by the first-level network, and output the precise location coordinates (x, y) and size estimation values (w, h) of the lesion.
[0167] For one sample, the true location and size are (x i , y i , w i , h i ), assumed to be (50, 60, 20, 30), and the predicted values are
[0168] It is trained through the mean squared error loss function Suppose after one training iteration, L reg = 10. As the training progresses, continuously optimize the network parameters to make the predicted values closer to the true values.
[0169] III. Post-processing and purification
[0170] (1) Morphological filtering
[0171] 1. Opening and closing operations:
[0172] Perform morphological filtering on the lesion area image output by the deep learning model. Suppose a binary lesion area image A is obtained, and the structuring element B is selected as a 5×5 square.
[0173] When performing the opening operation, first perform the erosion operation. For a certain pixel z in the image A, if the structuring element B completely contains within A after being translated by z (i.e., satisfying Then the pixel is retained in the eroded image; otherwise, it is removed. For example, in a certain area of Image A, when the structuring element B is translated, part of it exceeds the range of A, then the pixel corresponding to that translation position is removed in the eroded image.
[0174] After the erosion operation, the dilation operation is performed. For the eroded image, if the structuring element B intersects with the image after being translated by z (i.e., satisfying Then the pixel is retained in the dilated image. Through the opening operation, small objects in the image are removed.
[0175] The closing operation first performs the dilation operation and then the erosion operation, filling the small holes in the image.
[0176] 2. Connected component analysis:
[0177] The 4-connected labeling connected component algorithm is used to process the image after the opening operation and the closing operation. Suppose 10 independent lesion regions are marked after processing. Screening is performed according to the characteristics such as the area and perimeter of the lesion regions. For example, regions with an area less than 50 pixels are set as artifact or noise regions, and regions with a perimeter greater than 200 pixels may also be abnormal regions. After screening, 3 regions with too small an area and 2 regions with too long a perimeter are removed, and the lesion regions are purified.
[0178] 3. Boundary smoothing processing:
[0179] Suppose there are 8 discrete points on the lesion boundary For example, {(10,20),(15,22),(20,25),…}. Through the cubic spline interpolation function S (x) To fit these points, such that S (x) Is a cubic polynomial on each subinterval [x i ,x i:1 and satisfies the conditions of continuous first derivative and second derivative. Suppose after interpolation calculation, the smoothed lesion boundary curve is obtained, making the lesion boundary clearer and more accurate.
[0180] IV. Uncertainty evaluation and result output
[0181] (1) Uncertainty evaluation
[0182] 1. Monte Carlo Dropout operation:
[0183] During the model training process, Dropout layers are introduced in the fully connected layer and the convolutional layer, and the Dropout probability is set to 0.4. In the test phase, 10 forward propagations are performed on the input lung image sample (assumed to be a new image in the above dataset).
[0184] During each forward propagation, neurons are randomly discarded with a probability of 0.4. For example, during the first forward propagation, a certain neuron in the fully connected layer is randomly discarded, while it is retained during the second forward propagation. The prediction results for each time will vary due to different discarded neurons.
[0185] 2. Uncertainty calculation:
[0186] For this sample, 10 prediction results are obtained after 10 forward propagations Assume these prediction results are [0.6, 0.55, 0.62, 0.58, 0.65, 0.53, 0.59, 0.61, 0.57, 0.63] respectively.
[0187] First, calculate the mean of the prediction results (here m = 10), that is
[0188]
[0189] Then calculate the uncertainty of the sample For example, for Successively calculate the sum of the squares of the differences between other prediction results and the mean, and then divide by 9 (because m = 10, m - 1 = 9).
[0190]
[0191] Get U i = 0.001339 (approximate to 0.001, retain three decimal places, consistent with the previous text).
[0192] (2) Result output
[0193] Set the uncertainty threshold to 0.005. Since the uncertainty U i = 0.001339 (approximate to 0.001) of this sample is less than the threshold, it belongs to a high-confidence result. Directly output the final lesion recognition result including the lesion location, type (assumed to be a tumor according to the trained model), size, and confidence information (the confidence is 1 - U i = 1 - 0.001339 ≈ 0.999). If the uncertainty of a certain sample is greater than the threshold, it is a low-confidence result, and it will be manually judged by professionals or the imaging data will be re-collected for recognition to ensure the reliability of the diagnosis result.
[0194] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. All in all, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, creatively design structural manners and embodiments similar to the technical solution, they shall fall within the protection scope of the present invention.
Claims
1. A method for identifying lung lesions based on image deep learning, characterized in that: including the following steps, S1. Enhance and register the input lung images to construct an annotation dataset; S2. Construct a convolutional neural network model that can capture local details and global structural information in lung images and has multi-scale receptive fields. Introduce an attention mechanism to make the model focus on the lesion area, and construct a cascaded classification and regression network to identify the lesion location and estimate its size from coarse to fine; S3. Apply morphological filtering and connected component analysis to further purify the lesion area and remove artifacts and noise; S4. Evaluate the uncertainty of the recognition results to obtain the lesion recognition results.
2. The lung lesion recognition method based on image deep learning according to claim 1, wherein: In step S1, the clarity of lung images is improved through adaptive contrast enhancement and noise suppression; the registration of lung images is achieved through multi-scale feature matching and elastic deformation models to eliminate the influence of respiratory movement and individual differences on the recognition results; a semi-automatic annotation tool is used to construct a lung lesion annotation dataset.
3. The method for identifying lung lesions based on image deep learning according to claim 2, wherein: The method of adaptive contrast enhancement is as follows: The lung image is divided into several non-overlapping small blocks by using the adaptive histogram equalization method, histogram equalization is performed separately within each small block, and then bilinear interpolation is used to eliminate the boundary effect between small blocks.
4. A method for identifying lung lesions based on image deep learning according to claim 2, characterized in that: The method of noise suppression is as follows: For each pixel in the lung image, search for pixel blocks similar to this pixel in its neighborhood, assign weights to each similar pixel block according to the similarity between pixel blocks, and then perform weighted averaging on the pixel values of these similar pixel blocks to obtain the filtered value of this pixel.
5. A method for identifying lung lesions based on image deep learning according to claim 1, characterized in that: The convolutional neural network model with multi-scale receptive fields includes: Multiple convolutional layers, each convolutional layer is configured with convolutional kernels of different sizes to capture feature information of different scales; A pooling layer for reducing the resolution of the feature map and reducing the amount of computation; A feature fusion layer that splices the feature maps output by different convolutional layers in the channel dimension to generate a feature representation.
6. The method for identifying pulmonary lesions based on image deep learning according to claim 5, characterized in that: The method for obtaining the enhanced feature map after introducing the attention mechanism is as follows: Obtain a channel descriptor through global average pooling, obtain channel attention weights through two fully connected layers and activation functions, and multiply the channel attention weights by the input feature map to obtain the enhanced feature map.
7. A method for identifying lung lesions based on image deep learning according to claim 6, characterized in that: The method for constructing a cascaded classification and regression network is as follows: The first-level network: Build a convolutional neural network, the input is the multi-scale feature extraction network and the feature map enhanced by the attention mechanism, and output a heat map containing the probability of the lesion location; The second-level network: Perform secondary convolution and fully connected operations on the area with relatively high probability in the heat map output by the first-level network, and output the accurate position coordinates and size estimation value of the lesion.
8. The method for identifying lung lesions based on image deep learning according to claim 7, wherein: The morphological operations include: Use opening and closing operations to remove small objects and fill small holes in the image at the lesion location; use the labeled connected component algorithm to process the image after opening and closing operations to label independent lesion areas; remove areas with relatively too small area or relatively too long perimeter according to the area and perimeter characteristics of the lesion area; use the cubic spline interpolation method to smooth the lesion boundary.
9. A method for identifying lung lesions based on image deep learning according to claim 8, characterized in that: In step S4, the method for evaluating the uncertainty of the recognition results to obtain the lesion recognition results is as follows: Set an uncertainty threshold, and divide the recognition results into high-confidence and low-confidence results according to this threshold; for high-confidence results, directly output them as the final lesion recognition results; while for low-confidence results, they are judged manually by professionals or the imaging data is re-acquired for recognition. The result output of the uncertainty evaluation includes lesion location, type, size, and confidence information.
10. A lung lesion recognition system based on image deep learning, which is used to implement the lung lesion recognition method based on image deep learning described in any one of claims 1-9, and is characterized in that: It includes a preprocessing module, a deep learning model, and a postprocessing module. The preprocessing module is used to enhance, register, and construct an annotation dataset for the input lung images. The deep learning model module, including a multi-scale feature extraction network, an attention mechanism enhancement, and a cascaded classification and regression network, is used to train the deep learning model to identify lung lesions. The postprocessing module is used to perform morphological operations and uncertainty evaluation on the recognition results of the deep learning model and give the recognition results.
Citation Information
Patent Citations
Honeycomb lung lesion image segmentation identification method based on Transform semi-supervised algorithm
CN117523203A
Recognition method of capsular cavity type lung cancer image and training method of lung recognition model
CN119478482A
Cited By
Fry health state monitoring method based on image recognition
CN120787852A
Lung CT image focus intelligent segmentation and identification method and system
CN121391859A
A lung CT image lesion intelligent segmentation and identification method and system
CN121391859B
Dynamic image registration method for pulmonary atelectasis focus volume change monitoring
CN121414806A
A dynamic image registration method for monitoring atelectasis lesion volume change
CN121414806B