Lung Nodule Risk Assessment Method Based on Multi-Scale, Multi-Receptive Field and Multi-Attention Fusion

By introducing a deep learning network module TMFM with multi-scale, multi-receptive field and multi-attention fusion in lung nodules detection and classification, the problem of malignant nodules risk level assessment and feature extraction in the prior art is solved, and a higher accuracy and adaptive pulmonary nodules risk assessment is achieved.

CN119724588BActive Publication Date: 2025-06-13EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510223238.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the risk level of malignant nodules in the detection and classification of lung nodules, and the single-scale feature extraction and fixed receptive field design are difficult to fully capture the effective information in CT images, especially in complex backgrounds or multi-objective scenarios.

Method used

A pulmonary nodule risk assessment method based on multi-scale multi-receptive field and multi-attention fusion is proposed. By constructing a deep learning network module TMFM, including a multi-branch structure feature extraction module and feature fusion module, the spatial attention mechanism, channel attention mechanism and pixel attention mechanism are used for feature fusion.

Benefits of technology

Comprehensive feature analysis and optimization of CT scan images of pulmonary nodules is achieved, which significantly improves classification accuracy, enhances the generalization ability and noise resistance of the model, can capture key information in the image more accurately, and improves the adaptability and accuracy of pulmonary nodules risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724588B_ABST
    Figure CN119724588B_ABST
Patent Text Reader

Abstract

The present invention proposes a lung nodule risk assessment method based on multi-scale, multi-receptive field and multi-attention fusion. First, a lung nodule image classification model based on a deep learning network is constructed. The lung nodule image classification model based on the deep learning network includes an image preprocessing module, a deep learning network module TMFM, and a classification module. Then, a data set containing lung nodule CT scan images and class number labels corresponding to the malignant risk degree levels of the lung nodules in the images is obtained, and the lung nodule image classification model based on the deep learning network is trained. Then, the lung nodule CT scan image to be evaluated is input into the trained lung nodule image classification model based on the deep learning network for feature extraction, fusion and classification, and a lung nodule risk assessment classification result of the lung nodule CT scan image is obtained. This method has stronger adaptability and accuracy in the risk assessment of lung nodules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic information technology, and particularly relates to a lung nodule risk assessment method based on multi-scale, multi-receptive field and multi-attention fusion. Background Art

[0002] Lung cancer has always been one of the most common cancers threatening human health worldwide. Therefore, the early diagnosis of lung cancer is extremely important for the early treatment and cure of the disease. Early-stage lung cancer is characterized by small nodules, which can be detected as round or spherical structures in computed tomography scan images (CT scan images). The identification and diagnosis of benign and malignant nodules usually require doctors to extract multiple features of the nodules from CT scan images for analysis, such as size, shape, contour, interval growth, location and calcification of lung nodules during the time between two CT scan examinations. These features help to classify the nodules as benign (non-cancerous) or malignant (cancerous). For example, compared with benign nodules, malignant nodules usually have more irregular boundaries / edges, while benign nodules usually have smoother boundaries. The disadvantage of manual identification is that it requires a lot of time and effort, and it is easy to miss or misdiagnose due to relying solely on personal experience.

[0003] In recent years, convolutional neural networks (CNNs) have achieved great success in the field of image processing. As a core method, convolutional neural networks (CNNs) have been widely applied to tasks such as object detection, image segmentation and image classification, and have also been applied to the classification / detection of lung nodules in recent research. Existing classification methods include: (1) applying three Resnet network topologies / structures to the nodule classification task, where the three Resnet networks are respectively applied to different views / planes of the nodules, namely the axial, coronal and sagittal planes; then the outputs of the three different Resnet networks are passed to a fully connected multi-layer perceptron network. (2) Using three-dimensional (3D) CNNs for lung nodule risk stratification to utilize the volume information in CT scans. (3) A multi-view CNN applied to the lung nodule detection problem, using three scales and four views, generating a total of 12 different images; each image is trained with a separate convolutional neural network, and then all the models are combined and the convolutional neural network is retrained; in the preprocessing stage, linear interpolation, normalized spherical sampling using an icosahedron located at the center of the nodule and a threshold method are used to approximate the nodule radius, and this method achieved a classification rate of 92.1% for benign and malignant binary classification.

[0004] However, the current existing technologies still have the following problems:

[0005] (1) One of the main problems faced by computer-aided diagnosis (CAD) schemes for lung nodule detection and classification is the risk level classification of malignant nodules. Generally speaking, there are also differences in the risk levels among malignant nodules. However, the CAD schemes proposed in all previous studies did not provide solutions in the methods for solving the classification problem of malignant nodules, but only distinguished between benign and malignant nodules.

[0006] (2) Due to the different sizes and various shapes of lung nodules in CT images, it is difficult for single-scale feature extraction or the design of a fixed receptive field to comprehensively capture the effective information in the image. Especially when dealing with complex backgrounds or multi-object scenarios, the performance of the model is often limited.

[0007] (3) Existing methods often only optimize for a single attention mechanism or specific-scale features, and it is difficult to establish an effective feature fusion strategy between multi-scales and multi-attention mechanisms. Summary of the Invention

[0008] The main object of the present invention is to propose a lung nodule risk assessment method based on multi-scale multi-receptive field multi-attention fusion to solve the deficiencies mentioned in the above background technology.

[0009] A lung nodule risk assessment method based on multi-scale multi-receptive field multi-attention fusion proposed by the present invention includes the following steps:

[0010] Step S1, constructing a lung nodule image classification model based on a deep learning network, where the lung nodule image classification model based on the deep learning network includes an image preprocessing module, a deep learning network module TMFM, and a classification module.

[0011] The image preprocessing module is used to perform pre-feature extraction on the input image to obtain the initial feature tensor of the input image.

[0012] The deep learning network module TMFM includes a feature extraction module and a feature fusion module; the feature extraction module is used to perform multi-scale feature extraction on the feature tensor obtained by the image preprocessing module; the feature fusion module is used to fuse the feature tensors extracted by the feature extraction module.

[0013] Among them, the feature fusion module includes a spatial attention mechanism SA, a channel attention mechanism CA, a pixel attention mechanism PA and a feature weighted output module; the spatial attention mechanism SA is used to enhance the significant area in the spatial dimension of the feature tensor; the channel attention mechanism CA is used to assign weights on the channel dimension of the feature tensor, adaptively highlight channels with higher relevance, and reduce the influence of invalid channels; the pixel attention mechanism PA is used to refine the feature weights at the pixel level, perform weighted optimization according to the importance of each pixel point, and highlight the features of key pixels; the feature weighted output module is used to uniformly optimize the feature tensors processed and fused by the spatial attention mechanism SA, the channel attention mechanism CA, and the pixel attention mechanism PA, and generate a key feature tensor containing complete multi-scale, multi-receptive field and multi-attention optimization features;

[0014] The classification module is used to process the key feature tensor and output the pulmonary nodule risk assessment classification result;

[0015] Step S2, obtaining a data set including CT scan images of lung nodules and category number labels of the malignant risk levels of lung nodules corresponding to the images, training the lung nodule image classification model based on the deep learning network constructed in step S1, and obtaining a trained lung nodule image classification model based on the deep learning network;

[0016] In step S3, the CT scan image of the lung nodule to be evaluated is input into the lung nodule image classification model based on the deep learning network trained in step S2 for feature extraction, fusion and classification, so as to obtain the lung nodule risk assessment classification result of the lung nodule CT scan image.

[0017] Furthermore, the image preprocessing module includes an input convolution layer and four residual block groups connected in sequence, wherein the input convolution layer includes a convolution layer, a batch normalization layer, an activation layer and a maximum pooling layer; the four residual block groups respectively contain 3, 4, 6 and 3 convolution layers with skip connections.

[0018] Furthermore, the feature extraction module is a multi-branch structure, which includes six parallel branches, namely a convolution operation, three hole convolution operations, an average pooling layer and a wavelet convolution operation. The features output by the six branches are spliced ​​in the channel dimension as the output of the feature extraction module.

[0019] Furthermore, the feature fusion module extracts features of the feature tensor extracted by the feature extraction module in the spatial dimension and channel dimension through the spatial attention mechanism SA and the channel attention mechanism CA respectively, obtains a spatial feature tensor and a channel feature tensor, adds the spatial feature tensor and the channel feature tensor to obtain a prior feature tensor, and performs weighted optimization on the feature tensor extracted by the feature extraction module and the prior feature tensor through the pixel attention mechanism PA to obtain a pixel feature tensor. After performing dimensional normalization on the pixel feature tensor, it is fused with the spatial feature tensor and the channel feature tensor to obtain an attention fusion feature tensor, and then the attention fusion feature tensor is processed by the feature weighted output module to obtain a key feature tensor.

[0020] Furthermore, the spatial attention mechanism SA includes three parallel spatial pooling modules, a 1×1 convolution operation, and a Sigmoid function. The three parallel spatial pooling modules are a spatial average pooling module, a spatial maximum pooling module, and a spatial median pooling module respectively. The outputs of the three spatial pooling modules are concatenated in the channel dimension and then input into the 1×1 convolution operation for tensor dimension recovery processing, and then normalized through the Sigmoid function. The output of the Sigmoid function is multiplied element-wise with the input of the spatial attention mechanism SA and used as the output of the spatial attention mechanism SA.

[0021] Furthermore, the channel attention mechanism CA includes three parallel channel pooling modules and an expand function. Among them, the three parallel channel pooling modules are a channel average pooling module, a channel maximum pooling module, and a channel median pooling module respectively. The expand function includes a fully connected layer fc1, a Relu activation function, and a fully connected layer fc2 connected in sequence. The output of the fully connected layer fc2 is dot-multiplied with the input of the channel attention mechanism CA to obtain the output of the expand function. Among them, the outputs of the three channel pooling modules are respectively adjusted in weight through the expand function and fused with the input of the channel attention mechanism CA to obtain three feature tensors with adjusted weights. The three feature tensors with adjusted weights are added to obtain the final output of the channel attention mechanism CA.

[0022] Furthermore, the pixel attention mechanism PA includes a high-dimensional conversion operation, a concatenation operation, a rearrangement operation, a convolution operation, and a Sigmoid function.

[0023] Furthermore, the classification module includes an average pooling layer, a flattening operation layer, and a fully connected layer. The average pooling layer is used to organize the key feature tensors output by the deep learning network module TMFM. The flattening operation layer is used to reshape the feature dimensions of the feature tensors to conform to the input dimensions of the fully connected layer. The fully connected layer is used to calculate the probability scores of the lung nodule samples in the input image belonging to five lung nodule malignancy risk levels, generate a classification result vector, and convert the classification result vector into a probability distribution vector. The category number of the lung nodule malignancy risk level corresponding to the item with the largest probability value in the probability distribution vector is taken as the lung nodule risk assessment classification result of the lung nodule CT scan image.

[0024] Starting from multiple scales, multiple receptive fields, and multiple attentions, and considering the size, shape, and malignant characteristics of lung nodules, the present invention designs a multi-scale multi-receptive field multi-attention fusion module, named the deep learning network module TMFM. Based on the deep learning network module TMFM, a lung nodule image classification model based on the deep learning network is constructed, enabling the multi-attention mechanism to be guided by prior knowledge in the deep learning network, which is beneficial for the deep learning network to learn and classify lung nodules more meticulously.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] (1) Through the feature extraction method of multiple scales and multiple receptive fields, the present invention conducts a comprehensive feature analysis and optimization on the lung nodule CT scan image data. The core innovation of this design lies in that by introducing convolutional structures with different scales and receptive fields, the lung nodule image classification model based on the deep learning network can more comprehensively understand the detailed information and large-scale context information in the image, thereby greatly improving the classification accuracy. This feature extraction method avoids the defect of traditional methods that are prone to ignoring local detailed information at a single scale, and at the same time overcomes the problem of insufficient capture of large-range context information by overly simple convolutional structures.

[0027] (2) It combines three different attention mechanisms, namely channel attention, spatial attention, and pixel attention, to dynamically optimize and enhance features. Traditional convolutional neural networks often rely on fixed convolutional kernels to process images. Although this method is simple and effective, it may still overlook important features such as the boundaries and positions of pulmonary nodules. By introducing multiple attention mechanisms, the deep learning network module TMFM of the present invention can weight features according to the importance of different regions and scales in the image, enabling the pulmonary nodule image classification model based on the deep learning network to capture key information more accurately. It can not only focus on the key parts of the image but also maintain strong robustness in complex backgrounds and noises. Compared with traditional feature extraction methods, the attention mechanism of the deep learning network module TMFM significantly improves the generalization ability and anti-noise ability of the model, making the method of the present invention more adaptable and accurate in the risk assessment of pulmonary nodules.

[0028] (3) In the risk assessment of pulmonary nodules, important information in the image is often scattered in different levels and scales. Therefore, how to effectively fuse multi-level and multi-scale feature information has become one of the key issues in improving the performance of the classification model. The deep learning network module TMFM adopts an innovative feature fusion strategy. By using the pixel attention mechanism PA to generate prior information of the image, the prior information is used as a guide to fuse the outputs of the spatial attention mechanism SA and the channel attention mechanism CA. The feature tensors from different dimensions and scales are concatenated and optimized in combination with the attention mechanism, enabling the network to comprehensively understand the features of pulmonary nodules from multiple levels and perspectives, and finally obtaining a more accurate classification result, which also enhances the interpretability of the classification result. Since the pulmonary nodule image classification model based on the deep learning network constructed by the present invention can focus on the key features of the image at multiple levels and scales, the decision-making process of the model becomes clearer and more transparent. This is particularly important for deep learning models in medical image analysis because it can help doctors better understand the judgment basis of the model, thereby increasing the credibility of the model in actual clinical applications. Brief Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the overall structure of the deep learning network module TMFM in a pulmonary nodule risk assessment method based on multi-scale, multi-receptive field, and multi-attention fusion proposed by the present invention.

[0030] Figure 2 It is a schematic diagram of the structure of the spatial attention mechanism SA in the deep learning network module TMFM of the present invention.

[0031] Figure 3 It is a schematic diagram of the structure of the channel attention mechanism CA in the deep learning network module TMFM of the present invention.

[0032] Figure 4 It is a schematic structural diagram of the pixel attention mechanism PA in the deep learning network module TMFM of the present invention.

[0033] Figure 5 It is a distribution diagram of image samples corresponding to the malignant risk degree levels of pulmonary nodules in the LIDC-IDRI dataset used for model training of the present invention.

[0034] Figure 6 It is a heat map generated by benchmark testing of some pulmonary nodule image samples in the LIDC-IDRI dataset by the pulmonary nodule image classification model based on the deep learning network of the present invention.

[0035] Figure 7 It is a heat map generated by benchmark testing of the same pulmonary nodule image sample in the LIDC-IDRI dataset by the pulmonary nodule image classification model based on the deep learning network of the present invention and the existing deep learning network. Detailed implementation manners

[0036] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0037] Referring to the following description and drawings, the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0038] As Figure 1 shown, the embodiments of the present invention provide a pulmonary nodule risk assessment method based on multi-scale multi-receptive field multi-attention fusion. The method includes the following steps:

[0039] Step S1, constructing a pulmonary nodule image classification model based on a deep learning network. The pulmonary nodule image classification model based on the deep learning network includes an image preprocessing module, a deep learning network module TMFM, and a classification module.

[0040] Among them, the image preprocessing module is used to perform pre-feature extraction on the input image. In this embodiment, the image preprocessing module includes an input convolutional layer and four residual block groups. Among them, the input convolutional layer includes a convolutional layer with a convolutional kernel size of 7×7, a batch normalization layer, an activation layer, and a max pooling layer. The activation layer is composed of a Relu activation function; the four residual block groups are arranged in sequence according to the input order. The four residual block groups respectively contain 3, 4, 6, and 3 convolutional layers with skip connections. The skip connection means that the network can skip this layer to transmit data when reaching the threshold, enabling the network to alleviate the gradient disappearance problem through the identity mapping, so as to better train the deep network.

[0041] The deep learning network module TMFM is used for feature extraction. The deep learning network module TMFM includes a feature extraction module and a feature fusion module. The feature extraction module is a multi-branch structure and is used to perform multi-scale feature extraction on the feature tensor obtained by the image preprocessing module; the feature fusion module is used to fuse the feature tensors extracted by the feature extraction module. The feature fusion module includes a spatial attention mechanism SA, a channel attention mechanism CA, a pixel attention mechanism PA, and a feature weighted output module; among them, the spatial attention mechanism SA is used to enhance the significant regions in the spatial dimension of the feature tensor, strengthen the features related to the target region in the feature map through weighting, and at the same time suppress irrelevant or redundant background information to improve the attention to the target region; the channel attention mechanism CA is used to allocate weights in the channel dimension of the feature tensor, adaptively highlight the channels with higher correlation, reduce the influence of invalid channels, and thus improve the efficiency and accuracy of feature expression; the pixel attention mechanism PA is used to refine the feature weights at the pixel level, perform weighted optimization according to the importance of each pixel point, and highlight the features of key pixel points, which is particularly suitable for feature extraction of small targets or target regions with complex boundaries; the feature weighted output module is used to uniformly optimize the feature tensor processed and fused by the spatial attention mechanism SA, the channel attention mechanism CA, and the pixel attention mechanism PA, so as to be able to generate the final key feature tensor by combining the weights assigned by the spatial attention mechanism SA, the channel attention mechanism CA, and the pixel attention mechanism PA. The key feature tensor contains complete multi-scale, multi-receptive field, and multi-attention optimized features, which are displayed in the form of a heat map to provide higher-quality input features for subsequent classification or regression tasks.

[0042] The classification module is used to output the lung nodule risk assessment classification result of the lung nodule CT scan image. The classification module includes an average pooling layer, a flattening operation layer and a fully connected layer. The classification module performs feature sorting on the output key feature tensor obtained by the deep learning network module TMFM through the average pooling layer, and then reshapes the feature dimension to meet the input dimension of the fully connected layer through the flattening operation layer. Finally, the feature tensor with changed dimension after processing by the flattening operation layer is input into the fully connected layer. The probability score of the lung nodule samples in the input image belonging to the five lung nodule malignancy risk levels is calculated through the fully connected layer, and the lung nodule risk assessment classification result of the lung nodule CT scan image is obtained.

[0043] The specific process of processing the input image by the constructed lung nodule image classification model based on deep learning network is as follows:

[0044] Step S11, input the input image into the image preprocessing module for feature pre-extraction to obtain a feature tensor L; then perform multi-scale feature extraction on the feature tensor L through a multi-branch structure. The multi-branch structure includes six parallel branches, which are a convolution operation, three parallel dilated convolution operations, an average pooling layer, and a wavelet convolution operation. In the convolution operation, the size of the convolution kernel is 1×1, and in the three parallel dilated convolution operations, the size of the convolution kernel is 3×3, and the dilation rates are 3, 5, and 7, respectively. The features output by the six branches are concatenated in the channel dimension to obtain a feature tensor X. The dimension of the feature tensor X is B×C×H×W, where B is the batch size, C is the number of channels, and H and W are the height and width of the feature tensor, respectively.

[0045] Among them, the process of wavelet convolution operation is as follows: using db1 wavelet basis, the feature tensor L is decomposed in the frequency domain to extract multi-band feature information; then, the inverse wavelet transform is used to fuse the multi-band feature information in the frequency domain to obtain the fused feature information; finally, after a convolution operation with a convolution kernel size of 1×1, the features of the fused feature information are re-extracted.

[0046] Step S12, perform feature fusion operation on the feature tensor X obtained in step S11. The spatial dimension and channel dimension features of the feature tensor X are extracted by the spatial attention mechanism SA and the channel attention mechanism CA respectively to obtain the spatial feature tensor Y S and the channel feature tensor Y C , and transform the spatial feature tensor Y S and the channel feature tensor Y C Add them together to get the prior feature tensor P1; then use the pixel attention mechanism PA to perform a pixel-level dot multiplication operation on the feature tensor X and the prior feature tensor P1 to get the pixel feature tensor Y P , and the pixel feature tensor Y PPerform dimensional normalization to obtain the feature tensor P2; then multiply the feature tensor P2 element-wise with the spatial feature tensor Y S and the channel feature tensor Y C respectively, and then add the two feature tensors obtained after the element-wise multiplication operation to the feature tensor X obtained in step S11, so as to fuse the feature tensors obtained by guiding the spatial attention mechanism SA and the channel attention mechanism CA using the prior feature tensor P1, and input the fused feature tensor into a convolution operation with a convolution kernel size of 1×1 for processing to obtain the attention fusion feature tensor Y. The specific steps of the feature fusion operation are as follows:

[0047] Step S121, input the feature tensor X obtained in step S11 into the spatial attention mechanism SA and the channel attention mechanism CA respectively to perform feature extraction in the spatial dimension and the channel dimension, and obtain the spatial feature tensor Y S and the channel feature tensor Y C . Add the spatial feature tensor Y S and the channel feature tensor Y C to obtain the prior feature tensor P1.

[0048] As Figure 2 shown, the spatial attention mechanism SA includes three parallel spatial pooling modules, a 1×1 convolution operation, and a Sigmoid function. The three parallel spatial pooling modules are the spatial average pooling module, the spatial max pooling module, and the spatial median pooling module; the outputs of the three spatial pooling modules are concatenated in the channel dimension and then input into the 1×1 convolution operation for tensor dimension restoration processing. The output of the 1×1 convolution operation is normalized by the Sigmoid function, and the output of the Sigmoid function is multiplied element-wise with the input of the spatial attention mechanism SA and used as the output of the spatial attention mechanism SA, that is, the final spatial feature tensor Y S is obtained.

[0049] Among them, the specific steps of performing feature extraction and fusion on the feature tensor X through the spatial attention mechanism SA of the feature fusion module are as follows:

[0050] Step S1211, the feature tensor X obtained in step S11 is respectively extracted through three spatial pooling modules at the spatial level to obtain three spatially pooled feature tensors X avg 、X max 、X median . The three spatial pooling modules are: the spatial average pooling module, the spatial max pooling module, and the spatial median pooling module;

[0051] Among them, the spatial average pooling module is defined as:

[0052] ,

[0053] Among them, X represents the feature tensor, and X avg represents the feature tensor after spatial average pooling, C represents the number of channels of the feature tensor X, represents taking the spatial set of the c-th channel dimension of the feature tensor X. The feature tensor X obtained after being processed by the spatial average pooling module avg has the dimension of B×1×H×W.

[0054] The spatial max pooling module is defined as:

[0055] ,

[0056] Among them, X represents the feature tensor, and X max represents the feature tensor after spatial max pooling, C represents the number of channels of the feature tensor X, and max(X[c,:,:]) represents taking the maximum value of the spatial set of the c-th channel dimension of the feature tensor X. The feature tensor X obtained after being processed by the spatial max pooling module max has the dimension of B×1×H×W.

[0057] The spatial median pooling module is defined as:

[0058] ,

[0059] Among them, X represents the feature tensor, and X median represents the feature tensor after spatial median pooling, C represents the number of channels of the feature tensor X, and median(X[c,:,:]) represents taking the median value of the spatial set of the c-th channel dimension of the feature tensor X. The feature tensor X obtained after being processed by the spatial median pooling module median has the dimension of B×1×H×W.

[0060] Step S1212: Concatenate the spatially pooled feature tensor X avg , the feature tensor X max , and the feature tensor X median on the channel dimension, and then perform a 1×1 convolution operation and a Sigmoid function for tensor dimension recovery and normalization processing to generate the spatial feature tensor A, which is specifically expressed as:

[0061] ,

[0062] Among them, A represents the spatial feature tensor after normalization processing, and the spatial feature tensor A has the dimension of B×1×H×W; represents the Sigmoid function; Conv 1×1 represents a convolution operation with a convolution kernel size of 1×1; Represents an operation of concatenating feature tensors along the channel dimension.

[0063] Step S1213, multiply the spatial feature tensor A obtained in step S1212 element-wise with the feature tensor X to obtain the final spatial feature tensor Y S , specifically represented as:

[0064] ,

[0065] where Y S represents the final spatial feature tensor obtained by the spatial attention mechanism SA, and the spatial feature tensor Y S has dimensions B×1×H×W, represents the Hadamard product of tensors, that is, an operation of multiplying the elements at corresponding positions one by one.

[0066] For example Figure 3 as shown, the channel attention mechanism CA includes three parallel channel pooling modules and an expand function. Among them, the three parallel channel pooling modules are the channel average pooling module, the channel max pooling module, and the channel median pooling module. The expand function includes a fully connected layer fc1, a Relu activation function, and a fully connected layer fc2 connected in sequence. And the output of the fully connected layer fc2 is dot-multiplied with the input of the channel attention mechanism CA to obtain the output of the expand function; among them, the outputs of the three channel pooling modules are respectively adjusted in weight through the expand function and fused with the input of the channel attention mechanism CA, that is, the feature tensor X, to obtain three feature tensors Y 1 , Y 2 and Y 3 ; add the three feature tensors Y 1 、Y 2 and Y 3 to obtain the final output of the channel attention mechanism CA, that is, the channel feature tensor Y C .

[0067] Among them, the specific steps of feature extraction and fusion of the feature tensor X by the channel attention mechanism CA of the feature fusion module are as follows:

[0068] Step S1214, the feature tensor X extracted in step S11 is respectively extracted by three channel pooling modules at the channel level to obtain three channel-pooled feature tensors F avg 、F max 、F median . The three channel pooling modules are respectively: the channel average pooling module, the channel max pooling module, and the channel median pooling module;

[0069] Among them, the channel average pooling module is defined as:

[0070] ,

[0071] Among them, X represents the feature tensor, and F avg represents the feature tensor after channel average pooling. H and W represent the height and width of the feature tensor X, and X[:, i, j] represents the set of channels in the spatial dimension formed by the i-th height and j-th width of the feature tensor X. The feature tensor F avg obtained after being processed by the channel average pooling module has a dimension of B×C×1×1.

[0072] The channel max pooling module is defined as:

[0073] ,

[0074] Among them, X represents the feature tensor, and F max represents the feature tensor after channel max pooling, and max(X[:, i, j]) represents the maximum value of the set of channels in the spatial dimension formed by the i-th height and j-th width of the feature tensor X. The feature tensor F max obtained after being processed by the channel max pooling module has a dimension of B×C×1×1.

[0075] The channel median pooling module is defined as:

[0076] ,

[0077] Among them, X represents the feature tensor, and F median represents the feature tensor after channel median pooling, and median(X[:, i, j]) represents the median value of the set of channels in the spatial dimension formed by the i-th height and j-th width of the feature tensor X. The feature tensor F median obtained after being processed by the channel median pooling module has a dimension of B×C×1×1.

[0078] Step S1215: Create an expand function. The expand function is used to adjust the channel weights of the feature tensors after three-channel pooling and fuse the information of the feature tensor X. The specific representation of the expand function is:

[0079] ,

[0080] Among them, expand(·) represents the expand function, X represents the feature tensor, and F represents the feature tensor F avg 、F max 、F medianAny one of them, fc1 represents a fully connected layer that compresses the feature tensor after channel pooling in the channel dimension, and fc2 represents a fully connected layer that restores the feature tensor to the dimension of the feature tensor X; Relu represents the Relu activation function, X[:,i,j] represents taking the channel set of the spatial dimension formed by the i-th height and the j-th width of the feature tensor X, and F[:,i,j] represents taking the channel set of the spatial dimension formed by the i-th height and the j-th width of the feature tensor F after channel pooling.

[0081] Through the expand function, respectively, for the three feature tensors F after channel pooling obtained in step S1214 avg 、F max 、F median Adjust the channel weights to obtain three feature tensors Y with adjusted weights 1 、Y 2 and Y 3 , and finally add the three feature tensors Y with adjusted weights 1 、Y 2 and Y 3 to obtain the channel feature tensor Y C , specifically expressed as:

[0082] ,

[0083] where Y C represents the final channel feature tensor obtained by the channel attention mechanism CA, and the dimension of the channel feature tensor Y C is B×C×1×1.

[0084] Step S1216, add the spatial feature tensor Y S and the channel feature tensor Y C to obtain the prior feature tensor P1, and the dimension of the prior feature tensor P1 is B×C×H×W.

[0085] Step S122, through the pixel attention mechanism PA, perform pixel-level global information extraction on the feature tensor X obtained in step S11 and the feature tensor P1 obtained in step S121, and then through rearrangement operation, convolution operation, and normalization by the Sigmoid function, obtain the pixel feature tensor Y P .

[0086] As Figure 4 shown, the pixel attention mechanism PA includes two key steps: pixel-level global information extraction and feature fusion. First, perform high-dimensional conversion on the feature tensor X obtained in step S11 and the feature tensor P1 obtained in step S121 respectively, and then splice the two converted feature tensors to form a tensor S. Then, rearrange the dimensions of the tensor S through a rearrangement operation to obtain the rearranged tensor SR for subsequent processing of feature fusion operations. Then, a convolution operation with a convolution kernel size of 1×1 is performed on the rearranged tensor S R to extract global pixel features, and the Sigmoid function is used to normalize the convolution result to obtain the final output of the pixel attention mechanism PA, which is the pixel feature tensor Y P .

[0087] Among them, the specific steps for extracting pixel-level global information through the pixel attention mechanism PA of the feature fusion module are as follows:

[0088] Step S1221: Concatenate the feature tensor X obtained in step S11 and the feature tensor P1 obtained in step S121 (both with dimensions B×C×H×W) in the high dimension to form a new tensor S, and the dimension of the tensor S is B×C×2×H×W. Then, use a rearrangement operation to rearrange the dimensions of the tensor S, converting the tensor S from the dimension B×C×2×H×W to the dimension B×2C×H×W to obtain the rearranged tensor S R , the rearranged tensor S R has the dimension B×2C×H×W.

[0089] Step S1222: Perform a convolution operation on the rearranged tensor S R and finally use the Sigmoid function to normalize the convolution result to obtain the pixel feature tensor Y P , specifically expressed as:

[0090] ,

[0091] where Y P represents the pixel feature tensor, and the dimension of the pixel feature tensor Y P is B×C×H×W.

[0092] Step S1223: Then, perform dimensional normalization on the pixel feature tensor Y P through a Sigmoid function to obtain the feature tensor P2.

[0093] Step S123: Perform feature fusion on the feature tensor X obtained in step S11, the feature tensor P2 obtained in step S122, the spatial feature tensor Y S and the channel feature tensor Y C obtained in step S121, that is, multiply the feature tensor P2 with the spatial feature tensor Y S and the channel feature tensor Y CPerform a dot product operation, and then add the two feature tensors obtained after the dot product operation to the feature tensor X obtained in step S11. The obtained feature tensor after addition is processed through a convolution operation with a convolution kernel size of 1×1 to restore the tensor dimension and perform normalization, resulting in an attention fusion feature tensor Y; specifically expressed as:

[0094] ,

[0095] where Y represents the attention fusion feature tensor.

[0096] Step S13, the feature weighted output module includes a dot product operation and a convolution operation with a convolution kernel size of 1×1. Perform a dot product operation on the attention fusion feature tensor Y obtained in step S12 and the feature tensor X obtained in step S11 to obtain a feature weighted feature tensor, and then perform a convolution operation with a convolution kernel size of 1×1 on the feature weighted feature tensor to reduce the dimension and further extract features, finally generating a key feature tensor that can distinguish the malignant risk degree levels of pulmonary nodules and converting it into a heat map for evaluating the malignant risk degree levels of pulmonary nodules. Name the key feature tensor as feature tensor M, and the dimension of feature tensor M is B×C×H×W.

[0097] Step S14, the classification module consists of an average pooling layer, a flattening operation layer, and a fully connected layer. The classification module is used to calculate the probability scores of the pulmonary nodule samples in the input image belonging to five pulmonary nodule malignant risk degree levels according to the key feature tensor (i.e., feature tensor M). Input the feature tensor M output by the deep learning network module TMFM into the classification module, perform feature sorting through the average pooling layer, then reshape the feature dimension through the flattening operation layer to conform to the input dimension of the fully connected layer, and finally input the feature tensor with the changed dimension obtained after the flattening operation layer into the fully connected layer, and calculate through the fully connected layer to obtain the probability scores of the pulmonary nodule samples in the input image belonging to five pulmonary nodule malignant risk degree levels, and obtain the pulmonary nodule risk assessment classification result of the pulmonary nodule CT scan image. Among them, the specific steps for calculating the probability scores of the pulmonary nodule samples in the input image belonging to five pulmonary nodule malignant risk degree levels according to the feature tensor M are as follows:

[0098] Step S141, input the feature tensor M with a dimension of B×C×H×W obtained by the deep learning network module TMFM into the average pooling layer of the classification module, perform global average pooling on the feature tensor M in the channel dimension, and generate a feature sorted feature tensor M avg , feature tensor M avg has a dimension of B×C×1×1, and the specific process is expressed as:

[0099] ,

[0100] Among them, M represents the feature tensor M obtained by the deep learning network module TMFM, H and W represent the height and width of the feature tensor M, and M[:, i, j] represents the set of channels in the spatial dimension formed by the i-th height and the j-th width of the feature tensor M. M avg represents the feature tensor obtained by performing average pooling on the channel dimension through the average pooling layer of the classification module. The feature tensor M avg has dimensions of B×C×1×1.

[0101] Step S142: Input the feature tensor M obtained in step S141 avg into the flattening operation layer, reshape the dimensions of the feature tensor M avg to obtain the flattened feature tensor M flat . The flattened feature tensor M flat has dimensions of B×C.

[0102] Step S143: The fully connected layer initializes the weight matrix W R and the bias vector b r . Specifically, use a uniform distribution based on the number of channels C of the feature tensor M flat to initialize the weight matrix W R , so that the weight matrix W R is randomly generated within the range of . Initialize the bias vector b r as a random value of a uniform distribution based on the number of channels C of the feature tensor M flat or all zeros, that is, the bias vector b r is randomly generated within the range of or all zeros.

[0103] Input the flattened feature tensor M obtained in step S142 flat into the fully connected layer. The fully connected layer calculates the probability scores of each lung nodule sample in the input image belonging to five lung nodule malignancy risk levels through the weight matrix W R and the bias vector b r to generate the classification result vector R. The specific process is expressed as:

[0104] ,

[0105] Among them, R represents the classification result vector. The dimension of the classification result vector R is 5-dimensional. W R represents the weight matrix updated by each training network iteration. The W of the weight matrix R has dimensions of 5×C, and b r represents the 5-dimensional bias vector updated by each training network iteration.

[0106] Among them, the five malignant risk degree levels of pulmonary nodules include: benign pulmonary nodules, benign pulmonary nodules with a certain tendency to malignancy, pulmonary nodules with ambiguous malignancy degree, pulmonary nodules with general malignancy degree, and highly malignant pulmonary nodules. The corresponding category numbers are 1 to 5 in sequence.

[0107] Step S144: Input the classification result vector R obtained in step S143 into the Softmax function to convert the classification result vector R into a probability distribution vector R1, which is specifically expressed as:

[0108] ,

[0109] Among them, e represents the natural constant, R[:,i] represents the predicted value of the i-th category of the classification result vector R, R1 represents the probability distribution vector with a dimension of 5. The probability distribution vector R1 contains the probabilities that the pulmonary nodule samples in the input image belong to each malignant risk degree level of pulmonary nodules.

[0110] Step S145: Compare the probability distribution vector R1 obtained in step S144 item by item with the five malignant risk degree levels of pulmonary nodules, confirm the category number of the malignant risk degree level of the pulmonary nodule corresponding to the item with the largest probability value in the probability distribution vector R1, and use the category number of the malignant risk degree level of the pulmonary nodule as the final pulmonary nodule risk assessment classification result.

[0111] Step S2: Obtain a data set containing CT scan images of pulmonary nodules and the category number labels of the corresponding malignant risk degree levels of the pulmonary nodules, and train the pulmonary nodule image classification model based on the deep learning network constructed in step S1; after reaching the maximum number of iterations, obtain the model weight parameters of the last round of training; load the trained model weight parameters into the pulmonary nodule image classification model based on the deep learning network to obtain the trained pulmonary nodule image classification model based on the deep learning network, which is used as the pulmonary nodule risk assessment classification network model for the CT scan images of pulmonary nodules in the present invention.

[0112] In this embodiment, an existing data set is used to train the pulmonary nodule image classification model based on the deep learning network of the present invention. Figure 5Shows the number of samples of pulmonary nodules with different risk levels in the LIDC-IDRI dataset used for model training. In the LIDC-IDRI dataset, there are a total of 1018 cases of chest computed tomography images (CT scan images) including diagnosis and lung cancer screening. Each case is marked by 4 experienced chest radiologists for the lung CT scan image samples. Each lung CT scan image has at least one pulmonary nodule sample. In the LIDC-IDRI dataset, there are category number tags marked by radiologists for the pulmonary nodule malignant risk degree levels of the pulmonary nodule samples in the lung CT scan images. The category numbers include 1 to 5, which are used to indicate the suspicious degree of the malignant risk of the pulmonary nodules. Among them, category number 5 indicates highly malignant pulmonary nodules, category number 4 indicates moderately malignant pulmonary nodules, category number 3 indicates pulmonary nodules with ambiguous malignancy, category number 2 indicates benign pulmonary nodules with a certain malignant tendency, and category number 1 indicates benign pulmonary nodules.

[0113] Step S3, input the pulmonary nodule CT scan image to be evaluated into the trained pulmonary nodule image classification model based on the deep learning network in Step S2 for feature extraction, fusion and classification, and obtain the pulmonary nodule risk assessment classification result of the pulmonary nodule CT scan image.

[0114] This embodiment uses the LIDC-IDRI dataset for pulmonary nodule risk assessment tests, compares the pulmonary nodule image classification model based on the deep learning network of the present invention with the existing deep learning networks Resnet34, Resnet50, Resnet101, Resnet152, Densenet121, Densenet264 and ViT in terms of data and model attention respectively, and calculates their respective sample accuracy index, sample weighted accuracy index, area under the curve index (AUC) of the sample receiver operating characteristic ROC curve, and sample average correct rate index (AP). The performance comparison test results are shown in Table 1.

[0115] Table 1 Test result data table of the pulmonary nodule image classification model based on the deep learning network of the present invention and the existing deep learning networks in the LIDC-IDRI dataset for benchmark testing:

[0116]

[0117] According to the test results in Table 1, it can be seen that the method of the present invention not only has significant advantages in the overall classification accuracy, but also can show excellent performance in both the area under the ROC curve (AUC) and the average precision of samples (AP). Especially in the classification task of high-risk categories, it demonstrates extremely high robustness and sensitivity. For example, in the classification results of the difficult-to-distinguish relatively ambiguous malignancy grade category numbers 2 and 3, the method of the present invention greatly improves the accuracy of model classification and the corresponding average precision, verifying the superiority of the lung nodule image classification model based on the deep learning network of the present invention in lung nodule risk assessment.

[0118] Randomly select samples in the LIDC-IDRI dataset and use the lung nodule image classification model based on the deep learning network of the present invention for classification. Draw a heat map before the fully connected layer of the model to obtain Figure 6 the heat map results shown; Figure 6 The first row in Figure 6 is the original image of the lung nodule CT scan image sample in the LIDC-IDRI dataset, Figure 7 and the second row in Figure 7 is the heat map generated by the present invention. The places where the color is closer to red represent that the model pays more attention to the areas here, indicating that the deep learning network module TMFM pays more attention to this part of the image structure. Use the lung nodule image classification model based on the deep learning network of the present invention to generate heat maps on the same lung nodule image respectively and compare them with the existing deep learning networks Resnet34, Resnet50, Resnet101, Resnet152, Densenet121, Densenet264, and ViT. The comparison results are as Figure 7 shown. The results show that compared with the existing deep learning networks, the model of the present invention can not only accurately find the approximate position of the lung nodule, but also can Figure 7 find the pathological image signs of the lung nodule such as blood vessel penetration and adhesion to the lung wall shown in the original image of the lung nodule CT scan image sample in

[0119] The above embodiments only represent one implementation manner of the present invention, but should not be construed as limiting the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A lung nodule risk assessment method based on multi-scale, multi-receptive field and multi-attention fusion, characterized in that: The steps include: Step S1, constructing a lung nodule image classification model based on a deep learning network, wherein the lung nodule image classification model based on a deep learning network includes an image preprocessing module, a deep learning network module TMFM, and a classification module; The image preprocessing module is used to perform feature pre-extraction on the input image to obtain an initial feature tensor of the input image; The deep learning network module TMFM includes a feature extraction module and a feature fusion module; The feature extraction module is used to perform multi-scale feature extraction on the feature tensor obtained by the image preprocessing module; The feature fusion module is used to fuse the feature tensors extracted by the feature extraction module; Among them, the feature fusion module includes a spatial attention mechanism SA, a channel attention mechanism CA, a pixel attention mechanism PA and a feature weighted output module; the spatial attention mechanism SA is used to enhance the significant area in the spatial dimension of the feature tensor; the channel attention mechanism CA is used to assign weights on the channel dimension of the feature tensor, adaptively highlight channels with higher relevance, and reduce the influence of invalid channels; the pixel attention mechanism PA is used to refine the feature weights at the pixel level, perform weighted optimization according to the importance of each pixel point, and highlight the features of key pixels; the feature weighted output module is used to uniformly optimize the feature tensors processed and fused by the spatial attention mechanism SA, the channel attention mechanism CA, and the pixel attention mechanism PA, and generate a key feature tensor containing complete multi-scale, multi-receptive field and multi-attention optimization features; The classification module is used to process the key feature tensor and output the pulmonary nodule risk assessment classification result; Step S2, obtaining a data set including CT scan images of lung nodules and category number labels of the malignant risk levels of lung nodules corresponding to the images, training the lung nodule image classification model based on the deep learning network constructed in step S1, and obtaining a trained lung nodule image classification model based on the deep learning network; In step S3, the CT scan image of the lung nodule to be evaluated is input into the lung nodule image classification model based on the deep learning network trained in step S2 for feature extraction, fusion and classification, so as to obtain the lung nodule risk assessment classification result of the lung nodule CT scan image.

2. The pulmonary nodule risk assessment method based on multi-scale, multi-receptive field and multi-attention fusion according to claim 1 is characterized in that: The image preprocessing module includes an input convolution layer and four residual block groups connected in sequence, wherein the input convolution layer includes a convolution layer, a batch normalization layer, an activation layer and a maximum pooling layer; the four residual block groups respectively contain 3, 4, 6 and 3 convolution layers with skip connections.

3. The lung nodule risk assessment method based on multi-scale, multi-receptive field and multi-attention fusion according to claim 2 is characterized in that: The feature extraction module is a multi-branch structure, which includes six parallel branches, namely a convolution operation, three hole convolution operations, an average pooling layer and a wavelet convolution operation. The features output by the six branches are spliced ​​in the channel dimension as the output of the feature extraction module.

4. The lung nodule risk assessment method based on multi-scale multi-receptive field multi-attention fusion according to claim 3 is characterized in that: The feature fusion module extracts features of the spatial dimension and channel dimension of the feature tensor extracted by the feature extraction module through the spatial attention mechanism SA and the channel attention mechanism CA, respectively, to obtain the spatial feature tensor and the channel feature tensor, and adds the spatial feature tensor and the channel feature tensor to obtain the prior feature tensor, and performs weighted optimization on the feature tensor and the prior feature tensor extracted by the feature extraction module through the pixel attention mechanism PA to obtain the pixel feature tensor, and then normalizes the dimension of the pixel feature tensor and fuses it with the spatial feature tensor and the channel feature tensor to obtain the attention fusion feature tensor. The attention fusion feature tensor is then processed through the feature weighted output module to obtain the key feature tensor.

5. The pulmonary nodule risk assessment method based on multi-scale, multi-receptive field and multi-attention fusion according to claim 4 is characterized in that: The spatial attention mechanism SA includes three parallel spatial pooling modules, a 1×1 convolution operation and a Sigmoid function. The three parallel spatial pooling modules are a spatial average pooling module, a spatial maximum pooling module, and a spatial median pooling module. The outputs of the three spatial pooling modules are concatenated in the channel dimension and input into the 1×1 convolution operation for tensor dimension recovery, and then normalized by the Sigmoid function. The output of the Sigmoid function is multiplied element-by-element with the input of the spatial attention mechanism SA as the output of the spatial attention mechanism SA.

6. The lung nodule risk assessment method based on multi-scale multi-receptive field multi-attention fusion according to claim 5 is characterized in that: The channel attention mechanism CA includes three parallel channel pooling modules and an expand function, wherein the three parallel channel pooling modules are a channel average pooling module, a channel maximum pooling module and a channel median pooling module, respectively; the expand function includes a fully connected layer fc1, a Relu activation function and a fully connected layer fc2 connected in sequence, and the output of the fully connected layer fc2 is dot-multiplied with the input of the channel attention mechanism CA to obtain the output of the expand function; wherein the outputs of the three channel pooling modules are respectively weighted by the expand function and fused with the input of the channel attention mechanism CA to obtain three feature tensors after weight adjustment; and the three feature tensors after weight adjustment are added together to obtain the final output of the channel attention mechanism CA.

7. The lung nodule risk assessment method based on multi-scale, multi-receptive field and multi-attention fusion according to claim 6 is characterized in that: The pixel attention mechanism PA includes a high-dimensional conversion operation, a splicing operation, a rearrangement operation, a convolution operation and a Sigmoid function.

8. The lung nodule risk assessment method based on multi-scale, multi-receptive field, and multi-attention fusion according to claim 7, characterized in that: The classification module includes an average pooling layer, a flattening operation layer and a fully connected layer. The average pooling layer is used to perform feature sorting on the key feature tensor output by the deep learning network module TMFM. The flattening operation layer is used to reshape the feature dimension of the feature tensor to conform to the input dimension of the fully connected layer. The fully connected layer is used to calculate the probability score of the lung nodule samples in the input image belonging to five lung nodule malignant risk levels, generate a classification result vector, and convert the classification result vector into a probability distribution vector. The category number of the lung nodule malignant risk level corresponding to the item with the largest probability value in the probability distribution vector is taken as the lung nodule risk assessment classification result of the lung nodule CT scan image.

Citation Information

Patent Citations

  • Improved Mask-R-CNN pulmonary nodule auxiliary detection method fusing dual-path channel attention and cavity space attention

    CN116580017A

  • Juvenile fish limb identification method based on multi-scale cascaded perceptual convolutional neural network

    US20230343128A1