Nodule identification method, device, storage medium and electronic device
By using a preset nodule recognition model to fuse global and local features of lung CT scan images and combining it with a classifier to determine the nodule category, the problem of insufficient nodule recognition accuracy in existing technologies is solved, and the accuracy and reliability of nodule recognition are improved.
Patent Information
- Application Number
- CN202310829056.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-06
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-07-06
AI Technical Summary
Existing nodule recognition methods have deficiencies in accuracy and reliability, especially in early screening of lung cancer. The accuracy of lung nodule recognition is low, which affects the reliability of lung cancer diagnosis and treatment.
A preset nodule recognition model is adopted, which fuses global features and local features of feature maps of multiple scales through multiple serially connected global and local feature fusion modules, and combines the classifier to determine the nodule category information, including nodule class and non-nodule class.
It effectively improves the accuracy and reliability of nodule category information and improves the diagnostic efficiency of early screening for lung cancer.
Smart Images

Figure CN117036247B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image recognition technology, and in particular, to a nodule recognition method, device, storage medium, and electronic device. Background Art
[0002] With the rapid development of technology, image processing and artificial intelligence technologies are increasingly being used in the medical field. On the one hand, image processing technology provides new tools for medical research, accelerating the diagnosis and treatment of diseases. It also improves the quality and clarity of medical images, providing a more accurate basis for clinical diagnosis. On the other hand, artificial intelligence technology also provides medical professionals with more efficient and precise medical services, bringing better medical experience and treatment results to patients. However, when it comes to nodule identification, existing nodule recognition methods currently suffer from low accuracy and poor reliability. Summary of the Invention
[0003] The purpose of the present disclosure is to provide a nodule identification method, device, storage medium and electronic device.
[0004] In order to achieve the above-mentioned object, the present disclosure provides a first aspect of a nodule identification method, the method comprising:
[0005] Obtaining a scanned image to be identified;
[0006] Inputting the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes nodule category and non-nodule category;
[0007] In which, the preset nodule recognition model includes a classifier and multiple global and local feature fusion modules connected in series, the output end of the last global and local feature fusion module in the multiple global and local feature fusion modules connected in series is coupled with the classifier, and the multiple global and local feature fusion modules connected in series are used to fuse global features and local features of feature maps of multiple different scales to obtain target fusion features; the classifier is used to determine the nodule category information in the scanned image to be identified based on the target fusion features.
[0008] A second aspect of the present disclosure provides a nodule identification device, the device comprising:
[0009] An acquisition module is configured to acquire a scanned image to be identified;
[0010] A determination module is configured to input the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes a nodule category and a non-nodule category;
[0011] In which, the preset nodule recognition model includes a classifier and multiple global and local feature fusion modules connected in series, the output end of the last global and local feature fusion module in the multiple global and local feature fusion modules connected in series is coupled with the classifier, and the multiple global and local feature fusion modules connected in series are used to fuse global features and local features of feature maps of multiple different scales to obtain target fusion features; the classifier is used to determine the nodule category information in the scanned image to be identified based on the target fusion features.
[0012] A third aspect of the present disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processor.
[0013] A fourth aspect of the present disclosure provides an electronic device, including:
[0014] a memory having a computer program stored thereon;
[0015] A processor is used to execute the computer program in the memory to implement the steps of the method of the first aspect.
[0016] The above technical solution obtains a scanned image to be identified; inputs the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes nodule category and non-nodule category; since the preset nodule recognition model can fuse global features and local features of feature maps of multiple scales through the multiple serially connected global and local feature fusion modules to obtain target fusion features, it can effectively fuse multi-scale global features and local features; and then determines the nodule category information in the scanned image to be identified according to the target fusion features through the classifier, which can effectively improve the accuracy and reliability of the obtained nodule category information.
[0017] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings:
[0019] Figure 1 is a flow chart of a nodule identification method shown in an exemplary embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram of a preset nodule recognition model shown in an exemplary embodiment of the present disclosure;
[0021] Figure 3 is based on Figure 2 A schematic diagram of a preset nodule recognition model shown in the illustrated embodiment;
[0022] Figure 4 is based on Figure 3 A schematic diagram of a preset nodule recognition model shown in the illustrated embodiment;
[0023] Figure 5 is based on Figure 2 A schematic diagram of a preset nodule recognition model shown in the illustrated embodiment;
[0024] Figure 6 is based on Figure 1 The illustrated embodiment shows a flow chart of a nodule identification method;
[0025] Figure 7 is a flowchart of a model training method shown in an exemplary embodiment of the present disclosure;
[0026] Figure 8 is a block diagram of a nodule identification device according to an exemplary embodiment of the present disclosure;
[0027] Figure 9 is a block diagram of an electronic device according to an exemplary embodiment;
[0028] Figure 10 is a block diagram of another electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0029] The following describes the specific embodiments of the present disclosure in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure and are not intended to limit the present disclosure.
[0030] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0031] Before introducing the specific embodiments of the present disclosure in detail, the application scenarios of the present disclosure are first described as follows. The present disclosure can be applied to the identification of nodules, which may be lung nodules, lymph nodules, or nodules in other parts of the body. Lung cancer is one of the main diseases that lead to high cancer-related mortality rates worldwide. During the CT (Computed Tomography) scanning and analysis process for screening lung cancer, the suspected lung nodules detected may be early lung cancer or normal anatomical structures of the lungs. Using CT images to determine whether it is a lung nodule or normal tissue (non-nodule) suspected of being a lung nodule has important auxiliary significance for the diagnosis of lung cancer. Lung CT scans are widely used in the early screening of lung cancer, but due to the low accuracy of lung nodule recognition, their reliability has been challenged, and thus they cannot provide reliable data basis for the diagnosis and treatment of lung cancer.
[0032] In order to solve the above technical problems, the present disclosure provides a nodule recognition method, device, storage medium and electronic device. The nodule recognition method obtains a scanned image to be identified; inputs the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, and the nodule category information includes nodule class and non-nodule class; because the preset nodule recognition model can fuse global features and local features of multiple feature maps of different scales through the multiple serially connected global and local feature fusion modules to obtain target fusion features, it can effectively fuse multi-scale global features and local features; then the classifier is used to determine the nodule category information in the scanned image to be identified according to the target fusion features, which can effectively improve the accuracy and reliability of the obtained nodule category information.
[0033] Figure 1 is a flowchart of a nodule identification method shown in an exemplary embodiment of the present disclosure; Figure 1 As shown, the nodule identification method may include:
[0034] Step 101: Obtain a scanned image to be identified.
[0035] The scanned image to be identified may be a CT image.
[0036] Step 102: input the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes nodule category and non-nodule category.
[0037] In which, the preset nodule recognition model includes a classifier and multiple global and local feature fusion modules connected in series, the output end of the last global and local feature fusion module in the multiple global and local feature fusion modules connected in series is coupled with the classifier, and the multiple global and local feature fusion modules connected in series are used to fuse global features and local features of feature maps of multiple different scales to obtain target fusion features; the classifier is used to determine the nodule category information in the scanned image to be identified based on the target fusion features.
[0038] For example, Figure 2 is a schematic diagram of a preset nodule recognition model shown in an exemplary embodiment of the present disclosure; Figure 2 As shown, the preset nodule recognition model may include four global and local feature fusion modules connected in series, namely global and local feature fusion module 1, global and local feature fusion module 2, global and local feature fusion module 3 and global and local feature fusion module 4. Each of the global and local feature fusion modules includes at least a first convolution layer, a second convolution layer, a global and local feature fusion layer, a third convolution layer and a downsampling layer. The input end of the first convolution layer serves as the input end of the global and local feature fusion module; the output end of the first convolution layer is coupled with the input end of the second convolution layer, the output end of the second convolution layer is coupled with the input end of the global and local feature fusion layer, the output end of the second convolution layer is also coupled with the input end of the third convolution layer, the input end of the third convolution layer is also coupled with the output end of the global and local feature fusion layer, the output end of the third convolution layer is coupled with the input end of the downsampling layer, and the output end of the downsampling serves as the output end of the global and local feature fusion module;
[0039] The global and local feature fusion layer is used to perform global feature fusion on the standby feature map output by the second convolutional layer to obtain a target global feature map, perform local feature fusion on the standby feature map to obtain a target local feature map, and fuse the target global feature map and the target local feature map to obtain a first standby fused feature map;
[0040] The third convolutional layer is used to perform convolution processing on the first to-be-used fused feature map and the to-be-used feature map to obtain a second to-be-used fused feature map;
[0041] The downsampling layer is used to perform downsampling processing on the second to-be-used fused feature map to obtain a target to-be-used fused feature map.
[0042] The global and local feature fusion layer may include a local feature fusion submodule, a global feature fusion submodule and a feature fusion submodule. The structure of the global and local feature fusion layer may be as follows: Figure 3As shown, Figure 3 is based on Figure 2 The illustrated embodiment shows a schematic diagram of a preset nodule recognition model; the input end of the local feature fusion submodule, the input end of the global feature fusion submodule and the input end of the feature fusion submodule are coupled with the input end of the global and local feature fusion layer, and the input end of the feature fusion submodule is also coupled with the output end of the local feature fusion submodule and the output end of the global feature fusion submodule, and the output end of the feature fusion submodule serves as the output end of the global and local feature fusion layer.
[0043] The local feature fusion submodule is used to divide the stand-by feature map into multiple local blocks, perform attention convolution fusion processing on each local block to obtain a first local feature map corresponding to each local block, and splice multiple first local feature maps corresponding to the multiple local blocks to form the target local feature map.
[0044] The global feature fusion submodule is used to perform dimensionality reduction processing on the standby feature map to obtain an intermediate feature map after dimensionality reduction processing, and perform the attention convolution fusion processing on the intermediate feature map to obtain the target global feature map;
[0045] The feature fusion submodule is used to generate the first standby fused feature map according to the target global feature map, the target local feature map and the standby feature map.
[0046] It should be noted that the attention convolution fusion processing may include using the attention mechanism to perform feature extraction on the features to be processed to obtain target attention features, performing convolution processing on the features to be processed to obtain target intermediate features, fusing the features to be processed, the target attention features and the target intermediate features to obtain spatial attention features, inputting the spatial attention features into a preset pooling network and a preset activation network to obtain channel attention weights, and determining the processed feature map based on the channel attention weights and the spatial attention features.
[0047] For example, the attention convolution fusion process can be performed by Figure 3 The attention convolution feature fusion network shown in is performed. The structural diagram of the attention convolution feature fusion network is as follows Figure 4 As shown, Figure 4 is based on Figure 3The embodiment shown is a schematic diagram of a preset nodule recognition model; the attention convolution feature fusion network is mainly composed of two parts connected in series: a spatial attention sub-network and a channel attention sub-network. The spatial attention sub-network uses the attention mechanism to extract features from the features to be processed to obtain target attention features, performs convolution processing on the features to be processed to obtain target intermediate features, and fuses the features to be processed, the target attention features and the target intermediate features to obtain spatial attention features; the channel attention sub-network is used to input the spatial attention features into a preset pooling network and a preset activation network to obtain channel attention weights, and determines the processed feature map according to the channel attention weights and the spatial attention features; in this way, the spatial dimension and channel dimension of the input feature map are feature-fused by the connected spatial attention sub-network and the channel attention sub-network to finally obtain the output feature map. For example, the feature map to be used is The unused feature map is input into the local feature fusion submodule, which first divides the unused feature map into non-overlapping local blocks of size h×w
[0048]
[0049] Among them, c, h, w, i, and j are all positive integers, h is an integer less than H, w is an integer less than W, C is the number of channels, H is the height of the feature map, and W is the width of the feature map.
[0050] In the spatial attention sub-network, if the input feature map (which can be each local block) is right Perform self-attention calculation to obtain the target attention feature X selfattention ,right The convolution is performed to obtain the target intermediate feature X conv , using skip connections to Directly pass it to the output of the spatial attention sub-network, and then pass the target attention feature X selfattention , target intermediate feature X conv , features to be processed Add up to get the output spatial attention features of the spatial attention sub-network Right now:
[0051]
[0052]
[0053]
[0054] It should be pointed out that softmax and Relu are both activation functions, Q is the query vector in the attention mechanism, K is the key vector in the attention mechanism, V is the value vector in the attention mechanism, BN represents batch normalization, and Conv represents convolution.
[0055] In the channel attention sub-network, the spatial attention features are As the input of the channel attention sub-network (i.e. ),Will Input the preset pooling network and the preset activation network to obtain the channel attention weight, and determine the processed feature map according to the channel attention weight and the spatial attention feature; for example, first Through the average pooling layer and then through the Sigmoid function, the channel attention weight is obtained Then the channel attention weight is combined with the point-by-point convolution Multiply to get Right now:
[0056]
[0057] Among them, PWConv is point-by-point convolution processing, mean H,W It is an average pooling process. After the attention convolution feature fusion network is used to fusion the local blocks After extracting features, multiple first local feature maps corresponding to different local blocks can be obtained The target local feature map is formed by splicing the multiple first local feature maps. Combined into the output feature map of the local feature fusion submodule
[0058] In addition, the unused feature map is input into the global feature fusion submodule, which first reduces the dimensionality of the unused feature map through point-by-point convolution to reduce the computational complexity of downstream tasks. For example, the input feature map dimension can be reduced to half the original number of channels, that is, C′=1 / 2C, and then the intermediate feature map after dimensionality reduction is subjected to attention convolution fusion processing. The attention convolution processing process can refer to the relevant description of the attention convolution fusion processing in the local feature fusion submodule above, and will not be repeated in this disclosure.
[0059] Furthermore, the feature fusion submodule is used to perform channel-dimensional splicing on the target local feature map and the target global feature map to obtain a first spliced feature map, perform feature extraction on the first spliced feature map to obtain a first extracted feature, perform channel-dimensional splicing on the first extracted feature and the stand-by feature map to obtain a second spliced feature map, and perform feature extraction on the second spliced feature map to obtain the first stand-by fused feature map.
[0060] In one embodiment, the feature fusion submodule may include a fusion layer A and a fusion layer B. The fusion layer A performs fusion calculation on the target local feature map and the target global feature map. The fusion layer B further fuses the output result of the fusion layer A with a jump connection to obtain the first standby fusion feature map. For example, if the target local feature map is The target global feature map is The target local feature map and the target global feature map can be spliced together along the channel dimension to obtain Will After convolution processing, batch normalization processing and Relu activation, the output feature map of the fusion layer A is obtained Fusion layer B pair Connect with JumpX res Splicing along the channel dimension, we get Will After convolution processing, batch normalization processing and Relu activation, the output feature map of the fusion layer B is obtained During the convolution process, The global and local feature information in the
[15] is fused with the information of the original input feature map in the jump connection, and the number of input feature map channels is reduced to C.
[0061] It should also be noted that the downsampling layer may include a pooling layer, for example, it may be composed of an average pooling layer with a step size of 2; the classifier includes 3 fully connected layers, 2 random inactivation layers and 1 softmax layer, and the specific structure can be as follows Figure 5 As shown, Figure 5 is based on Figure 2 The illustrated embodiment shows a schematic diagram of a preset nodule recognition model, in which fully connected layers and random dropout layers are alternately arranged, and a softmax layer is connected after the last fully connected layer. In this way, random dropout layers are applied after the first two fully connected layers to reduce model overfitting. Finally, the softmax layer calculates the nodule probability obtained by the final classification of the input image, and determines whether it belongs to the nodule class or non-nodule class based on the node probability. For example, the softmax layer outputs a probability of belonging to the node class of 89%, and the preset probability threshold is 85%. Since 89% is greater than the preset probability threshold, it can be determined that it belongs to the node class.
[0062] The above technical solution obtains a scanned image to be identified; inputs the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model. Since the preset nodule recognition model can fuse global features and local features of feature maps of multiple scales through the multiple serially connected global and local feature fusion modules to obtain target fusion features, it can effectively fuse multi-scale global features and local features; and then determines the nodule category information in the scanned image to be identified according to the target fusion features through the classifier, which can effectively improve the accuracy and reliability of the obtained nodule category information.
[0063] Figure 6 is based on Figure 1 The embodiment shown is a flowchart of a nodule identification method; Figure 6 As shown, Figure 1 Step 102 may include:
[0064] Step 1021: input the scanned image to be identified into a first global and local feature fusion module to obtain the target fused feature map to be used at a first scale output by the first global and local feature fusion module.
[0065] The structure of the preset nodule recognition model can be as follows: Figure 2 As shown, the present disclosure will not be repeated here.
[0066] Step 1022: Input the target unused fused feature map of the first scale into a second global and local feature fusion module to obtain the target unused fused feature map of the second scale output by the second global and local feature fusion module.
[0067] Step 1023: Determine the target fusion feature according to the target fusion feature map at the second scale.
[0068] One possible implementation of this step is: inputting the target fusion feature map of the second scale into the third global and local feature fusion module to obtain the target stand-by fusion feature map of the third scale output by the third global and local feature fusion module, and determining the target fusion feature according to the target fusion feature map of the third scale.
[0069] It should be noted that if the preset nodule recognition model includes three global and local feature fusion modules, the above-mentioned implementation method of determining the target fusion feature based on the target fusion feature map of the third scale may be to determine the target fusion feature based on the target fusion feature map of the third scale; if the preset nodule recognition model includes more than three global and local feature fusion modules, after obtaining the target fusion feature map of the third scale, it is also necessary to obtain the target fusion feature map of the fourth scale through the fourth global and local feature fusion module in sequence.
[0070] Another possible implementation is: using the target fusion feature map of the second scale as the target fusion feature.
[0071] Step 1024 : Input the target fusion feature into the classifier to obtain the nodule category information corresponding to each nodule region in the scanned image to be identified.
[0072] It should be noted that Figure 2 Taking the preset nodule recognition model shown in as an example, the scales of the feature maps of different scales corresponding to the global and local feature fusion module 1, the global and local feature fusion module 2, the global and local feature fusion module 3, and the global and local feature fusion module 4 can be shown in Table 1:
[0073] Table 1
[0074] Module Input scale Output scale Global and local feature fusion module 1 3×H×W 64×1 / 2H×1 / 2W Global and local feature fusion module 2 64×1 / 2H×1 / 2W 128×1 / 4H×1 / 4W Global and local feature fusion module 3 128×1 / 4H×1 / 4W 256×1 / 8H×1 / 8W Global and local feature fusion module 4 256×1 / 8H×1 / 8W 512×1 / 16H×1 / 16W
[0075] In Table 1 above, H is the height of the feature map, and W is the width of the feature map.
[0076] The above technical solution can fuse global features and local features of feature maps of multiple scales through the multiple serially connected global and local feature fusion modules to obtain target fusion features, which can effectively fuse multi-scale global features and local features; and then determine the nodule category information in the scanned image to be identified based on the target fusion features through the classifier, which can effectively improve the accuracy and reliability of the obtained nodule category information.
[0077] Figure 7 is a flow chart of a model training method shown in an exemplary embodiment of the present disclosure; Figure 7 As shown, Figure 1 The training process of the preset nodule recognition model may include:
[0078] S1, obtaining multiple scanned sample images and the nodule category of each tissue region in each scanned sample image.
[0079] The scanned sample image may be a lung CT image sample, or a CT image or ultrasound image of another part of the body. The tissue region may be an image region in the scanned sample image where tissue greater than or equal to 3 mm is located.
[0080] It should be noted that, in order to meet the network training data requirements, the original collected data can be flipped horizontally and / or vertically to generate more tissue areas suspected of nodules to meet the data volume requirements during the training process.
[0081] S2, taking the nodule category of each tissue region as label data of the tissue region.
[0082] S3, using the multiple scan sample images and the label data as training data, iteratively training the preset initial model to obtain the preset nodule recognition model.
[0083] The preset initial model includes an initial classifier and multiple initial global and local feature fusion modules connected in series. The model structure of the preset initial model can be Figures 2 to 5 The model structure shown in any one of the above will not be described in detail in this disclosure.
[0084] The above technical solution uses the nodule category of each tissue area as the label data of the tissue area, and the multiple scanned sample images and the label data as training data to iteratively train the preset initial model to obtain the preset nodule recognition model, which can effectively train the preset nodule recognition model with higher recognition accuracy.
[0085] Figure 8 is a block diagram of a nodule identification device shown in an exemplary embodiment of the present disclosure. Figure 8 As shown, the device may include:
[0086] An acquisition module 801 is configured to acquire a scanned image to be identified;
[0087] The determination module 802 is configured to input the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes a nodule category and a non-nodule category;
[0088] In which, the preset nodule recognition model includes a classifier and multiple global and local feature fusion modules connected in series, the output end of the last global and local feature fusion module in the multiple global and local feature fusion modules connected in series is coupled with the classifier, and the multiple global and local feature fusion modules connected in series are used to fuse global features and local features of feature maps of multiple different scales to obtain target fusion features; the classifier is used to determine the nodule category information in the scanned image to be identified based on the target fusion features.
[0089] The above technical solution obtains a scanned image to be identified; inputs the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes nodule category and non-nodule category; since the preset nodule recognition model can fuse global features and local features of feature maps of multiple scales through the multiple serially connected global and local feature fusion modules to obtain target fusion features, it can effectively fuse multi-scale global features and local features; and then determines the nodule category information in the scanned image to be identified according to the target fusion features through the classifier, which can effectively improve the accuracy and reliability of the obtained nodule category information.
[0090] Optionally, the global and local feature fusion module includes at least a first convolution layer, a second convolution layer, a global and local feature fusion layer, a third convolution layer and a downsampling layer, the input end of the first convolution layer serves as the input end of the global and local feature fusion module; the output end of the first convolution layer is coupled with the input end of the second convolution layer, the output end of the second convolution layer is coupled with the input end of the global and local feature fusion layer, the output end of the second convolution layer is also coupled with the input end of the third convolution layer, the input end of the third convolution layer is also coupled with the output end of the global and local feature fusion layer, the output end of the third convolution layer is coupled with the input end of the downsampling layer, and the output end of the downsampling serves as the output end of the global and local feature fusion module;
[0091] The global and local feature fusion layer is used to perform global feature fusion on the standby feature map output by the second convolutional layer to obtain a target global feature map, perform local feature fusion on the standby feature map to obtain a target local feature map, and fuse the target global feature map and the target local feature map to obtain a first standby fused feature map;
[0092] The third convolutional layer is used to perform convolution processing on the first to-be-used fused feature map and the to-be-used feature map to obtain a second to-be-used fused feature map;
[0093] The downsampling layer is used to perform downsampling processing on the second to-be-used fused feature map to obtain a target to-be-used fused feature map.
[0094] Optionally, the global and local feature fusion layer includes a local feature fusion submodule, a global feature fusion submodule and a feature fusion submodule. The input end of the local feature fusion submodule, the input end of the global feature fusion submodule and the input end of the feature fusion submodule are coupled to the input end of the global and local feature fusion layer. The input end of the feature fusion submodule is also coupled to the output end of the local feature fusion submodule and the output end of the global feature fusion submodule. The output end of the feature fusion submodule serves as the output end of the global and local feature fusion layer.
[0095] The local feature fusion submodule is used to divide the unused feature map into multiple local blocks, perform attention convolution fusion processing on each local block to obtain a first local feature map corresponding to each local block, and splice multiple first local feature maps corresponding to the multiple local blocks to form the target local feature map; wherein, the attention convolution fusion processing includes using the attention mechanism to extract features of the to-be-processed features to obtain target attention features, performing convolution processing on the to-be-processed features to obtain target intermediate features, fusing the to-be-processed features, the target attention features and the target intermediate features to obtain spatial attention features, inputting the spatial attention features into a preset pooling network and a preset activation network to obtain channel attention weights, and determining the processed feature map according to the channel attention weights and the spatial attention features;
[0096] The global feature fusion submodule is used to perform dimensionality reduction processing on the standby feature map to obtain an intermediate feature map after dimensionality reduction processing, and perform the attention convolution fusion processing on the intermediate feature map to obtain the target global feature map;
[0097] The feature fusion submodule is used to generate the first standby fused feature map according to the target global feature map, the target local feature map and the standby feature map.
[0098] Optionally, the feature fusion submodule is used to perform channel-dimensional splicing on the target local feature map and the target global feature map to obtain a first spliced feature map, perform feature extraction on the first spliced feature map to obtain a first extracted feature, perform channel-dimensional splicing on the first extracted feature and the stand-by feature map to obtain a second spliced feature map, and perform feature extraction on the second spliced feature map to obtain the first stand-by fused feature map.
[0099] Optionally, the determining module 802 is configured to:
[0100] Inputting the scanned image to be identified into a first global and local feature fusion module to obtain the target fused feature map to be used at a first scale output by the first global and local feature fusion module;
[0101] Inputting the target unused fused feature map of the first scale into a second global and local feature fusion module to obtain the target unused fused feature map of the second scale output by the second global and local feature fusion module;
[0102] Determining the target fusion feature according to the target fusion feature map of the second scale;
[0103] The target fusion feature is input into the classifier to obtain the nodule category information corresponding to each nodule area in the scanned image to be identified.
[0104] Optionally, the determining module 802 is configured to:
[0105] Inputting the target fusion feature map of the second scale into a third global and local feature fusion module to obtain the target standby fusion feature map of the third scale output by the third global and local feature fusion module, and determining the target fusion feature according to the target fusion feature map of the third scale; or
[0106] The target fusion feature map at the second scale is used as the target fusion feature.
[0107] Optionally, the device further includes a model training module, and the model training module is configured to:
[0108] acquiring a plurality of scanned sample images and a nodule category of each tissue region in each scanned sample image;
[0109] The nodule category of each tissue region is used as label data of the tissue region;
[0110] The preset initial model is iteratively trained using the multiple scan sample images and the label data as training data to obtain the preset nodule recognition model; wherein the preset initial model includes an initial classifier and multiple serially connected initial global and local feature fusion modules.
[0111] The above technical solution uses the nodule category of each tissue area as the label data of the tissue area, and the multiple scanned sample images and the label data as training data to iteratively train the preset initial model to obtain the preset nodule recognition model, which can effectively train the preset nodule recognition model with higher recognition accuracy.
[0112] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0113] Figure 9 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Figure 9 As shown, the first electronic device 900 may include: a first processor 901 and a first memory 902. The first electronic device 900 may also include one or more of a multimedia component 903, a first input / output interface 904, and a first communication component 905.
[0114] The first processor 901 is used to control the overall operation of the first electronic device 900 to complete all or part of the steps in the above-mentioned nodule identification method. The first memory 902 is used to store various types of data to support the operation of the first electronic device 900. For example, this data may include instructions for any application or method operating on the first electronic device 900, as well as application-related data, such as contact information, sent and received messages, pictures, audio, video, etc. The first memory 902 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 903 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the first memory 902 or transmitted via the first communication component 905. The audio component also includes at least one speaker for outputting audio signals. The first input / output interface 904 provides an interface between the first processor 901 and other interface modules. The above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The first communication component 905 is used for wired or wireless communication between the first electronic device 900 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more thereof, is not limited here. Therefore, the corresponding first communication component 905 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0115] In an exemplary embodiment, the first electronic device 900 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above-mentioned nodule identification method.
[0116] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described nodule identification method. For example, the computer-readable storage medium may be the above-described first memory 902 including the program instructions. The above-described program instructions may be executed by the first processor 901 of the first electronic device 900 to perform the above-described nodule identification method.
[0117] Figure 10 1 is a block diagram of another electronic device according to an exemplary embodiment. For example, the second electronic device 1000 can be provided as a server. Figure 10 The second electronic device 1000 includes a second processor 1022, which may be one or more, and a second memory 1032 for storing a computer program executable by the second processor 1022. The computer program stored in the second memory 1032 may include one or more modules, each corresponding to a set of instructions. In addition, the second processor 1022 may be configured to execute the computer program to perform the above-mentioned nodule identification method.
[0118] In addition, the second electronic device 1000 may further include a power supply component 1026 and a second communication component 1050. The power supply component 1026 may be configured to perform power management of the second electronic device 1000, and the second communication component 1050 may be configured to implement communication, for example, wired or wireless communication, of the second electronic device 1000. In addition, the second electronic device 1000 may further include a second input / output interface 1058. The second electronic device 1000 may operate based on an operating system stored in the second memory 1032.
[0119] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described nodule identification method. For example, the non-transitory computer-readable storage medium may be the second memory 1032 including the program instructions. The program instructions may be executed by the second processor 1022 of the second electronic device 1000 to perform the above-described nodule identification method.
[0120] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and the computer program has a code portion for performing the above-mentioned nodule identification method when executed by the programmable device.
[0121] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.
[0122] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.
[0123] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.
Claims
1. A nodule identification method, characterized in that: The method comprises: Obtaining a scanned image to be identified; Inputting the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes nodule category and non-nodule category; The preset nodule recognition model includes a classifier and multiple global and local feature fusion modules connected in series, wherein the output end of the last global and local feature fusion module in the multiple global and local feature fusion modules connected in series is coupled to the classifier, and the multiple global and local feature fusion modules connected in series are used to fuse global features and local features of feature maps of multiple different scales to obtain target fusion features; the classifier is used to determine the nodule category information in the scanned image to be identified based on the target fusion features; The global and local feature fusion module includes at least a first convolution layer, a second convolution layer, a global and local feature fusion layer, a third convolution layer and a downsampling layer. The input end of the first convolution layer serves as the input end of the global and local feature fusion module; the output end of the first convolution layer is coupled with the input end of the second convolution layer, the output end of the second convolution layer is coupled with the input end of the global and local feature fusion layer, the output end of the second convolution layer is also coupled with the input end of the third convolution layer, the input end of the third convolution layer is also coupled with the output end of the global and local feature fusion layer, the output end of the third convolution layer is coupled with the input end of the downsampling layer, and the output end of the downsampling serves as the output end of the global and local feature fusion module; The global and local feature fusion layer is used to perform global feature fusion on the standby feature map output by the second convolutional layer to obtain a target global feature map, perform local feature fusion on the standby feature map to obtain a target local feature map, and fuse the target global feature map and the target local feature map to obtain a first standby fused feature map; The third convolutional layer is used to perform convolution processing on the first to-be-used fused feature map and the to-be-used feature map to obtain a second to-be-used fused feature map; The downsampling layer is used to perform downsampling processing on the second to-be-used fused feature map to obtain a target to-be-used fused feature map; The global and local feature fusion layer includes a local feature fusion submodule, a global feature fusion submodule and a feature fusion submodule. The input end of the local feature fusion submodule, the input end of the global feature fusion submodule and the input end of the feature fusion submodule are coupled to the input end of the global and local feature fusion layer. The input end of the feature fusion submodule is also coupled to the output end of the local feature fusion submodule and the output end of the global feature fusion submodule. The output end of the feature fusion submodule serves as the output end of the global and local feature fusion layer. The local feature fusion submodule is used to divide the unused feature map into multiple local blocks, perform attention convolution fusion processing on each local block to obtain a first local feature map corresponding to each local block, and splice multiple first local feature maps corresponding to the multiple local blocks to form the target local feature map; wherein, the attention convolution fusion processing includes using the attention mechanism to extract features of the to-be-processed features to obtain target attention features, performing convolution processing on the to-be-processed features to obtain target intermediate features, fusing the to-be-processed features, the target attention features and the target intermediate features to obtain spatial attention features, inputting the spatial attention features into a preset pooling network and a preset activation network to obtain channel attention weights, and determining the processed feature map according to the channel attention weights and the spatial attention features; The global feature fusion submodule is used to perform dimensionality reduction processing on the standby feature map to obtain an intermediate feature map after dimensionality reduction processing, and perform the attention convolution fusion processing on the intermediate feature map to obtain the target global feature map; The feature fusion submodule is used to generate the first standby fused feature map according to the target global feature map, the target local feature map and the standby feature map; The feature fusion submodule is configured to perform channel-dimensional splicing on the target local feature map and the target global feature map to obtain a first spliced feature map, perform feature extraction on the first spliced feature map to obtain a first extracted feature, perform channel-dimensional splicing on the first extracted feature and the standby feature map to obtain a second spliced feature map, and perform feature extraction on the second spliced feature map to obtain the first standby fused feature map; Inputting the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model includes: Inputting the scanned image to be identified into a first global and local feature fusion module to obtain a target fused feature map of a first scale output by the first global and local feature fusion module; Inputting the target unused fused feature map of the first scale into a second global and local feature fusion module to obtain the target unused fused feature map of the second scale output by the second global and local feature fusion module; Determining the target fusion feature according to the target fusion feature map of the second scale; The target fusion feature is input into the classifier to obtain the nodule category information corresponding to each nodule area in the scanned image to be identified.
2. The method according to claim 1, characterized in that The determining the target fusion feature according to the target fusion feature map at the second scale includes: Inputting the target fusion feature map of the second scale into a third global and local feature fusion module to obtain the target standby fusion feature map of the third scale output by the third global and local feature fusion module, and determining the target fusion feature according to the target fusion feature map of the third scale; or The target fusion feature map at the second scale is used as the target fusion feature.
3. The method according to claim 1, characterized in that The training process of the preset nodule recognition model includes: acquiring a plurality of scanned sample images and a nodule category of each tissue region in each scanned sample image; The nodule category of each tissue region is used as label data of the tissue region; The preset initial model is iteratively trained using the multiple scan sample images and the label data as training data to obtain the preset nodule recognition model; wherein the preset initial model includes an initial classifier and multiple serially connected initial global and local feature fusion modules.
4. A nodule identification device, characterized in that: The device comprises: An acquisition module is configured to acquire a scanned image to be identified; A determination module is configured to input the scanned image to be identified into a preset nodule recognition model to obtain nodule category information output by the preset nodule recognition model, wherein the nodule category information includes a nodule category and a non-nodule category; The preset nodule recognition model includes a classifier and multiple global and local feature fusion modules connected in series, wherein the output end of the last global and local feature fusion module in the multiple global and local feature fusion modules connected in series is coupled to the classifier, and the multiple global and local feature fusion modules connected in series are used to fuse global features and local features of feature maps of multiple different scales to obtain target fusion features; the classifier is used to determine the nodule category information in the scanned image to be identified based on the target fusion features; The global and local feature fusion module at least includes a first convolution layer, a second convolution layer, a global and local feature fusion layer, a third convolution layer and a downsampling layer. The input end of the first convolution layer serves as the input end of the global and local feature fusion module; the output end of the first convolution layer is coupled with the input end of the second convolution layer, the output end of the second convolution layer is coupled with the input end of the global and local feature fusion layer, the output end of the second convolution layer is also coupled with the input end of the third convolution layer, the input end of the third convolution layer is also coupled with the output end of the global and local feature fusion layer, the output end of the third convolution layer is coupled with the input end of the downsampling layer, and the output end of the downsampling layer is coupled. The output end serves as the output end of the global and local feature fusion module; the global and local feature fusion layer is used to perform global feature fusion on the standby feature map output by the second convolution layer to obtain a target global feature map, and perform local feature fusion on the standby feature map to obtain a target local feature map, and fuse the target global feature map and the target local feature map to obtain a first standby fused feature map; the third convolution layer is used to perform convolution processing on the first standby fused feature map and the standby feature map to obtain a second standby fused feature map; the downsampling layer is used to perform downsampling processing on the second standby fused feature map to obtain a target standby fused feature map; The global and local feature fusion layer includes a local feature fusion submodule, a global feature fusion submodule and a feature fusion submodule. The input end of the local feature fusion submodule, the input end of the global feature fusion submodule and the input end of the feature fusion submodule are coupled with the input end of the global and local feature fusion layer. The input end of the feature fusion submodule is also coupled with the output end of the local feature fusion submodule and the output end of the global feature fusion submodule. The output end of the feature fusion submodule serves as the output end of the global and local feature fusion layer. The local feature fusion submodule is used to divide the unused feature map into multiple local blocks, perform attention convolution fusion processing on each local block to obtain a first local feature map corresponding to each local block, and splice the multiple first local feature maps corresponding to the multiple local blocks to form the target local feature map. The product fusion processing includes using the attention mechanism to extract features from the features to be processed to obtain target attention features, performing convolution processing on the features to be processed to obtain target intermediate features, fusing the features to be processed, the target attention features and the target intermediate features to obtain spatial attention features, inputting the spatial attention features into a preset pooling network and a preset activation network to obtain channel attention weights, and determining a processed feature map based on the channel attention weights and the spatial attention features; the global feature fusion submodule is used to perform dimensionality reduction processing on the feature map to be used to obtain an intermediate feature map after dimensionality reduction processing, and performing the attention convolution fusion processing on the intermediate feature map to obtain the target global feature map; the feature fusion submodule is used to generate the first standby fused feature map based on the target global feature map, the target local feature map and the standby feature map; The feature fusion submodule is configured to perform channel-dimensional splicing on the target local feature map and the target global feature map to obtain a first spliced feature map, perform feature extraction on the first spliced feature map to obtain a first extracted feature, perform channel-dimensional splicing on the first extracted feature and the standby feature map to obtain a second spliced feature map, and perform feature extraction on the second spliced feature map to obtain the first standby fused feature map; The determination module is configured to: input the scanned image to be identified into a first global and local feature fusion module to obtain the target fused feature map for use at a first scale output by the first global and local feature fusion module; input the target fused feature map for use at the first scale into a second global and local feature fusion module to obtain the target fused feature map for use at a second scale output by the second global and local feature fusion module; determine the target fusion feature based on the target fusion feature map at the second scale; and input the target fusion feature into the classifier to obtain the nodule category information corresponding to each nodule area in the scanned image to be identified.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.
6. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Medical image-based nodule detection method, device, and electronic equipment
CN111798424A
Image category judgment method and device and electronic equipment
CN115187579A