Chest X-ray Image Recognition Method, Device, Computer Equipment and Storage Medium
By designing a chest X-ray image recognition network combining convolution and Transformer's ConvTransBlock module, the problem of large amounts of parameters and slow convergence speed of existing models is solved, and the lesion information in chest X-ray images is efficiently recognized.
Patent Information
- Application Number
- CN202111542388.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-14
AI Technical Summary
The existing chest X-ray image recognition model has large parameters and slow convergence speed, making it difficult to efficiently identify lesion information in chest X-ray images.
A chest X-ray image recognition network including stem module, feature extraction network and classification network was designed. The feature extraction network adopts the ConvTransBlock module, combining convolutional branches and Transformer branches, and features fusion through the FCU module to reduce the amount of parameters and speed up the convergence speed.
The model parameters are small, the convergence speed is fast, and the chest X-ray image recognition accuracy is improved, especially when identifying chest X-ray images of pneumonia patients, it has high recognition ability.
Smart Images

Figure CN115249228B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and particularly to a method, an apparatus, a computer device, and a storage medium for identifying chest X-ray images. Background Art
[0002] The chest cavity is metaphorically referred to as a mirror of human health and diseases, because it contains many important tissue structures of the human body and can provide multi-faceted information of the human body. For example, the diagnosis of lung diseases, rib fractures and injuries, symptoms of heart enlargement, and cardio-pulmonary coefficients can all be identified and confirmed by chest X-ray images. Chest X-ray is still a routine examination item for hospital imaging diagnosis. Even for cases that require CT and MRI examinations, the manifestations of chest X-ray examinations are often needed for reference. Chest X-ray films account for more than 40% of all radiological imaging diagnoses due to their low price and low radiation dose. This reflects the important application value of chest X-ray films in the medical field. However, how to analyze chest X-ray films is a very challenging task, and even experienced experts often find it difficult. In recent years, deep learning has become one of the popular research fields in artificial intelligence. Deep Convolutional Neural Network (DCNN) has excellent performance in computer vision tasks such as image recognition, image segmentation, and object detection, and is widely used in the field of medical image processing to solve complex computer vision problems in the medical image field.
[0003] Chest X-ray images have high similarity between classes and low variability within classes, which increases the difficulty for network models to identify lesion information in chest X-ray images. After the outbreak of the COVID-19 pandemic, using DCNN to detect COVID-19 has become a recognized research direction, and many excellent research results have emerged in this direction. However, in the existing research results, the number of parameters of the model increases with the improvement of the classification accuracy, and the network convergence speed is low. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, an apparatus, a computer device, and a storage medium for identifying chest X-ray images to solve the above technical problems.
[0005] A method for identifying chest X-ray images, the method comprising:
[0006] Obtain a plurality of labeled chest X-ray images, and preprocess the chest X-ray images to obtain training samples.
[0007] Construct a chest X-ray image recognition network; the chest X-ray image recognition network includes a stem module, a feature extraction network, and a classification network; the Stem module is used to extract the initial local features of the training samples through convolution and max pooling operations; the feature extraction network includes a number of consecutive ConvTransBlock modules, and the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module for bridging the two branches; the convolutional branch is used to extract the local features of the input features using an inverted residual structure composed of convolutional layers; the Transformer branch is used to extract the complex spatial transformation and long-range feature dependencies of the input features to obtain global features; the FCU module is used to fuse the local features of the convolutional branch with the global features of the Transformer branch; the classification network is used to perform image recognition based on the output features of the convolutional branch and the Transformer branch of the last ConvTransBlock module respectively, and fuse the two classification prediction results obtained to obtain the sample prediction classification result.
[0008] Train the chest X-ray image recognition network according to the identification of the training samples and the sample prediction classification results obtained by inputting the training samples into the chest X-ray image recognition network to obtain a trained chest X-ray image recognition network.
[0009] Input the chest X-ray image to be measured into the trained chest X-ray image recognition network to obtain the category of the chest X-ray image.
[0010] A chest X-ray image recognition device, the device includes:
[0011] A training sample determination module, configured to obtain a number of labeled chest X-ray images and preprocess the chest X-ray images to obtain training samples.
[0012] A chest X-ray image recognition network construction module for constructing a chest X-ray image recognition network; the chest X-ray image recognition network includes a stem module, a feature extraction network, and a classification network; the Stem module is used to extract initial local features of training samples through convolution and max pooling operations; the feature extraction network includes several consecutive ConvTransBlock modules; the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module for bridging the two branches; the convolutional branch is used to extract local features of input features using an inverse residual structure composed of convolutional layers; the Transformer branch is used to extract complex spatial transformations and long-range feature dependencies of input features to obtain global features; the FCU module is used to fuse the local features of the convolutional branch with the global features of the Transformer branch; the classification network is used to perform image recognition based on the output features of the convolutional branch and the Transformer branch of the last ConvTransBlock module respectively, and fuse the two classification prediction results obtained to obtain a sample prediction classification result.
[0013] A chest X-ray image recognition network training module for training the chest X-ray image recognition network according to the identification of the training samples and the sample prediction classification results obtained by inputting the training samples into the chest X-ray image recognition network, to obtain a trained chest X-ray image recognition network.
[0014] A chest X-ray image category determination module for inputting a chest X-ray image to be measured into the trained chest X-ray image recognition network to obtain the category of the chest X-ray image.
[0015] The above-mentioned chest X-ray image recognition method, device, computer device, and storage medium, the method includes: obtaining a number of labeled chest X-ray images, and preprocessing them to obtain training samples; constructing a chest X-ray image recognition network; which includes a stem module, several consecutive ConvTransBlock modules, and a classification network; the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module for bridging the two branches; the convolutional branch is used to extract local features of input features; the Transformer branch is used for global features; the FCU module is used to fuse the local features of the convolutional branch with the global features of the Transformer branch; using the training samples to train the network, and using the trained chest X-ray image recognition network to recognize the chest X-ray image to be measured to obtain the category of the chest X-ray image. The model of this method has few parameters, a fast convergence speed, and high recognition accuracy. Description of the Drawings
[0016] Figure 1 It is a schematic flowchart of a chest X-ray image recognition method in an embodiment;
[0017] Figure 2 It is a structural diagram of chest X-ray image recognition in an embodiment;
[0018] Figure 3 It is a structural diagram of the ConvTransBlock module in an embodiment;
[0019] Figure 4 It is the FCU structure in another embodiment;
[0020] Figure 5 It is a structural block diagram of a chest X-ray image recognition device in an embodiment;
[0021] Figure 6 It is the internal structural diagram of a computer device in an embodiment. Detailed implementation manners
[0022] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] In one embodiment, as Figure 1 shown, a chest X-ray image recognition method is provided, and the method includes the following steps:
[0024] A chest X-ray image recognition method, the method includes:
[0025] Step 100: Obtain a number of labeled chest X-ray images, and preprocess the chest X-ray images to obtain training samples.
[0026] Specifically, the chest X-ray images include chest X-ray images of bacterial pneumonia (Pneu-Bacteria), viral pneumonia (Pneu-Viral), COVID-19, and normal (Noremal).
[0027] Step 102: Construct a chest X-ray image recognition network.
[0028] The chest X-ray image recognition network includes a stem module, a feature extraction network, and a classification network.
[0029] The Stem module is used to extract the initial local features of the training samples through convolution and maximum pooling operations.
[0030] The feature extraction network includes several consecutive ConvTransBlock modules; the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module that bridges the two branches; the convolutional branch is used to extract local features of the input features using an inverse residual structure composed of convolutional layers; the Transformer branch is used to extract complex spatial transformations and long-range feature dependencies of the input features to obtain global features; the FCU module is used to fuse the local features of the convolutional branch with the global features of the Transformer branch. The convolutional network is good at extracting local features of images, while the transformer structure is good at extracting global features of images.
[0031] The classification network is used to perform image recognition based on the output features of the convolutional branch and the Transformer branch of the last ConvTransBlock module respectively, and fuse the two obtained classification prediction results to obtain the sample prediction classification result.
[0032] The convolutional network is good at extracting local features of images, while the transformer structure is good at extracting global features of images. According to the characteristics of the convolutional network and the transformer structure in extracting image features respectively, a new parallel network module for chest X-ray image recognition network is proposed. When the convolutional network and the transformer network extract features, the feature dimensions are inconsistent. This module uses the FCU structure for information interaction. Aiming at the fact that the transformer structure often has a large number of parameters and a slow convergence speed, this structure uses an inverse residual structure with a smaller number of parameters for convolutional feature extraction to reduce the number of parameters and speed up the convergence speed.
[0033] Step 104: Train the chest X-ray image recognition network according to the identifier of the training sample and the sample prediction classification result obtained by inputting the training sample into the chest X-ray image recognition network, to obtain a trained chest X-ray image recognition network.
[0034] Step 106: Input the chest X-ray image to be measured into the trained chest X-ray image recognition network to obtain the category of the chest X-ray image.
[0035] In the above chest X-ray image recognition method, the method includes: obtaining a number of labeled chest X-ray images, preprocessing them to obtain training samples; constructing a chest X-ray image recognition network; the network includes a stem module, a number of consecutive ConvTransBlock modules, and a classification network; the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module for bridging the two branches; the convolutional branch is used to extract local features of the input features; the Transformer branch is used for global features; the FCU module is used to fuse the local features of the convolutional branch and the global features of the Transformer branch; training the network with the training samples, and using the trained chest X-ray image recognition network to recognize the chest X-ray image to be measured, so as to obtain the category of the chest X-ray image. The model of this method has few parameters and a fast convergence speed, and has a high recognition accuracy when recognizing chest X-ray images, especially has a strong recognition ability for the chest X-ray images of various pneumonia patients, and has important significance in the auxiliary diagnosis process of this disease.
[0036] In one of the embodiments, the Stem module consists of 1 convolutional layer and 1 max pooling layer; the feature extraction network further includes 1 convolutional branch and 1 Transformer branch; step 104 includes: inputting the training samples into the convolutional layer of the Stem module for feature extraction, and inputting the extracted features into the max pooling layer to obtain initial local features; respectively inputting the initial local features and the initial local features after passing through the Project module and normalization processing into the convolutional branch and the Transformer branch of the volume feature extraction network to obtain the input features of the convolutional branch and the input features of the Transformer branch of the first ConvTransBlock module; inputting the input features of the convolutional branch and the input features of the Transformer branch into the first ConvTransBlock module to obtain the first local feature and the first global feature; respectively using the first local feature and the first global feature as the inputs of the convolutional branch and the Transformer branch of the second ConvTransBlock module for feature extraction, and so on, inputting the (N-1)th local feature and the (N-1)th global feature output by the (N-1)th ConvTransBlock module into the convolutional branch and the Transformer branch of the Nth ConvTransBlock module respectively to obtain local features and global features; inputting the local features and the global features into the classification network to obtain the sample prediction classification result; training the chest X-ray image recognition network according to the labels of the training samples, the sample prediction classification result, and a preset loss function to obtain the trained chest X-ray image recognition network.
[0037] The Procject module includes: patch embedding and class embedding operations. The Procject module is used to segment the given input feature image into small patches of a fixed size, expand them into a one-dimensional vector through the patch embedding operation, and then, through the class embedding operation, add the class_token vector to the one-dimensional vector and output the one-dimensional vector; where class_token is a vector used for classification in the transformer branch. For example: a 14*14 feature map becomes a 196*1 feature vector after patch embedding, and becomes a 197*1 feature vector after adding class_token through class embedding. The main features of class embedding are: (1) not based on image content; (2) fixed position encoding.
[0038] Specifically, the structure of the chest X-ray image recognition network is as Figure 2 shown. The chest X-ray image recognition network includes a stem module, N ConvTransBlock modules, and two classifiers. The Stem module consists of a 7×7 convolution with a stride of 2 and a 3×3 max pooling layer with a stride of 2, and is used to extract initial local features. The N ConvTransBlock modules are connected in sequence to form a dual-branch feature extraction network composed of a convolutional branch and a Transformer branch respectively. The Transformer branch maps the initial local features through the project module, converting the feature map into a feature vector of a specified dimension, while the convolutional branch directly inputs the initial local feature map into the branch. The convolutional branch and the Transformer branch of the ConvTransBlock module are respectively composed of multiple convolutional layers and Transformer blocks. The structure of the ConvTransBlock module is as Figure 3 shown. This concurrent structure means that the convolutional branch and the Transformer branch can respectively maximize the retention of local features and global representations. The FCU, as a bridging module, fuses the local features of the convolutional branch with the global features of the Transformer branch. Along these branch structures, the FCU will gradually fuse the feature map and patch embedding in an interactive manner.
[0039] Finally, for the branches composed of convolutional branches, after all features are merged, they are input to a classifier. For the branches composed of Transformer branches, trainable class encodings are embedded in the Transformer. The class encodings are embedded by adding the class number to the input dimension. Since the convolutional branches already contain position information, the Transformer does not need to embed position encodings and finally is sent to the classifier of the upper branch network for classification.
[0040] In one embodiment, as Figure 3The structural diagram of the ConvTransBlock module shown, the convolutional branch consists of two inverse residual modules; the inverse residual module consists of 1 point convolutional layer, 1 depth convolutional layer and 1 point convolutional layer, and the convolutional kernel of the depth convolutional layer is 3×3; the Transformer branch includes a multi-head attention mechanism and a multi-layer perceptron; the FCU module includes a downsampling fusion branch and an upsampling fusion branch; step 104 further includes: inputting the feature of the convolutional branch into the first point convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module for feature extraction, and normalizing the extracted feature to obtain the first normalized local feature; inputting the first normalized local feature into the depth convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module, and normalizing the obtained result to obtain the first depth convolutional feature; inputting the first depth convolutional feature into the second point convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module to obtain the second convolutional feature; adding the second convolutional feature and the initial local feature, and inputting the obtained result into the first point convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module to obtain the third convolutional feature; inputting the first depth convolutional feature into the downsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain the local fusion feature; adding the local fusion feature and the feature map obtained by processing the initial local feature through the Project module to obtain the input feature of the Transformer branch; normalizing the input feature of the Transformer branch and then inputting it into the multi-head attention mechanism of the Transformer branch of the first ConvTransBlock module to obtain the global attention feature; adding the input feature of the Transformer branch and the global attention feature and then normalizing, and inputting the obtained normalized result into the multi-layer perceptron of the Transformer branch of the first ConvTransBlock module to obtain the first global feature; inputting the first global feature into the upsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain the global fusion feature; adding the third convolutional feature and the global fusion feature and then inputting it into the depth convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module, and normalizing the obtained result to obtain the second depth convolutional feature; inputting the second depth convolutional feature into the second point convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module, and adding the obtained convolutional result and the second convolutional feature to obtain the first local feature.
[0041] Specifically, the ConvTransBlock module consists of three parts: a convolutional branch, a Transformer branch, and a feature coupling unit (FCU module). The convolutional branch in this module is used to collect local features and retain local cues as features. Compared with the classical convolutional layer, the residual structure has fewer parameters and is easier to converge.
[0042] The convolutional branch consists of two inverse residual structures. Among them, the 1×1 convolution is used to change the number of channels, and the 3×3 depthwise convolution (DW convolution) is used to extract the features of each channel.
[0043] The Transformer branch is used to extract complex spatial transformations and long-range feature dependencies, thereby obtaining a global feature representation.
[0044] The Transformer branch consists of two parts: a multi-head attention mechanism and a multi-layer perceptron (MLP). LayerNorm is used for normalization before both parts. The given input feature image is segmented into small patches of a fixed size and then unfolded into a one-dimensional vector (patch embedding) and input into the Transformer branch. Since the convolutional branch (3×3 DW convolution) encodes both local features and spatial position information at the same time, the Transformer structure in this branch no longer requires position encoding.
[0045] The role of the FCU module is to eliminate the feature misalignment between the feature map in the given convolutional branch and the patch embedding in the Transformer branch. Because the feature dimensions generated by the convolutional branch and the Transformer branch are inconsistent and need to be dimensionally aligned. Therefore, when input to the Transformer branch, it is first necessary to align the channel dimensions through a 1×1 convolution, and then perform a downsampling operation to transform the feature map into the feature vector corresponding to the Transformer branch, as Figure 4 shown. When the one-dimensional feature vector from the Transformer branch is fed into the convolutional branch, it is also necessary to align the channel dimensions through a 1×1 convolution and perform an upsampling operation to transform the feature vector into a feature map. Since when the feature vector undergoes dimensional transformation, the 1×196 feature vector can only be transformed into the corresponding 14×14 feature map, so when transforming the 1×196 into a larger feature map, an upsampling process is required for vector padding, as Figure 4 shown.
[0046] In addition, LayerNorm and Batch Normalization are respectively used for feature normalization in the transformer branch and the convolutional branch.
[0047] Figure 3 In this, BN represents Batch Normalization, which is a data normalization method. Its function can accelerate the convergence speed during model training, make the model training process more stable, and avoid gradient explosion or gradient disappearance.
[0048] In one embodiment, the downsampling fusion branch includes a point convolutional layer and a downsampling module; step 104 further includes: inputting the first depth convolutional feature into the point convolutional layer of the downsampling branch of the FCU module of the first ConvTransBlock module for channel dimension alignment, and inputting the obtained feature with aligned channel dimensions into the downsampling module in the downsampling branch of the FCU module of the first ConvTransBlock module for downsampling; normalizing the obtained downsampling result to obtain a local fusion feature.
[0049] In one embodiment, the upsampling fusion branch includes a point convolutional layer and a downsampling module; step 104 further includes: inputting the first global feature into the upsampling module in the upsampling branch of the FCU module of the first ConvTransBlock module to obtain an upsampling result; inputting the upsampling result into the point convolutional layer of the FCU module of the first ConvTransBlock module, and inputting the obtained feature with aligned channel dimensions; inputting it into the downsampling module in the upsampling branch of the FCU module of the first ConvTransBlock module for downsampling; normalizing the obtained upsampling result to obtain a global fusion feature.
[0050] In one embodiment, the classification network includes two classifiers; step 104 further includes: pooling the local features and then inputting them into the first classifier to obtain a first predicted classification result; inputting the global features into the second classifier to obtain a second predicted classification result; adding the first predicted classification result and the second predicted classification result to obtain a sample predicted classification result.
[0051] In one embodiment, the preset loss function is two cross-entropy loss functions.
[0052] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,Figure 1 At least a part of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages do not necessarily need to be executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turns with at least a part of other steps or sub-steps or stages of other steps.
[0053] In one embodiment, as Figure 5 shown, a chest X-ray image recognition device is provided, including: a training sample determination module, a chest X-ray image recognition network construction module, a chest X-ray image recognition network training module, and a chest X-ray image category determination module, where:
[0054] The training sample determination module is used to obtain a number of labeled chest X-ray images and preprocess the chest X-ray images to obtain training samples.
[0055] The chest X-ray image recognition network construction module is used to construct a chest X-ray image recognition network; the chest X-ray image recognition network includes a stem module, a feature extraction network, and a classification network; the Stem module is used to extract initial local features of the training samples through convolution and maximum pooling operations; the feature extraction network includes a number of consecutive ConvTransBlock modules; the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module for bridging the two branches; the convolutional branch is used to extract local features of the input features using an inverted residual structure composed of convolutional layers; the Transformer branch is used to extract complex spatial transformations and long-distance feature dependencies of the input features to obtain global features; the FCU module is used to fuse the local features of the convolutional branch with the global features of the Transformer branch; the classification network is used to perform image recognition according to the output features of the convolutional branch and the Transformer branch of the last ConvTransBlock module respectively, and fuse the two classification prediction results obtained to obtain a sample prediction classification result.
[0056] The chest X-ray image recognition network training module is used to train the chest X-ray image recognition network according to the labels of the training samples and the sample prediction classification results obtained by inputting the training samples into the chest X-ray image recognition network to obtain a trained chest X-ray image recognition network.
[0057] The chest X-ray image category determination module is used to input the chest X-ray image to be measured into the trained chest X-ray image recognition network to obtain the category of the chest X-ray image.
[0058] In one of the embodiments, the Stem module consists of 1 convolutional layer and 1 max pooling layer; the feature extraction network further includes 1 convolutional branch and 1 Transformer branch; the chest X-ray image recognition network training module is further configured to input the training samples into the convolutional layer of the Stem module for feature extraction, and input the extracted features into the max pooling layer to obtain initial local features; input the initial local features and the initial local features after passing through the Project module and normalization processing into the convolutional branch and the Transformer branch of the convolutional feature extraction network respectively to obtain the convolutional branch input features and the Transformer branch input features of the first ConvTransBlock module; input the convolutional branch input features and the Transformer branch input features into the first ConvTransBlock module to obtain the first local feature and the first global feature; use the first local feature and the first global feature as the inputs of the convolutional branch and the Transformer branch of the second ConvTransBlock module respectively for feature extraction, and so on, input the (N-1)th local feature and the (N-1)th global feature output by the (N-1)th ConvTransBlock module into the convolutional branch and the Transformer branch of the Nth ConvTransBlock module respectively to obtain the local feature and the global feature; input the local feature and the global feature into the classification network to obtain the sample prediction classification result; train the chest X-ray image recognition network according to the identification of the training samples, the sample prediction classification result, and a preset loss function to obtain the trained chest X-ray image recognition network.
[0059] In one of the embodiments, the convolutional branch consists of two inverse residual modules; the inverse residual module consists of 1 layer of point convolutional layer, 1 layer of depth convolutional layer and 1 layer of point convolutional layer, and the convolutional kernel of the depth convolutional layer is 3×3; the Transformer branch includes a multi-head attention mechanism and a multi-layer perceptron; the FCU module includes a downsampling fusion branch and an upsampling fusion branch; the chest X-ray image recognition network training module is further configured to input the features of the convolutional branch into the first point convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module for feature extraction, and normalize the extracted features to obtain the first normalized local features; input the first normalized local features into the depth convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module, and normalize the obtained result to obtain the first depth convolutional features; input the first depth convolutional features into the second point convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module to obtain the second convolutional features; add the second convolutional features and the initial local features, and input the obtained result into the first point convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module to obtain the third convolutional features; input the first depth convolutional features into the downsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain local fusion features; add the local fusion features and the feature map obtained by processing the initial local features through the Project module to obtain the input features of the Transformer branch; normalize the input features of the Transformer branch and then input them into the multi-head attention mechanism of the Transformer branch of the first ConvTransBlock module to obtain global attention features; add the input features of the Transformer branch and the global attention features, normalize the result, and input the obtained normalized result into the multi-layer perceptron of the Transformer branch of the first ConvTransBlock module to obtain the first global features; input the first global features into the upsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain global fusion features; add the third convolutional features and the global fusion features and input them into the depth convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module, and normalize the obtained result to obtain the second depth convolutional features; input the second depth convolutional features into the second point convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module, and add the obtained convolutional result and the second convolutional features to obtain the first local features.
[0060] In one embodiment, the downsampling fusion branch includes a point convolution layer and a downsampling module; the chest X-ray image recognition network training module is further configured to input the first depth convolution feature into the point convolution layer of the downsampling branch of the FCU module of the first ConvTransBlock module for channel dimension alignment, and input the obtained feature with aligned channel dimensions into the downsampling module in the downsampling branch of the FCU module of the first ConvTransBlock module for downsampling; perform normalization processing on the obtained downsampling result to obtain a local fusion feature.
[0061] In one embodiment, the upsampling fusion branch includes a point convolution layer and a downsampling module; the chest X-ray image recognition network training module is further configured to input the first global feature into the upsampling module in the upsampling branch of the FCU module of the first ConvTransBlock module to obtain an upsampling result; input the upsampling result into the point convolution layer of the FCU module of the first ConvTransBlock module, and input the obtained feature with aligned channel dimensions; input it into the downsampling module in the upsampling branch of the FCU module of the first ConvTransBlock module for downsampling; perform normalization processing on the obtained upsampling result to obtain a global fusion feature.
[0062] In one embodiment, the classification network includes two classifiers; the chest X-ray image recognition network training module is further configured to input the local feature after pooling into the first classifier to obtain a first predicted classification result; input the global feature into the second classifier to obtain a second predicted classification result; add the first predicted classification result and the second predicted classification result to obtain a sample predicted classification result.
[0063] In one embodiment, the preset loss function in the chest X-ray image recognition network training module is two cross-entropy loss functions. Preferably, the importance weights of the two cross-entropy loss functions are set to be the same.
[0064] For the specific limitations of the chest X-ray image recognition device, reference can be made to the limitations on the chest X-ray image recognition method in the foregoing text, which will not be elaborated here. Each module in the above-mentioned chest X-ray image recognition device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0065] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for chest X-ray image recognition. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0066] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0067] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method in the above embodiment.
[0068] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of the method in the above embodiment.
[0069] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0070] In a validation embodiment, the experimental data set was experimented with the medical image classification data set Dataset-B in the MCFF-Ne paper. Dataset-B collected CXR images from five different public databases, namely: (1) Actualmed-COVID-chestxray-dataset; (2) COVID-19 Radiography Database; (3) Figure1-COVID-chestxray-dataset; (4) Pneumonia Virus vs Pneumonia Bacteria; (5) Chest X-ray Image. Dataset-B contains four categories of CXR images, namely COVID-19, normal, bacterial pneumonia, and viral pneumonia, with a total of 5,985 images. The training set has a total of 5,300 images, including 800 images of COVID-19 patients, 1,300 normal images, 1,600 viral pneumonia images, and 1,600 bacterial pneumonia images. The test set has a total of 741 images, including 142 images of COVID-19 patients, 200 normal images, 202 bacterial pneumonia images, and 197 viral pneumonia images.
[0071] (1) Model complexity and effectiveness
[0072] The number of parameters and computational complexity of this network, the Conformer network, and the ResNet101 network are shown in Table 1. The network structure proposed in the present invention alleviates the problem of the large number of parameters and computational complexity of the transformer network structure, and the network structure proposed in this paper is easier to converge from the comparison of various index results on the test set.
[0073] Table 1 Number of parameters and computational complexity of the model
[0074]
[0075] To reflect the recognition effect of this network on the medical image classification dataset, a comparative experiment was conducted using the ResNet101 network model and the Conformer-Ti network model. Table 2 shows the best accuracy rates of the ResNet101 network model, the Conformer-Ti network model, and this network on the dataset recorded after running 120 epochs under the same hardware environment. It can be seen from Table 2 that the classification accuracy of this network is about 1% higher than that of the ResNet101 network model and 2% higher than that of the Conformer-Ti network model.
[0076] It can be found that the numerical stability of the accuracy rate of this network model is better than that of the ResNet101 network model and the Conformer-Ti network model. At the same time, the values of other model evaluation indexes of these three networks, such as sensitivity, specificity, precision, F 1 value, etc.
[0077] Table 2 Results of various indexes on the test set
[0078]
[0079] The above results show that the best accuracy rate of this network model on the experimental chest X-ray image dataset can reach 95.33%, and the precision and sensitivity for identifying pneumonia are 95.30% and 95.33% respectively. Compared with other methods, this network model can detect various pneumonias from chest X-ray images more quickly, which can help improve the accuracy of manual film reading and save the energy of radiologists.
[0080] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0081] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for chest X-ray image recognition, characterized in that, the method includes: Obtain a number of labeled chest X-ray images, and preprocess the chest X-ray images to obtain training samples; Construct a chest X-ray image recognition network; the chest X-ray image recognition network includes a Stem module, a feature extraction network, and a classification network; the Stem module is used to extract initial local features of the training samples through convolution and max pooling operations; the feature extraction network includes several consecutive ConvTransBlock modules, and the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module that bridges the two branches; the convolutional branch is used to extract local features of the input features using an inverted residual structure composed of convolutional layers; the Transformer branch is used to extract the complex spatial transformation and long-distance feature dependencies of the input features to obtain global features; the FCU module is used to fuse the local features of the convolutional branch with the global features of the Transformer branch; the classification network is used to perform image recognition based on the output features of the convolutional branch and the Transformer branch of the last ConvTransBlock module respectively, and fuse the two classification prediction results obtained to obtain a sample prediction classification result; Train the chest X-ray image recognition network according to the labels of the training samples and the sample prediction classification results obtained by inputting the training samples into the chest X-ray image recognition network to obtain a trained chest X-ray image recognition network; Input the chest X-ray image to be measured into the trained chest X-ray image recognition network to obtain the category of the chest X-ray image; wherein, the convolutional branch is composed of two inverted residual modules; the inverted residual module is composed of 1 layer of point convolutional layer, 1 layer of depth convolutional layer, and 1 layer of point convolutional layer, and the convolutional kernel of the depth convolutional layer is 3×3; the Transformer branch includes a multi-head attention mechanism and a multi-layer perceptron; the FCU module includes a downsampling fusion branch and an upsampling fusion branch; In the first ConvTransBlock module: Input the input features of the convolutional branch into the first point convolutional layer of the first inverted residual module of the convolutional branch of the first ConvTransBlock module for feature extraction, and normalize the extracted features to obtain the first normalized local features; Input the first normalized local features into the depth convolutional layer of the first inverted residual module of the convolutional branch of the first ConvTransBlock module, and normalize the obtained result to obtain the first depth convolutional features; Input the first depth convolutional features into the second point convolutional layer of the first inverted residual module of the convolutional branch of the first ConvTransBlock module to obtain the second convolutional features; Add the second convolutional feature and the initial local feature, and input the obtained result into the first point convolutional layer of the second inverse residual module in the convolutional branch of the first ConvTransBlock module to obtain a third convolutional feature; Input the first depth convolutional feature into the downsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain a local fusion feature; Input the Transformer branch input feature into the multi-head attention mechanism of the Transformer branch of the first ConvTransBlock module to obtain a global attention feature; Add the Transformer branch input feature and the global attention feature and then perform normalization processing. Input the obtained normalized result into the multi-layer perceptron of the Transformer branch of the first ConvTransBlock module to obtain a first global feature; Input the first global feature into the upsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain a global fusion feature; Add the third convolutional feature and the global fusion feature and input the result into the depth convolutional layer of the second inverse residual module in the convolutional branch of the first ConvTransBlock module, and perform normalization processing on the obtained result to obtain a second depth convolutional feature; Input the second depth convolutional feature into the second point convolutional layer of the second inverse residual module in the convolutional branch of the first ConvTransBlock module, and add the obtained convolutional result and the second convolutional feature to obtain a first local feature.
2. The method according to claim 1, wherein, the Stem module is composed of 1 convolutional layer and 1 max pooling layer; the feature extraction network further includes 1 convolutional branch and 1 Transformer branch; According to the identifier of the training sample and the sample prediction classification result obtained by inputting the training sample into the chest X-ray image recognition network, train the chest X-ray image recognition network to obtain a trained chest X-ray image recognition network, including: Input the training sample into the convolutional layer of the Stem module for feature extraction, and input the extracted feature into the max pooling layer to obtain an initial local feature; Input the initial local feature and the initial local feature after passing through the Project module and normalization processing into the convolutional branch and the Transformer branch of the volume feature extraction network respectively to obtain the convolutional branch input feature and the Transformer branch input feature of the first ConvTransBlock module; Input the convolutional branch input feature and the Transformer branch input feature into the first ConvTransBlock module to obtain a first local feature and a first global feature; The first local feature and the first global feature are respectively used as the inputs of the convolutional branch and the Transformer branch of the second ConvTransBlock module for feature extraction. By analogy, the (N-1)th local feature and the (N-1)th global feature output by the (N-1)th ConvTransBlock module are respectively input into the convolutional branch and the Transformer branch of the Nth ConvTransBlock module to obtain local features and global features; The local features and the global features are input into a classification network to obtain a sample prediction classification result; According to the identifier of the training sample, the sample prediction classification result, and a preset loss function, the chest X-ray image recognition network is trained to obtain a trained chest X-ray image recognition network.
3. The method according to claim 1, wherein, the downsampling fusion branch includes a point convolutional layer and a downsampling module; Inputting the first depth convolutional feature into the downsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain local fusion features, including: Inputting the first depth convolutional feature into the point convolutional layer of the downsampling branch of the FCU module of the first ConvTransBlock module for channel dimension alignment, and inputting the obtained feature with aligned channel dimensions into the downsampling module in the downsampling branch of the FCU module of the first ConvTransBlock module for downsampling; Performing normalization processing on the obtained downsampling result to obtain local fusion features.
4. The method according to claim 1, wherein, the upsampling fusion branch includes a point convolutional layer and a downsampling module; Inputting the first global feature into the upsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain global fusion features, including: Inputting the first global feature into the upsampling module in the upsampling branch of the FCU module of the first ConvTransBlock module to obtain an upsampling result; Inputting the upsampling result into the point convolutional layer of the FCU module of the first ConvTransBlock module, and inputting the obtained feature with aligned channel dimensions into the downsampling module in the upsampling branch of the FCU module of the first ConvTransBlock module for downsampling; Performing normalization processing on the obtained upsampling result to obtain global fusion features.
5. The method according to claim 2, wherein, the classification network includes two classifiers; Inputting the local features and the global features into the classification network to obtain a sample prediction classification result, including: Pooling the local features and then inputting them into the first classifier to obtain a first prediction classification result; Inputting the global features into the second classifier to obtain a second prediction classification result; Adding the first prediction classification result and the second prediction classification result to obtain a sample prediction classification result.
6. The method according to claim 2, wherein, the chest X-ray image recognition network is trained according to the identifier of the training sample, the sample prediction classification result, and a preset loss function, to obtain a trained chest X-ray image recognition network, and the preset loss function in the step is two cross-entropy loss functions.
7. A chest X-ray image recognition device, wherein, the device includes: a training sample determination module, configured to obtain a plurality of labeled chest X-ray images, and preprocess the chest X-ray images to obtain training samples; a chest X-ray image recognition network construction module, configured to construct a chest X-ray image recognition network; the chest X-ray image recognition network includes a Stem module, a feature extraction network, and a classification network; the Stem module is configured to extract initial local features of the training sample through convolution and max pooling operations; the feature extraction network includes a plurality of consecutive ConvTransBlock modules; the ConvTransBlock module includes a convolutional branch, a Transformer branch, and an FCU module for bridging the two branches; the convolutional branch is configured to extract local features of the input features by using an inverted residual structure composed of convolutional layers; the Transformer branch is configured to extract complex spatial transformations and long-range feature dependencies of the input features to obtain global features; the FCU module is configured to fuse the local features of the convolutional branch with the global features of the Transformer branch; the classification network is configured to perform image recognition according to the output features of the convolutional branch and the Transformer branch of the last ConvTransBlock module respectively, and fuse the two obtained classification prediction results to obtain a sample prediction classification result; a chest X-ray image recognition network training module, configured to train the chest X-ray image recognition network according to the identifier of the training sample and the sample prediction classification result obtained by inputting the training sample into the chest X-ray image recognition network, to obtain a trained chest X-ray image recognition network; a chest X-ray image category determination module, configured to input a chest X-ray image to be measured into the trained chest X-ray image recognition network to obtain the category of the chest X-ray image; wherein, the convolutional branch is composed of two inverted residual modules; the inverted residual module is composed of 1 layer of point convolutional layer, 1 layer of depth convolutional layer, and 1 layer of point convolutional layer, and the convolutional kernel of the depth convolutional layer is 3×3; the Transformer branch includes a multi-head attention mechanism and a multi-layer perceptron; the FCU module includes a downsampling fusion branch and an upsampling fusion branch; The chest X-ray image recognition network construction module is further configured to, in the first ConvTransBlock module: input the convolutional branch input features into the first point convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module for feature extraction, and normalize the extracted features to obtain the first normalized local features; input the first normalized local features into the depth convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module, and normalize the obtained result to obtain the first depth convolutional features; input the first depth convolutional features into the second point convolutional layer of the first inverse residual module of the convolutional branch of the first ConvTransBlock module to obtain the second convolutional features; add the second convolutional features and the initial local features, and input the obtained result into the first point convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module to obtain the third convolutional features; input the first depth convolutional features into the downsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain local fusion features; input the Transformer branch input features into the multi-head attention mechanism of the Transformer branch of the first ConvTransBlock module to obtain global attention features; add the Transformer branch input features and the global attention features and then normalize them, and input the obtained normalized result into the multi-layer perceptron of the Transformer branch of the first ConvTransBlock module to obtain the first global features; input the first global features into the upsampling fusion branch of the FCU module of the first ConvTransBlock module to obtain global fusion features; add the third convolutional features and the global fusion features and input them into the depth convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module, and normalize the obtained result to obtain the second depth convolutional features; input the second depth convolutional features into the second point convolutional layer of the second inverse residual module of the convolutional branch of the first ConvTransBlock module, and add the obtained convolutional result and the second convolutional features to obtain the first local features.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable memory, having a computer program stored thereon, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
X-ray image recognition method and device, computer equipment and storage medium
CN112419321A
Channel attention feature extraction method and recognition method for chest X-ray image
CN112784856A