An image classification method, apparatus, terminal, and storage medium
By introducing convolutional layer, attention layer and classification layer into deep convolutional neural networks, the problem that deep convolutional neural networks only have local context information capture capabilities is solved, and the precise classification of images is achieved.
Patent Information
- Application Number
- CN202111583881.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-22
AI Technical Summary
Existing deep convolutional neural networks only have the ability to capture local context information, but do not have the ability to global model, resulting in poor classification performance.
By introducing convolutional layer, attention layer and classification layer into the target classification model, using the convolutional layer to obtain local feature information, the attention layer performs global modeling, and finally image classification is performed through the classification layer, specifically including cascading convolutional blocks and maximum pooling layer, attention module and hierarchical multi-head attention module, to capture and utilize global features.
The precise classification of images is realized, and the local feature information of the image can be captured and globally modeled, thereby improving the classification performance of the model.
Smart Images

Figure CN114219044B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an image classification method, apparatus, terminal, and storage medium. Background Art
[0002] A deep convolutional neural network (DCNN) has a powerful ability to learn the subtle differences between classes and the large differences within a class. Therefore, DCNN is considered the mainstream paradigm for various image-related tasks, such as image classification, semantic segmentation, and object detection. The multi-level structure of DCNN enables it to extract low, medium, and high-level features and automatically learn the semantic differences of digital images. However, the receptive field of DCNN is limited by the size of the convolutional kernel, and it only has the ability to capture local context information and does not have the ability of global modeling, resulting in poor classification performance of DCNN.
[0003] Therefore, the prior art still needs to be improved and developed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an image classification method, apparatus, terminal, and storage medium for the above-mentioned defects of the prior art, aiming to solve the problem that the deep convolutional neural network in the prior art only has the ability to capture local context information and does not have the ability of global modeling, resulting in poor classification performance of the model.
[0005] The technical solution adopted by the present invention to solve the problem is as follows:
[0006] In a first aspect, an embodiment of the present invention provides an image classification method, where the method includes:
[0007] Obtain an image to be classified, and input the image to be classified into a target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer;
[0008] Obtain local feature information of the image to be classified through the convolutional layer to obtain a first feature map;
[0009] Perform global modeling on the first feature map through the attention layer to obtain a second feature map;
[0010] Perform image classification on the second feature map through the classification layer to obtain an image category corresponding to the image to be classified.
[0011] In an implementation manner, the convolutional layer includes a plurality of cascaded first convolutional blocks and a max pooling layer. The obtaining of local feature information of the image to be classified through the convolutional layer to obtain a first feature map includes:
[0012] Perform a convolution operation on the image to be classified through several cascaded first convolution blocks to obtain an initial feature map;
[0013] Perform downsampling on the initial feature map through the max pooling layer to obtain the first feature map.
[0014] In one implementation, the attention layer includes several cascaded attention modules and a hierarchical multi-head attention module. The global modeling of the first feature map through the attention layer to obtain the second feature map includes:
[0015] Input the first feature map into the first attention module, and obtain the position weight calibrated feature map output by the last attention module, where the position weight calibrated feature map includes several regions, each region has a position weight value, and each position weight value is used to reflect the level of spatial attention and channel attention corresponding to a region;
[0016] Input the position weight calibrated feature map into the hierarchical multi-head attention module, and output the second feature map through the hierarchical multi-head attention module.
[0017] In one implementation, each attention module includes a segmentation attention module and a coordinate attention module,
[0018] The segmentation attention module is used to output a weight calibrated feature map according to the input feature map, where the weight calibrated feature map includes several regions, each region has a weight value, and the size of the weight value is used to reflect the level of channel attention corresponding to the region;
[0019] The coordinate attention module is used to output the position weight calibrated feature map according to the weight calibrated feature map.
[0020] In one implementation, the segmentation attention module includes a global average pooling layer, a first fully connected layer, and an r-Softmax layer. The output of the weight calibrated feature map according to the input feature map includes:
[0021] Perform feature mapping on the input feature map to obtain several feature mapping maps, where several feature mapping maps respectively correspond to different mapping paths;
[0022] Fuse several feature mapping maps to obtain a group of feature mapping maps;
[0023] Input the group of feature mapping maps into the global average pooling layer to obtain global context information;
[0024] Input the global context information into the first fully connected layer to obtain first-channel weight value information;
[0025] Input the first-channel weight value information into the r-Softmax layer to obtain several groups of attention weight value information;
[0026] Perform weight calibration on several feature maps one by one according to several groups of the attention weight value information to obtain several initial weight-calibrated feature maps;
[0027] Fuse several initial weight-calibrated feature maps to obtain the weight-calibrated feature map.
[0028] In one implementation, the coordinate attention module includes a horizontal global average pooling layer, a vertical global average pooling layer, a second fully connected layer, and an activation function layer. Obtaining the position weight-calibrated feature map corresponding to the weight-calibrated feature map through the coordinate attention module includes:
[0029] Input the weight-calibrated feature map into the horizontal global average pooling layer to obtain a horizontal perception attention map, and input the weight-calibrated feature map into the vertical global average pooling layer to obtain a vertical perception attention map;
[0030] Input the horizontal perception attention map and the vertical perception attention map into the second fully connected layer to obtain second-channel weight value information;
[0031] Divide the second-channel weight value information into horizontal-channel weight value information and vertical-channel weight value information along the horizontal and vertical directions respectively;
[0032] Input the horizontal-channel weight value information and the vertical-channel weight value information into the activation function layer respectively to obtain horizontal attention weight information corresponding to the horizontal-channel weight value information and vertical attention weight information corresponding to the vertical-channel weight value information;
[0033] Determine the position weight value of each region on the weight-calibrated feature map according to the horizontal attention weight information and the vertical attention weight information to obtain the position weight-calibrated feature map.
[0034] In one implementation, the hierarchical multi-head attention module includes several self-attention modules with a hierarchical relationship. Inputting the position weight-calibrated feature map into the hierarchical multi-head attention module and outputting the second feature map through the hierarchical multi-head attention module includes:
[0035] Use the position weight calibrated feature map as the input image of the first self-attention module, and downsample the input image of the previous self-attention module to obtain the input image of the next self-attention module;
[0036] Obtain the attention feature maps output by each self-attention module to obtain a plurality of the attention feature maps;
[0037] Fuse the plurality of attention feature maps to obtain the second feature map.
[0038] In a second aspect, an embodiment of the present invention further provides an image classification device, where the device includes:
[0039] An image input module, configured to obtain an image to be classified and input the image to be classified into a target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer;
[0040] A local extraction module, configured to obtain local feature information of the image to be classified through the convolutional layer to obtain a first feature map;
[0041] A global modeling module, configured to perform global modeling on the first feature map through the attention layer to obtain a second feature map;
[0042] An image classification module, configured to perform image classification on the second feature map through the classification layer to obtain an image category corresponding to the image to be classified.
[0043] In a third aspect, an embodiment of the present invention further provides a terminal, where the terminal includes a memory and one or more processors; the memory stores one or more programs; the programs include instructions for executing the image classification method as described in any one of the above; the processor is configured to execute the programs.
[0044] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a plurality of instructions are stored, where the instructions are suitable for being loaded and executed by a processor to implement the steps of the image classification method as described in any one of the above.
[0045] Advantages of the present invention: In the embodiments of the present invention, by obtaining an image to be classified and inputting the image to be classified into a target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer; obtaining local feature information of the image to be classified through the convolutional layer to obtain a first feature map; performing global modeling on the first feature map through the attention layer to obtain a second feature map; and performing image classification on the second feature map through the classification layer to obtain the image category corresponding to the image to be classified. The target classification model in the present invention can capture local feature information of the image to be classified and can also perform global modeling, so that the image category corresponding to the image to be classified can be accurately predicted. It solves the problem in the prior art that the deep convolutional neural network only has the ability to capture local context information and does not have the ability of global modeling, resulting in poor classification performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 is a schematic flowchart of the image classification method provided by the embodiments of the present invention.
[0048] Figure 2 is a schematic structural diagram of the target classification model provided by the embodiments of the present invention.
[0049] Figure 3 is a schematic structural diagram of the split attention module and the coordinate attention module provided by the embodiments of the present invention.
[0050] Figure 4 is a schematic structural diagram of the hierarchical multi-head attention module provided by the embodiments of the present invention.
[0051] Figure 5 is a schematic diagram of the ROC curve of the comparison method and the ablation experiment provided by the embodiments of the present invention.
[0052] Figure 6 is the Grad-CAM result graph provided by the embodiments of the present invention.
[0053] Figure 7 is a schematic diagram of the visualized data of t-SNE provided by the embodiments of the present invention.
[0054] Figure 8 is a schematic diagram of the modules of the image classification device provided by the embodiments of the present invention.
[0055] Figure 9It is a schematic block diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners
[0056] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention, but not to limit the present invention.
[0057] It should be noted that if there are directional indications (such as up, down, left, right, front, back, etc.) involved in the embodiments of the present invention, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0058] A deep convolutional neural network (DCNN) has a powerful ability to learn the subtle differences between classes and the large differences within classes. Therefore, DCNN is considered the mainstream paradigm for various image-related tasks, such as image classification, semantic segmentation, and object detection. The multi-level structure of DCNN enables it to extract low-, medium-, and high-level features and automatically learn the semantic differences of digital images. However, the receptive field of DCNN is limited by the size of the convolutional kernel, and it only has the ability to capture local context information and does not have the ability of global modeling, resulting in poor classification performance of DCNN.
[0059] Aiming at the above defects of the prior art, the present invention provides an image classification method. The method includes obtaining an image to be classified and inputting the image to be classified into a target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer; obtaining local feature information of the image to be classified through the convolutional layer to obtain a first feature map; performing global modeling on the first feature map through the attention layer to obtain a second feature map; and performing image classification on the second feature map through the classification layer to obtain the image category corresponding to the image to be classified. The target classification model in the present invention can capture the local feature information of the image to be classified and can also perform global modeling, so it can accurately predict the image category corresponding to the image to be classified. This solves the problem in the prior art that the deep convolutional neural network only has the ability to capture local context information and does not have the ability of global modeling, resulting in poor classification performance of the model.
[0060] As Figure 1 shown, the method includes the following steps:
[0061] Step S100: Obtain the image to be classified, and input the image to be classified into the target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer.
[0062] Specifically, the image to be classified in this embodiment can be any image that needs to predict the image category. For example, it can be a fluorescence image of Pseudomonas aeruginosa to be predicted for the bacterial category (sensitive or drug-resistant). In order to obtain the image category of the image to be classified, a target classification model is pre-constructed and trained in this embodiment. As Figure 2 shown (H and W respectively represent the height and width of the input image, and C i represents the number of channels of the feature map), the target classification model includes a convolutional layer, an attention layer, and a classification layer. Among them, the convolutional layer can extract local features of the input image, the attention layer can perform global modeling on the input image, and the classification layer can accurately classify the image according to the extracted local feature information and global feature information, and finally obtain the image category corresponding to the image to be classified.
[0063] As Figure 1 shown, the method further includes the following steps:
[0064] Step S200: Obtain the local feature information of the image to be classified through the convolutional layer, and obtain a first feature map.
[0065] Specifically, input the image to be classified into the target classification model, and the image to be classified is first used as the input image of the convolutional layer. The convolutional layer can extract the local feature information of the image to be classified and output a first feature map containing its local feature information.
[0066] In one implementation, the convolutional layer includes a number of cascaded first convolutional blocks and a max pooling layer, and step S200 specifically includes the following steps:
[0067] Step S201: Perform a convolution operation on the image to be classified through a number of cascaded first convolutional blocks to obtain an initial feature map;
[0068] Step S202: Perform downsampling on the initial feature map through the max pooling layer to obtain the first feature map.
[0069] Specifically, the convolutional layer in this embodiment includes multiple first convolutional blocks. The multiple first convolutional blocks have a cascading relationship. Each first convolutional block performs a convolution operation on the input image to extract its local features. The image to be classified serves as the input image of the first first convolutional block. Then, the output of the previous first convolutional block is the input of the next first convolutional block. Then, the output of the last first convolutional block is obtained to get the initial feature map. A max pooling layer is connected after the last first convolutional block in the convolutional layer. The initial feature map output by the last convolutional block is the input image of this max pooling layer. The max pooling layer downsamples the input initial feature map and outputs the first feature map.
[0070] For example, the convolutional layer in this embodiment consists of three consecutive 3×3 convolutional layers (stride 2, stride 1, stride 1) and a max pooling layer (stride 2).
[0071] As Figure 1 shown, the method further includes the following steps:
[0072] Step S300: Perform global modeling on the first feature map through the attention layer to obtain the second feature map.
[0073] Specifically, to add the global modeling ability of the model, an attention layer is set in the target classification model in this embodiment. The attention layer can capture long-range feature dependencies, thereby learning global feature representations. Therefore, it can perform global modeling on the first feature map and output the second feature map.
[0074] In one implementation, the attention layer includes several cascaded attention modules and a hierarchical multi-head attention module. Step S300 specifically includes the following steps:
[0075] Step S301: Input the first feature map into the first attention module, and obtain the position weight calibration feature map output by the last attention module. Among them, the position weight calibration feature map includes several regions, and each region has a position weight value. Each position weight value is used to reflect the level of spatial attention and channel attention corresponding to a region;
[0076] Step S302: Input the position weight calibration feature map into the hierarchical multi-head attention module, and output the second feature map through the hierarchical multi-head attention module.
[0077] Specifically, the attention layer in this embodiment includes multiple attention modules, and the multiple attention modules have a cascading relationship, that is, the first feature map serves as the input image of the first attention module, and the output image of the previous attention module serves as the input image of the next attention module. Each attention module determines the levels of spatial attention and channel attention for each region in the input image and outputs a position weight calibration feature map. A hierarchical multi-head attention module is connected after the last attention module, and the position weight calibration feature map output by the last attention module serves as the input image of this hierarchical multi-head attention module. The hierarchical multi-head attention module models the global feature relationship of the input image in a hierarchical manner and outputs a second feature map.
[0078] In one implementation, each of the attention modules includes a segmentation attention module and a coordinate attention module.
[0079] The segmentation attention module is used to output a weight calibration feature map according to the input feature map, where the weight calibration feature map includes several regions, and each region has a weight value, and the magnitude of the weight value is used to reflect the level of channel attention corresponding to the region.
[0080] The coordinate attention module is used to output the position weight calibration feature map according to the weight calibration feature map.
[0081] Specifically, as Figure 3 shown (where H, W, and C represent the height, width, and number of channels of the feature map respectively), each attention module in this embodiment includes two parts, one is a segmentation attention module ( Figure 3 a), and the other is a coordinate attention module ( Figure 3 b). In an attention module, the input image of this attention module is the input image of the segmentation attention module, the output image of the segmentation attention module is the input image of the coordinate attention module, and the output image of the coordinate attention module is the output image of this attention module. Among them, the segmentation attention module is used to determine the level of channel attention for each region in the input image, and the coordinate attention module is used to determine the level of spatial attention for each region in the input image. Therefore, the position weight calibration feature map output by the attention module can reflect the levels of channel attention and spatial attention for each region in the input image of this attention module.
[0082] In one implementation, the segmentation attention module includes a global average pooling layer, a first fully connected layer, and an r-Softmax layer. The outputting a weight calibration feature map according to the input feature map includes:
[0083] Step S10: Perform feature mapping on the input feature map to obtain a number of feature mapping maps, where the number of feature mapping maps respectively correspond to different mapping paths;
[0084] Step S11: Fuse the number of feature mapping maps to obtain a feature mapping map group;
[0085] Step S12: Input the feature mapping map group into the global average pooling layer to obtain global context information;
[0086] Step S13: Input the global context information into the first fully connected layer to obtain first channel weight value information;
[0087] Step S14: Input the first channel weight value information into the r-Softmax layer to obtain a number of groups of attention weight value information;
[0088] Step S15: Calibrate the weights of the number of feature mapping maps one-to-one according to the number of groups of attention weight value information to obtain a number of initial weight-calibrated feature maps;
[0089] Step S16: Fuse the number of initial weight-calibrated feature maps to obtain the weight-calibrated feature map.
[0090] Specifically, the split attention module in this embodiment includes a global average pooling layer, a first fully connected layer, and an r-Softmax layer. For a split attention module, the feature map input to the split attention module will first perform feature mapping through multiple mapping paths, and a feature mapping map will be generated based on each mapping path, resulting in multiple feature mapping maps. After these feature mapping maps are fused by element-wise addition, a feature mapping group is obtained. Then, the feature mapping group is input to the global average pooling layer, and the global average pooling layer will calculate the average value of all pixels in each feature map in the feature mapping group and output a value, which summarizes the global context information (more robust to spatial translation of the input). Then, the global context information is input to the first fully connected layer (FC layer). Since the first fully connected layer can achieve information interaction between feature channels, it can output the first channel weight value information based on the input global context information. Then, the first channel weight value information is input to the r-Softmax layer, and the r-Softmax layer normalizes the first channel weight value information to generate several groups of attention weight value information. Among them, the number of groups of the several groups of attention weight value information is the same as the number of mapping paths. For example, if the number of mapping paths is 2, then two groups of attention weight value information are obtained. The last feature mapping map is weighted and calibrated using a group of attention weight value information. After calibration, several initial weight-calibrated feature maps are obtained, and then all the initial weight-calibrated feature maps are fused by element-wise addition to obtain the weight-calibrated feature map output by this split attention module. The split attention module in this embodiment can identify discriminative regions with rich visual information by performing information interaction between channels and using attention weights for feature recalibration.
[0091] In one implementation, the calculation method of r-Softmax is as follows:
[0092]
[0093] where R represents the number of split paths in each feature cardinality group, and in this embodiment, R can be set to 2. Based on the global context representation S k , represents the weight of the c-th channel in each split path.
[0094] In one implementation, the coordinate attention module includes a horizontal global average pooling layer, a vertical global average pooling layer, a second fully connected layer, and an activation function layer. Obtaining the position weight-calibrated feature map corresponding to the weight-calibrated feature map through the coordinate attention module includes:
[0095] Step S20: Input the weight-calibrated feature map into the horizontal global average pooling layer to obtain a horizontal perception attention map, and input the weight-calibrated feature map into the vertical global average pooling layer to obtain a vertical perception attention map;
[0096] Step S21: Input the horizontal perception attention map and the vertical perception attention map into the second fully connected layer to obtain second-channel weight value information;
[0097] Step S22: Divide the second-channel weight value information into horizontal-channel weight value information and vertical-channel weight value information along the horizontal and vertical directions respectively;
[0098] Step S23: Input the horizontal-channel weight value information and the vertical-channel weight value information into the activation function layer respectively to obtain horizontal attention weight information corresponding to the horizontal-channel weight value information and vertical attention weight information corresponding to the vertical-channel weight value information;
[0099] Step S24: Determine the position weight value of each region on the weight-calibrated feature map according to the horizontal attention weight information and the vertical attention weight information to obtain the position weight-calibrated feature map.
[0100] Specifically, in this embodiment, two global average pooling layers are set up. One is a horizontal global average pooling layer that performs global average pooling operations on the input weight-calibrated feature map along the X direction; the other is a vertical global average pooling layer that performs global average pooling operations on the input feature map along the Y direction. Through these two global average pooling layers, a pair of direction perception attention maps, namely a horizontal perception attention map and a vertical perception attention map, are obtained. Then, this pair of direction perception attention maps are concatenated and input into the second fully connected layer. After information interaction between channels through the second fully connected layer, second-channel weight value information is output. Then, the output second-channel weight value information is divided into two separate tensors along the spatial dimensions (horizontal and vertical), namely horizontal-channel weight value information and vertical-channel weight value information. Finally, the horizontal-channel weight value information and the vertical-channel weight value information are respectively input into the activation function layer (such as the sigmoid activation function) to obtain horizontal attention weight information in the X direction and vertical attention weight information in the Y direction, and the attention weight information in these two directions is superimposed on the input weight-calibrated feature map to obtain the position weight-calibrated feature map. In this embodiment, the position information is embedded into the channel attention through the coordinate attention module, which helps the target classification model capture the position information of the target of interest.
[0101] In one implementation, the hierarchical multi-head attention module includes a number of self-attention modules with a hierarchical relationship. Inputting the position-weight calibrated feature map into the hierarchical multi-head attention module and outputting the second feature map through the hierarchical multi-head attention module includes:
[0102] Step S3021: Use the position-weight calibrated feature map as the input image of the first-layer self-attention module, and downsample the input image of the previous-layer self-attention module to obtain the input image of the next-layer self-attention module;
[0103] Step S3022: Obtain the attention feature maps output by each layer of the self-attention module to obtain a number of attention feature maps;
[0104] Step S3023: Fuse the number of attention feature maps to obtain the second feature map.
[0105] Specifically, as Figure 3 (a) shows, the hierarchical multi-head attention module (H-MHSA) in this embodiment includes a number of self-attention modules (MHSA) with a cascading relationship. Use the position-weight calibrated feature map as the input image of the first-layer self-attention module, and then downsample the input image of the previous-layer self-attention module to obtain the input image of the next-layer self-attention module, that is, the size of the input image of each layer of the self-attention module decreases in turn, thereby reducing the computational complexity of the target classification model. Then, fuse the attention feature maps output by each layer of the self-attention module by element-wise addition to obtain the second feature map.
[0106] In one implementation, each self-attention module includes a point convolution module, a position encoding module, and a softmax layer. Obtaining the attention feature map output by each self-attention module includes:
[0107] Step S30: Input the input image of each self-attention module into the point convolution module to obtain a Q weight matrix, a K weight matrix, and a V weight matrix;
[0108] Step S31: Input the input image of each self-attention module into the position encoding module to obtain a position encoding map;
[0109] Step S32: Multiply the Q weight matrix by the position encoding map to obtain a first matrix, and multiply the Q weight matrix by the K weight matrix to obtain a second matrix;
[0110] Step S33: Perform element-wise addition on the first matrix and the second matrix to obtain a third matrix, and input the third matrix into the softmax layer;
[0111] Step S34: Matrix-multiply the output result of the softmax layer with the V weight matrix to obtain the attention feature map.
[0112] For example, Figure 4 (b) shows the structure of the self-attention module in this embodiment. First, the input feature map is divided into small tiles of size G x G, that is, each tile contains G x G pixels (for example, G is set to 4), and then the dimension is adjusted to:
[0113]
[0114] A = Softmax(QK T + QP T )V
[0115] where Q = X'W Q , K = X'W K and V = X'W V are the weight matrices generated by the input feature map through pointwise convolution, where W Q , W K and W V are learnable parameters and will be updated together with the model parameters during training. The simple self-attention mechanism cannot capture the input order, but the spatial position information between the input image patches plays a crucial role in the model's understanding of the context information. In this embodiment, relative position encoding is used to incorporate the position information into the self-attention structure. Specifically, two trainable matrices P h and P w are introduced, representing the position encodings along the height and width of the feature map respectively. First, P h and P w are added element-wise to obtain the relative position information P. When calculating the correlation between pixel points x j and x i , the relationship of the position information of x j to x i also needs to be considered additionally. Therefore, the calculation of the association degree (attention score) is α = QK + QP, and then the softmax operation is performed on all the calculated α values to obtain α'. Then, the formula for calculating A is applied to extract the important context information based on the association degree, enhance the effective information, and suppress the invalid information. Specifically, the attention features are calculated twice through the MHSA layer to obtain the attention feature maps A0 and A1 respectively, and then their dimensions are adjusted to the shape of the input feature map X. Finally, A0, A1, and X are added element-wise to obtain the output of this module.
[0116] The traditional MHSA module calculates attention across the entire input image, and its computational complexity is proportional to the square of the number of image patches (N), i.e.:
[0117] Ω time (MHSA) = 4NC 2 + 2N 2 C
[0118] In contrast, the H-MHSA module in this embodiment calculates attention in a hierarchical manner, so that only a limited number of image patches are processed in each step. A0 and A1 are calculated within each G x G small tile, and the computational cost is significantly reduced, i.e.:
[0119] Ω time (H-MHSA) = 4NC 2 + 2G 2 NC
[0120] As Figure 1 shown, the method further includes the following steps:
[0121] Step S400: Classify the second feature map through the classification layer to obtain the image category corresponding to the image to be classified.
[0122] Specifically, in the target classification model of this embodiment, a classification layer is further connected after the attention layer. The input image of this classification layer is the second feature map output by the attention layer. Since the second feature map can reflect the local feature information of the image to be classified and the information obtained after global modeling, the classification layer can accurately perform image classification based on the input second feature map, and then obtain the image category corresponding to the image to be classified.
[0123] To prove the technical effect of the present invention, the inventors conducted the following experiments:
[0124] Data and Experimental Settings:
[0125] Fluorescent images of Pseudomonas aeruginosa were obtained. Specifically, 48 clinical strains were selected, and resistance to 6 common antibiotics (ceftazidime, ciprofloxacin, imipenem, levofloxacin, moxifloxacin, and tobramycin) was detected. According to the minimum inhibitory concentration, 12 multi-drug resistant strains (i.e., showing resistance to all six antibiotics) and 11 sensitive strains were screened out. These bacteria were cultured and stained in vitro. Finally, about 100 images were taken of each sample with a fluorescence microscope. A database containing 2625 fluorescent images of sensitive Pseudomonas aeruginosa, namely the PAFI database, was established. 1233 images were of sensitive Pseudomonas aeruginosa (PA), and 1392 images were of multi-drug resistant Pseudomonas aeruginosa (MDRPA). Each image was randomly assigned to three sets, including 1683 images in the training set, 421 images in the validation set, and 521 images in the test set.
[0126] To save computing resources, the inventors adjusted the size of the original image to 320x320x3. To prevent the network from overfitting, data augmentation was performed on the training data through various transformations such as translation, rotation, flipping, affine, color jitter, and grayscale. The amplitude of each transformation can be controlled by a relative parameter (e.g., rotation angle), and each transformation will be executed with different probabilities. For fair comparison, all settings in the experiment are consistent for all comparison methods.
[0127] To measure the prediction performance of the proposed CTN, accuracy (Acc), precision (Pre), recall (Re), F1-score (F1), kappa (Kap), and area under the curve (AUC) were used to evaluate the prediction results. The inventors selected adaptive moment estimation (Adam) as the optimizer to iteratively optimize the model, and the learning rate was set to 0.0001, that is, training the model from scratch instead of using a pre-trained model. The number of training epochs was set to 200. In addition, cross-entropy loss was used as the loss function of the model, and the algorithm was implemented on the PyTorch platform with two NVIDIA TITAN X GPUs.
[0128] The experimental results are as follows:
[0129] As shown in Table 1, the classification performance of ResNeXt is slightly better than that of ResNet. This indicates that ResNeXt with a wider network structure can extract more fine-grained information. As lightweight networks, Shufflenetv2 and Mobilenetv2 are fast but lack the ability to represent fine-grained features. DenseNet shows similar performance to ResNeXt because the dense connection mechanism can achieve feature reuse. ViT is the first non-convolutional transformer network with performance comparable to CNN models. However, ViT needs to be trained on a very large dataset to perform well. Therefore, ViT performs poorly on our small PAFI database.
[0130] Table 1. Prediction results (%) of different methods on the test set of the PAFI database
[0131]
[0132]
[0133] Ablation experiment:
[0134] The inventors evaluated various design options for CTN. The experimental results are shown in Table 2. The inventors started with the original ResNeSt-50 and set different numbers of groups. The results showed that when the number of groups was 2, the performance of the model was better. Therefore, subsequent experiments were all based on a group number of 2. In addition, the inventors compared two attention mechanisms, CBAM and CA. Under the same experimental conditions, the performance of CA was better. As can be seen from (i) and (j) in Figure 6 , CA can extract discriminant regions more accurately than CBAM. Finally, the inventors applied the test-time augmentation (TTA) strategy, performed five data augmentations on the prediction samples, and then averaged these augmented predictions, which also improved the results of the model.
[0135] Table 2. Ablation experiments on the PAFI database, "g" represents the number of groups in ResNeSt-50 (%)
[0136]
[0137] To more intuitively evaluate the classifier, the inventors plotted the ROC curve, as shown in Figure 5 . By comparing the values of the area under the curve (AUC), it can be observed that CTN has excellent classification performance. Figure 6 shows the visualization of Grad-CAM (Grad-CAM highlights the discriminant regions for predicting sensitive PA and MDRPA. The red regions correspond to the classes with high scores. The default number of groups for ResNeSt-50 is 2). It can be seen that the H-MHSA module can fuse non-local information and help the network more precisely locate the class-related regions in the image. As can be seen from Figure 7 , the proposed CTN can effectively identify sensitive Pseudomonas aeruginosa and multidrug-resistant Pseudomonas aeruginosa.
[0138] Therefore, the target classification model in the present invention can automatically identify sensitive Pseudomonas aeruginosa and multidrug-resistant Pseudomonas aeruginosa. Specifically, the coordinate attention module can locate the position of the attention object, which helps the network extract fine-grained information from the target region of interest. The H-MHSA module can make up for the defect that DCNN cannot effectively capture long-range dependencies. Replace the last three 3×3 spatial convolutions in the network with H-MHSA. In this way, H-MHSA can learn global feature representations from the feature maps captured by the convolution and is not as computationally intensive as the traditional MHSA. The experimental results show that the present invention is effective in predicting Pseudomonas aeruginosa drug resistance and can help clinicians make decisions.
[0139] Based on the above embodiments, the present invention also provides an image classification device, as shown in Figure 8 . The device includes:
[0140] The image input module 01 is used to obtain the image to be classified and input the image to be classified into the target classification model. Among them, the target classification model includes a convolutional layer, an attention layer, and a classification layer;
[0141] The local extraction module 02 is used to obtain the local feature information of the image to be classified through the convolutional layer to obtain a first feature map;
[0142] The global modeling module 03 is used to perform global modeling on the first feature map through the attention layer to obtain a second feature map;
[0143] The image classification module 04 is used to perform image classification on the second feature map through the classification layer to obtain the image category corresponding to the image to be classified.
[0144] Based on the above embodiments, the present invention also provides a terminal, and its principle block diagram can be as Figure 9 shown. The terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the image classification method. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.
[0145] Those skilled in the art can understand that Figure 9 the principle block diagram shown in
[0146] In one implementation, one or more programs are stored in the memory of the terminal, and are configured to be executed by one or more processors. The one or more programs include instructions for performing the image classification method.
[0147] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0148] In summary, the present invention discloses an image classification method, apparatus, terminal, and storage medium. The method includes obtaining an image to be classified and inputting the image to be classified into a target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer; obtaining local feature information of the image to be classified through the convolutional layer to obtain a first feature map; performing global modeling on the first feature map through the attention layer to obtain a second feature map; and performing image classification on the second feature map through the classification layer to obtain an image category corresponding to the image to be classified. The target classification model in the present invention can capture local feature information of the image to be classified and can also perform global modeling, so it can accurately predict the image category corresponding to the image to be classified. This solves the problem in the prior art that a deep convolutional neural network only has the ability to capture local context information and does not have the ability to perform global modeling, resulting in poor classification performance of the model.
[0149] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. An image classification method, characterized in that, The method includes: Obtain an image to be classified, and input the image to be classified into a target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer; Obtain local feature information of the image to be classified through the convolutional layer to obtain a first feature map; Perform global modeling on the first feature map through the attention layer to obtain a second feature map, including: the attention layer includes a plurality of cascaded attention modules and a hierarchical multi-head attention module, and the hierarchical multi-head attention module includes a plurality of self-attention modules with a hierarchical relationship; input the first feature map into the first attention module, and obtain a position weight calibration feature map output by the last attention module, where the position weight calibration feature map includes a plurality of regions, each region has a position weight value, and each position weight value is used to reflect the level of spatial attention and channel attention corresponding to a region; input the position weight calibration feature map into the hierarchical multi-head attention module, and output the second feature map through the hierarchical multi-head attention module; Perform image classification on the second feature map through the classification layer to obtain the image category corresponding to the image to be classified.
2. The image classification method according to claim 1, wherein The convolutional layer includes a plurality of cascaded first convolutional blocks and a max pooling layer. The obtaining of the local feature information of the image to be classified through the convolutional layer to obtain a first feature map includes: Perform a convolution operation on the image to be classified through a plurality of cascaded first convolutional blocks to obtain an initial feature map; Perform downsampling on the initial feature map through the max pooling layer to obtain the first feature map.
3. The image classification method according to claim 1, wherein Each attention module includes a split attention module and a coordinate attention module. The split attention module is used to output a weight calibration feature map according to the input feature map, where the weight calibration feature map includes a plurality of regions, each region has a weight value, and the size of the weight value is used to reflect the level of channel attention corresponding to the region; The coordinate attention module is used to output the position weight calibration feature map according to the weight calibration feature map.
4. The image classification method according to claim 3, wherein The split attention module includes a global average pooling layer, a first fully connected layer, and an r-Softmax layer. The outputting of the weight calibration feature map according to the input feature map includes: Perform feature mapping on the input feature map to obtain a plurality of feature mapping maps, where the plurality of feature mapping maps respectively correspond to different mapping paths; Fuse the plurality of feature mapping maps to obtain a group of feature mapping maps; Input the group of feature mapping maps into the global average pooling layer to obtain global context information; Input the global context information into the first fully connected layer to obtain first channel weight value information; Input the first channel weight value information into the r-Softmax layer to obtain a plurality of groups of attention weight value information; Perform weight calibration on the plurality of feature mapping maps one by one according to the plurality of groups of attention weight value information to obtain a plurality of initial weight calibration feature maps; Fuse the plurality of initial weight calibration feature maps to obtain the weight calibration feature map.
5. The image classification method according to claim 3, characterized in that, The coordinate attention module includes a horizontal global average pooling layer, a vertical global average pooling layer, a second fully connected layer, and an activation function layer. Obtaining the position weight calibration feature map corresponding to the weight calibration feature map through the coordinate attention module includes: Inputting the weight calibration feature map into the horizontal global average pooling layer to obtain a horizontal perception attention map, and inputting the weight calibration feature map into the vertical global average pooling layer to obtain a vertical perception attention map; Inputting the horizontal perception attention map and the vertical perception attention map into the second fully connected layer to obtain second channel weight value information; Dividing the second channel weight value information into horizontal channel weight value information and vertical channel weight value information along the horizontal and vertical directions respectively; Inputting the horizontal channel weight value information and the vertical channel weight value information into the activation function layer respectively to obtain horizontal attention weight information corresponding to the horizontal channel weight value information and vertical attention weight information corresponding to the vertical channel weight value information; Determining the position weight value of each region on the weight calibration feature map according to the horizontal attention weight information and the vertical attention weight information to obtain the position weight calibration feature map.
6. The image classification method according to claim 1, wherein Inputting the position weight calibration feature map into the hierarchical multi-head attention module, and outputting the second feature map through the hierarchical multi-head attention module, including: Using the position weight calibration feature map as the input image of the first self-attention module in the first layer, and performing downsampling on the input image of the previous self-attention module to obtain the input image of the next self-attention module; Obtaining the attention feature maps output by each layer of the self-attention module to obtain a plurality of the attention feature maps; Fusing the plurality of attention feature maps to obtain the second feature map.
7. An image classification device, characterized in that, The device includes: An image input module, configured to obtain an image to be classified and input the image to be classified into a target classification model, where the target classification model includes a convolutional layer, an attention layer, and a classification layer; A local extraction module, configured to obtain local feature information of the image to be classified through the convolutional layer to obtain a first feature map; A global modeling module, configured to perform global modeling on the first feature map through the attention layer to obtain a second feature map, including: the attention layer includes a plurality of cascaded attention modules and a hierarchical multi-head attention module, and the hierarchical multi-head attention module includes a plurality of self-attention modules with a hierarchical relationship; inputting the first feature map into the first attention module, and obtaining the position weight calibration feature map output by the last attention module, where the position weight calibration feature map includes a plurality of regions, each region has a position weight value, and each position weight value is used to reflect the level of spatial attention and channel attention corresponding to a region; inputting the position weight calibration feature map into the hierarchical multi-head attention module, and outputting the second feature map through the hierarchical multi-head attention module; An image classification module, configured to perform image classification on the second feature map through the classification layer to obtain the image category corresponding to the image to be classified.
8. A terminal, characterized in that, The terminal includes a memory and one or more processors; the memory stores one or more programs; the programs include instructions for executing the image classification method according to any one of claims 1-6; the processor is configured to execute the programs.
9. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that, The instructions are adapted to be loaded and executed by a processor to implement the steps of the image classification method according to any one of claims 1-6 above.
Citation Information
Patent Citations
Multi-class image classification method and device, terminal equipment and storage medium
CN112651438A
Image classification method based on attention mechanism
CN113408577A