Image classification method and device and related equipment
Through the combination of a dual-channel encoder and a self-attention mechanism, the problem of low accuracy of fine-grained image classification in the prior art is solved, and higher image classification accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510608287.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-15
AI Technical Summary
The existing fine-grained image classification method cannot effectively learn the correlation between image feature points, resulting in low classification accuracy.
The dual-channel encoder module is used to divide the image data into image feature blocks of varying sizes, and a self-attention mechanism is introduced into the encoder. The relationship between different regions of the image is paid attention to through the dual-channel encoder architecture, and the feature extraction is performed using a multi-layer series-connected Transformer encoder.
It improves the accuracy and efficiency of fine-grained image classification and significantly improves the accuracy of image classification.
Smart Images

Figure CN120495770A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image classification method, apparatus, and related equipment. Background Art
[0002] Fine-grained image classification is a fundamental task in computer vision. Unlike simpler image classification tasks, fine-grained image classification datasets contain subtle differences that are difficult to distinguish. For example, the difference between a goldfinch and a green-throated vireo is so small, consisting solely of wing color, that common image classification models struggle to discern the difference.
[0003] Existing fine-grained image classification methods primarily utilize convolutional neural networks (CNNs) to extract image features, assist in model training to learn the characteristics of the target domain, and then import the trained weight file into the classification model to classify new image data. CNN-based classification detection models can only be trained on the image's own features and are unable to further learn the relationships between image feature points, making it difficult to achieve high classification performance and resulting in low accuracy for fine-grained image classification. Summary of the Invention
[0004] The present invention provides an image classification method for improving the accuracy of fine-grained image classification.
[0005] In a first aspect, the present application provides an image classification method, the method comprising:
[0006] Inputting the target image to be classified into a pre-trained image classification model; wherein the image classification model includes a dual-channel image segmentation module, a dual-channel encoder and a classifier;
[0007] Using the dual-channel image segmentation module, the target image is segmented into two sets of image blocks; wherein the sizes of the image blocks in the same set of image blocks are the same, and the sizes of any two image blocks in different sets of image blocks are different;
[0008] Using the dual-channel encoder to perform feature extraction on the two image block sets respectively to obtain two classification vectors of the target image, wherein an attention mechanism is added to each encoder in the dual-channel encoder;
[0009] The classifier is used to perform image classification based on two classification vectors to obtain the category of the target image.
[0010] This method uses a dual-channel encoder module, adding another encoder channel to the single-channel one. These two channels divide the image data into image feature blocks of varying sizes, thus avoiding the loss of edge semantics associated with single-channel segmentation. The self-attention mechanism within the dual-channel encoder architecture guides the model to focus on the relationships between different image regions, enabling higher-precision image classification and improving the accuracy of fine-grained image classification.
[0011] In one possible implementation, the image classification model further includes a dual-channel embedding layer; and before using the dual-channel encoder to perform feature extraction on the two image block sets respectively, the method further includes:
[0012] Using the dual-channel embedding layer to respectively embed the two image block sets, obtaining vector sets corresponding to the two image block sets, wherein the vector set corresponding to any one of the image block sets includes an initial classification vector and a corresponding position vector of each image block in the any one of the image block sets, and the position vector is used to represent the position of the image block in the target image;
[0013] The two image block sets are updated respectively using the vector sets corresponding to the two image block sets.
[0014] The above method adds an initial classification vector and a position vector in the embedding layer, locates the position of the image block in the original image through the position vector, and uses the classification vector in the encoder for the classification task of the classifier, thereby ensuring the accuracy of fine-grained image classification.
[0015] In a possible implementation, the encoder of any one channel of the dual-channel encoder includes multiple layers of serially connected transformer encoders;
[0016] The step of using the dual-channel encoder to extract features from the two image block sets to obtain a classification vector for the target image includes:
[0017] For any one layer of transformer encoders in the multi-layer transformer encoders in the first channel, use the transformer encoder to perform feature extraction on each input vector to obtain each intermediate classification vector, wherein the input vector is each initial classification vector in the image block set corresponding to the first channel or each intermediate classification vector obtained by the transformer encoder in the previous layer of the any one transformer encoder, and the first channel is any one channel of the dual channels;
[0018] Each intermediate classification vector obtained by the transformer encoder located at the last layer in the multi-layer transformer encoder is determined as the classification vector of the target image.
[0019] In the above method, feature extraction is performed on two sets of image blocks respectively through a multi-layer serially connected transformer encoder, which ensures that more fine-grained features can be obtained and further improves the accuracy of fine-grained image classification.
[0020] In one possible implementation, any one layer of the transformer encoder includes a first normalization layer and a multi-head attention mechanism layer; the input of the first normalization layer and the output of the multi-head attention mechanism layer are residually connected;
[0021] The transformer encoder is used to extract features from each input vector to obtain each intermediate classification vector, including:
[0022] Normalizing the input vectors using the first normalization layer to obtain first eigenvectors;
[0023] Utilizing the multi-head attention mechanism layer to calculate each first eigenvector, to obtain each second eigenvector;
[0024] The first eigenvectors and the second eigenvectors are added together to obtain intermediate classification vectors.
[0025] The above method adds a discriminative attention mechanism to the Transformer encoder to locate discriminative areas with subtle differences in the image, thereby obtaining high-quality fine-grained image features and ensuring the accuracy of fine-grained image classification.
[0026] In a possible implementation, if the transformer encoder of any one layer is the transformer encoder of the penultimate layer;
[0027] The multi-head attention mechanism layer is used to calculate each first eigenvector to obtain each second eigenvector, including:
[0028] Receive the second eigenvectors output by the transformer encoder of the previous layer; wherein the number of the second eigenvectors is the same as the number of attention mechanisms in the multi-head attention mechanism layer, each attention mechanism corresponds to one second eigenvector, the number of attention mechanisms corresponding to each transformer encoder layer is the same, and the number of feature parameters in the second eigenvector is the same as the number of image blocks in the image block set corresponding to the transformer encoder;
[0029] Obtain a target weight matrix based on the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the penultimate layer and the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the target layer, where the target layer is the layers before the penultimate layer and the weight matrix includes the weights corresponding to each attention mechanism in the multi-head attention mechanism layer;
[0030] For any attention mechanism in the multi-head attention mechanism layer, use the target weight corresponding to the attention mechanism to process the second target feature vector to obtain the attention weight of the second target feature vector, and determine the feature parameter with the largest value of the attention weight in the second target feature vector as the target feature parameter, wherein the second target feature vector is a second feature vector in which the identifier of the attention mechanism in each second feature vector output by the transformer encoder of the previous layer is the same as the identifier of the any attention mechanism, and the target weight corresponding to the attention mechanism is obtained based on the target weight matrix;
[0031] The eigenvector composed of the target feature parameters corresponding to each attention weight in the target weight matrix is determined as the second eigenvector of the transformer encoder of the penultimate layer.
[0032] The above method significantly reduces the amount of attention calculation by only taking the features of image blocks with larger attention weights as input in the last Transformer encoder layer, thereby increasing the calculation speed and improving the efficiency of fine-grained image classification.
[0033] In a possible implementation, before adding the first eigenvectors and the second eigenvectors to obtain the intermediate classification vectors, the method further includes:
[0034] For any first eigenvector, using the position vector corresponding to the first eigenvector, searching in the position vectors corresponding to the second eigenvectors to determine whether there is a target position vector identical to the position vector corresponding to the first eigenvector in the position vectors corresponding to the second eigenvectors;
[0035] If it exists, retain the first eigenvector;
[0036] If it does not exist, the first eigenvector is deleted.
[0037] The above method screens the first eigenvector to ensure that the screened first eigenvector can be fused with the second eigenvector, thereby ensuring that the image classification can be performed normally.
[0038] In a possible implementation, performing image classification based on two classification vectors using the classifier to obtain the category of the target image includes:
[0039] Perform feature fusion on the two classification vectors to obtain a target classification vector;
[0040] Normalizing the target classification vector to obtain a normalized target classification vector;
[0041] The normalized target classification vector is input into a classifier to obtain the category of the target image.
[0042] The above method first fuses the features of the two classification vectors to obtain the target classification vector, and then classifies the normalized target classification vector, ensuring that a classification vector with edge semantic information of different image feature blocks is obtained. Classification is performed using this classification vector, thereby improving the accuracy of image classification.
[0043] In a second aspect, the present application provides an image classification device, the device comprising:
[0044] An input module, configured to input a target image to be classified into a pre-trained image classification model; wherein the image classification model comprises a dual-channel image segmentation module, a dual-channel encoder, and a classifier;
[0045] an image segmentation module, configured to segment the target image into two sets of image blocks using the dual-channel image segmentation module; wherein the image blocks in the same set of image blocks have the same size, and any two image blocks in different sets of image blocks have different sizes;
[0046] a feature extraction module, configured to perform feature extraction on the two image block sets respectively using the dual-channel encoder to obtain two classification vectors of the target image, wherein the dual-channel encoder includes an attention mechanism;
[0047] The classification module is used to perform image classification based on two classification vectors using the classifier to obtain the category of the target image.
[0048] In a possible implementation, the device further includes:
[0049] An embedding module, wherein the image classification model further includes a dual-channel embedding layer; before using the dual-channel encoder to perform feature extraction on the two image block sets, the dual-channel embedding layer is used to perform embedding operations on the two image block sets to obtain vector sets corresponding to the two image block sets, wherein the vector set corresponding to any one of the image block sets includes an initial classification vector and a corresponding position vector for each image block in the image block set, and the position vector is used to represent the position of the image block in the target image;
[0050] An updating module is configured to update the two image block sets respectively by using the vector sets corresponding to the two image block sets.
[0051] In a possible implementation, the encoder of any one channel of the dual-channel encoder includes multiple layers of serially connected transformer encoders;
[0052] The feature extraction module is specifically used to:
[0053] For any one layer of transformer encoders in the multi-layer transformer encoders in the first channel, use the transformer encoder to perform feature extraction on each input vector to obtain each intermediate classification vector, wherein the input vector is each initial classification vector in the image block set corresponding to the first channel or each intermediate classification vector obtained by the transformer encoder in the previous layer of the any one transformer encoder, and the first channel is any one channel of the dual channels;
[0054] Each intermediate classification vector obtained by the transformer encoder located at the last layer in the multi-layer transformer encoder is determined as the classification vector of the target image.
[0055] In one possible implementation, any one layer of the transformer encoder includes a first normalization layer and a multi-head attention mechanism layer; the input of the first normalization layer and the output of the multi-head attention mechanism layer are residually connected;
[0056] The feature extraction module is further used to:
[0057] Normalizing the input vectors using the first normalization layer to obtain first eigenvectors;
[0058] Utilizing the multi-head attention mechanism layer to calculate each first eigenvector, to obtain each second eigenvector;
[0059] The first eigenvectors and the second eigenvectors are added together to obtain intermediate classification vectors.
[0060] In a possible implementation, if the transformer encoder of any one layer is the transformer encoder of the penultimate layer;
[0061] The feature extraction module is further used to:
[0062] Receive the second eigenvectors output by the transformer encoder of the previous layer; wherein the number of the second eigenvectors is the same as the number of attention mechanisms in the multi-head attention mechanism layer, each attention mechanism corresponds to one second eigenvector, the number of attention mechanisms corresponding to each transformer encoder layer is the same, and the number of feature parameters in the second eigenvector is the same as the number of image blocks in the image block set corresponding to the transformer encoder;
[0063] Obtain a target weight matrix based on the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the penultimate layer and the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the target layer, where the target layer is the layers before the penultimate layer and the weight matrix includes the weights corresponding to each attention mechanism in the multi-head attention mechanism layer;
[0064] For any attention mechanism in the multi-head attention mechanism layer, use the target weight corresponding to the attention mechanism to process the second target feature vector to obtain the attention weight of the second target feature vector, and determine the feature parameter with the largest value of the attention weight in the second target feature vector as the target feature parameter, wherein the second target feature vector is a second feature vector in which the identifier of the attention mechanism in each second feature vector output by the transformer encoder of the previous layer is the same as the identifier of the any attention mechanism, and the target weight corresponding to the attention mechanism is obtained based on the target weight matrix;
[0065] The eigenvector composed of the target feature parameters corresponding to each attention weight in the target weight matrix is determined as the second eigenvector of the transformer encoder of the penultimate layer.
[0066] In a possible implementation, the device further includes:
[0067] a search module configured to, before adding the first eigenvectors and the second eigenvectors to obtain the intermediate classification vectors, search, for any first eigenvector, using the position vector corresponding to the first eigenvector, the position vectors corresponding to the second eigenvectors to determine whether there is a target position vector identical to the position vector corresponding to the first eigenvector in the position vectors corresponding to the second eigenvectors;
[0068] a retaining module, configured to retain the first feature vector if it exists;
[0069] A deleting module is used to delete the first feature vector if it does not exist.
[0070] In a possible implementation, the classification module is specifically configured to:
[0071] Perform feature fusion on the two classification vectors to obtain a target classification vector;
[0072] Normalizing the target classification vector to obtain a normalized target classification vector;
[0073] The normalized target classification vector is input into a classifier to obtain the category of the target image.
[0074] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor implements the steps in the image classification method.
[0075] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-mentioned image classification method of the present application.
[0076] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which is stored in a computer-readable storage medium; when the processor of a memory access device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the memory access device to execute the steps in the above-mentioned image classification method of the present application.
[0077] For each aspect from the second to the fifth aspect and the technical effects that may be achieved by each aspect, please refer to the above description of the technical effects that can be achieved by various possible solutions in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0079] Figure 1 This is a flow chart of an image classification method according to an embodiment of the present application;
[0080] Figure 2 A schematic diagram of the structure of an image classification model provided in an embodiment of the present application;
[0081] Figure 3 A schematic diagram of a process for extracting features from an image block set according to an embodiment of the present application;
[0082] Figure 4 The embodiment of the present application provides a schematic diagram of the structure of a transformer encoder of any layer;
[0083] Figure 5 A schematic diagram of a flow chart of obtaining a second feature vector for a transformer encoder of the penultimate layer provided in an embodiment of the present application;
[0084] Figure 6 A schematic diagram illustrating the accuracy of image classification of each model provided in the embodiments of the present application;
[0085] Figure 7 A schematic diagram of an ablation experiment of an image classification model provided in an embodiment of the present application;
[0086] Figure 8 The second flowchart of the image classification method provided in the embodiment of the present application;
[0087] Figure 9 A schematic diagram of an image classification device provided in an embodiment of the present application;
[0088] Figure 10 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0089] In order to make the purpose, technical solutions and advantages of this application more clear, the application will be further described in detail below with reference to the accompanying drawings. The specific operation methods in the method embodiments can also be applied to the device embodiments or system embodiments.
[0090] In the description of this application, "multiple" is understood to mean "at least two." "And / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. A and B are connected, which can mean: A and B are directly connected, and A and B are connected through C. In addition, in the description of this application, words such as "first" and "second" are used only for the purpose of distinguishing descriptions and should not be understood as indicating or implying relative importance or order.
[0091] Existing fine-grained image classification methods primarily utilize convolutional neural networks to extract image features, assist in model training to learn the characteristics of the target domain, and then import the trained weight file into the classification model to classify new image data. CNN-based classification detection models can only be trained based on the image's inherent features and are unable to further learn the relationships between image feature points, making it difficult to achieve high classification performance and resulting in low accuracy for fine-grained image classification.
[0092] To address this issue, an embodiment of the present application provides an image classification method that uses a dual-channel encoder module to add another encoder channel to a single-channel one. The two channels divide the image data into image feature blocks of varying sizes, thereby avoiding the loss of edge semantics of the image feature blocks caused by single-channel division. Because the self-attention mechanism in the dual-channel encoder architecture can guide the model to focus on the relationship between different areas in the image, the model can achieve higher-precision image classification, improving the accuracy of fine-grained image classification.
[0093] The present application is described in further detail below with reference to the accompanying drawings. Figure 1 FIG. 1 is a flow chart of an image classification method provided in an embodiment of the present application. The specific implementation process of the method is as follows:
[0094] Step 101: Inputting a target image to be classified into a pre-trained image classification model; wherein the image classification model includes a dual-channel image segmentation module, a dual-channel encoder and a classifier;
[0095] like Figure 2 As shown, Figure 2 This is a structural diagram of the image classification model. Figure 2 As can be seen in FIG, the image classification model 200 includes a dual-channel image segmentation module 201, a dual-channel embedding layer 202, a dual-channel encoder 203 and a classifier 204.
[0096] In order to ensure the accuracy of image classification, in one possible implementation, the dual-channel embedding layer 202 is used to perform embedding operations on the two image block sets respectively to obtain vector sets corresponding to the two image block sets, wherein the vector set corresponding to any one of the image block sets includes the initial classification vector and the corresponding position vector of each image block in the any one of the image block sets, and the position vector is used to represent the position of the image block in the target image; the two image block sets are updated respectively using the vector sets corresponding to the two image block sets.
[0097] Any image block set in the embodiments of the present application includes each image block, a position vector corresponding to each image block, and an initial classification vector corresponding to each image block.
[0098] Step 102: using the dual-channel image segmentation module to segment the target image into two sets of image blocks; wherein the sizes of the image blocks in the same set of image blocks are the same, and the sizes of any two image blocks in different sets of image blocks are different;
[0099] In the embodiment of the present application, the size of each image block in one image block set is A*A, and the size of each image block in another image block set is (A+b)×(A+b) ,in b It is a hyperparameter of the image classification model and can be set manually before training the image classification model.
[0100] It should be noted that the values of the size A and b of the image block in the embodiment of the present application can be set according to actual conditions, and the embodiment of the present application does not limit the specific values of A and b.
[0101] Step 103: performing feature extraction on the two image block sets respectively using the dual-channel encoder to obtain two classification vectors of the target image, wherein an attention mechanism is added to each encoder in the dual-channel encoder;
[0102] The encoder of any one channel of the dual-channel encoder in the embodiment of the present application includes a transformer encoder with multiple layers connected in series. The following describes the specific method of extracting features from the image block set in step 103. Figure 3 FIG. 1 is a flow chart showing a process of extracting features from an image block set, which may include the following steps:
[0103] Step 301: for any one of the multiple layers of transformer encoders in a first channel, use the transformer encoder to perform feature extraction on each input vector to obtain each intermediate classification vector, wherein the input vector is each initial classification vector in the image block set corresponding to the first channel or each intermediate classification vector obtained by the transformer encoder of the previous layer of any one transformer encoder, and the first channel is any one of the dual channels;
[0104] Step 302: Determine each intermediate classification vector obtained by the transformer encoder located at the last layer in the multi-layer transformer encoder as the classification vector of the target image.
[0105] It should be noted that the number of layers of the transformer encoder is not limited in the embodiment of the present application. The number of layers of each transformer encoder in the embodiment of the present application can be set according to specific actual conditions.
[0106] like Figure 4 As shown in the figure, it is a schematic diagram of the structure of any layer of transformer encoder. Figure 4 As can be seen, the transformer encoder of any layer includes a first normalization layer 401 and a multi-head attention mechanism layer 402; the input of the first normalization layer 401 and the output of the multi-head attention mechanism layer 402 are residually connected. The following describes the specific method of using the transformer encoder to extract features from each input vector in combination with the structure of the transformer encoder:
[0107] The first normalization layer 401 is used to normalize the input vectors to obtain first eigenvectors; the multi-head attention mechanism layer 402 is used to calculate the first eigenvectors to obtain second eigenvectors; the first eigenvectors and the second eigenvectors are added to obtain intermediate classification vectors. The intermediate classification vectors can be obtained by formula (1):
[0108] Z l =MSA(LN(Z l-1 ))+Z l-1 …(1);
[0109] Among them, Z l is the intermediate classification vector obtained by the l-th layer transformer encoder, Z l-1The intermediate classification vector output by the l-1 layer transformer encoder, LN is the normalization operation, and MSA represents the multi-head attention mechanism operation.
[0110] In order to improve the efficiency of image classification, the multi-head attention mechanism layer 402 in the embodiment of the present application calculates each first eigenvector to obtain each second eigenvector in two ways. Among them, the way the transformer encoder of the penultimate layer obtains the second eigenvector is different from the way the transformer encoders of other layers obtain the second eigenvector. The following is a detailed description:
[0111] like Figure 5 As shown in FIG, it is a flow chart of the second feature vector obtained by the transformer encoder of the penultimate layer, which may specifically include the following steps:
[0112] Step 501: Receive each second eigenvector output by the transformer encoder of the previous layer; wherein the number of the second eigenvectors is the same as the number of attention mechanisms in the multi-head attention mechanism layer, each attention mechanism corresponds to one second eigenvector, the number of attention mechanisms corresponding to each transformer encoder layer is the same, and the number of feature parameters in the second eigenvector is the same as the number of image blocks in the image block set corresponding to the transformer encoder;
[0113] For example, if each layer of the transformer encoder includes K attention mechanisms, the number of second eigenvectors is K. If the number of image blocks in the embodiment of the present application is N, the number of feature parameters in any second eigenvector in the embodiment of the present application is N.
[0114] It should be noted that the specific value of K is not limited in the embodiments of the present application, and the specific value of K in the embodiments of the present application can be set according to specific actual conditions.
[0115] Step 502: Obtain a target weight matrix based on the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the penultimate layer and the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the target layer, wherein the target layer is each layer located before the penultimate layer, and the weight matrix includes the weights corresponding to each attention mechanism in the multi-head attention mechanism layer;
[0116] For example, the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of layer l is
[0117] In one possible implementation, step 502 may be specifically implemented as follows: multiplying each weight matrix in sequence according to the number of layers of the encoder to obtain the target weight matrix, wherein each weight matrix includes the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the penultimate layer and the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the target layer; wherein the target weight matrix can be obtained by formula (2):
[0118]
[0119] Among them, a final is the target weight matrix, a l is the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of layer l, and m is the total number of layers in the transformer encoder. m is a positive integer.
[0120] It should be noted that the embodiment of the present application does not limit the specific value of m. The specific value of m in the embodiment of the present application can be set according to specific actual conditions.
[0121] Step 503: For any attention mechanism in the multi-head attention mechanism layer, use the target weight corresponding to the attention mechanism to process the second target feature vector to obtain the attention weight of the second target feature vector, and determine the feature parameter with the largest value of the attention weight in the second target feature vector as the target feature parameter, wherein the second target feature vector is the second feature vector in which the identifier of the attention mechanism in each second feature vector output by the transformer encoder of the previous layer is the same as the identifier of the any attention mechanism, and the target weight corresponding to the attention mechanism is obtained based on the target weight matrix;
[0122] The number of target weights in the target weight matrix obtained in the embodiment of the present application is also k, that is, is the target weight corresponding to the k-th multi-head attention mechanism. In the embodiment of the present application, the identifier corresponding to each second eigenvector is also 1 to k.
[0123] In the embodiment of the present application, the target weight corresponding to the attention mechanism is used to process the second target feature vector to obtain the attention weight of the second target feature vector, which belongs to the prior art and is not described in detail in the embodiment of the present application.
[0124] Step 504: Determine the eigenvector composed of the target feature parameters corresponding to each attention weight in the target weight matrix as the second eigenvector of the transformer encoder of the penultimate layer.
[0125] In the embodiment of the present application, each target feature parameter is arranged in ascending order of identification to form the second feature vector.
[0126] Therefore, by taking only the features of image blocks with larger attention weights as input in the last Transformer encoder layer, the amount of attention calculation is greatly reduced, thereby increasing the calculation speed and improving the efficiency of fine-grained image classification.
[0127] The transformer encoders of other layers in the embodiment of the present application have the same process as the transformer encoder in the prior art, that is, the feature parameters of the second feature vector are not screened. Since this does not belong to the invention point of the present application, the embodiment of the present application will not be described in detail here.
[0128] Since the second eigenvectors in the penultimate layer are screened, in order to ensure that image classification can proceed smoothly, in one possible implementation, for any first eigenvector, the position vector corresponding to the first eigenvector is used to search among the position vectors corresponding to the second eigenvectors to determine whether there is a target position vector identical to the position vector corresponding to the first eigenvector in the position vectors corresponding to the second eigenvectors; if so, the first eigenvector is retained; if not, the first eigenvector is deleted.
[0129] In this way, it can be ensured that the number of second eigenvectors input to the last layer is the same as the number of first eigenvectors, ensuring smooth image classification.
[0130] Step 104: Utilize the classifier to perform image classification based on the two classification vectors to obtain the category of the target image.
[0131] In one possible implementation, step 104 can be specifically implemented as follows: performing feature fusion on the two classification vectors to obtain a target classification vector; normalizing the target classification vector to obtain a normalized target classification vector; and inputting the normalized target classification vector into a classifier to obtain the category of the target image. The target classification vector can be obtained by formula (3):
[0132] σ=ω1β1+ω2β2……(3);
[0133] Wherein, σ is the target classification vector, β1 is one of the two classification vectors, β2 is the other of the two classification vectors, ω1 is the fusion weight of one of the classification vectors, and ω2 is the fusion weight of the other classification vector.
[0134] It should be noted that the specific values of the two fusion weights in the embodiment of the present application can be set according to the specific actual situation, and the embodiment of the present application does not limit the specific values of the two fusion weights.
[0135] The classifier used in the embodiment of the present application is an MLP classifier, but the embodiment of the present application does not limit the classifier. The classifier in the embodiment of the present application can be set according to the specific actual situation. Among them, the category of the target image can be obtained by formula (4):
[0136] P = MLP(LN(σ))…(4);
[0137] Wherein, P is the category of the target image, MLP() is the MLP classifier, and LN() is the normalization layer.
[0138] like Figure 6 As shown in the table, using the fine-grained bird classification dataset as an example, the proposed image classification model significantly outperforms the currently popular CNN-based image classification model, ResNet101. The image classification accuracy improves from 84.9% to 92.1%. The Transformer-based image classification model outperforms the CNN-based model on the three datasets. This demonstrates the powerful self-attention mechanism of the Transformer, which effectively learns the relationships between image feature patches, making it more suitable for fine-grained image classification. Compared to the baseline model, ViT, the proposed image classification model achieves a 1.8% increase in image classification accuracy. This demonstrates that the proposed dual-channel encoder module effectively captures semantic information at the edges of image feature patches, thereby improving model performance. Compared to all other methods, the proposed model achieves state-of-the-art performance of 95.2% and 92.9% on the fine-grained car classification dataset and fine-grained dog classification dataset, respectively. These image classification accuracies surpass those of existing mainstream image classification models and baseline models, further demonstrating the effectiveness of the proposed innovation.
[0139] This paper conducts ablation experiments on a fine-grained bird classification dataset to verify the impact of different modules on model performance. Figure 7As shown in the figure, when the image classification model does not use the dual-channel encoder, attention mechanism, and classification vector feature fusion strategy, the model's classification accuracy is the lowest, at 90.3%. When the dual-channel encoder is added, the model's classification accuracy increases to 91.1%, indicating that the model learns the missing edge semantic information of image feature blocks of different sizes, thereby improving classification accuracy. When only the attention mechanism is added, the model's classification accuracy increases to 91.6%, indicating that the model now learns fine-grained image features. When all three modules are used together, the model's performance reaches the best 92.1%.
[0140] Next, combine Figure 8 The image classification method in the embodiment of the present application is described, which may specifically include the following steps:
[0141] Step 801: Inputting a target image to be classified into a pre-trained image classification model; wherein the image classification model includes a dual-channel image segmentation module, a dual-channel encoder, a dual-channel embedding layer and a classifier;
[0142] Step 802: Using the dual-channel image segmentation module, segment the target image into two sets of image blocks; wherein the image blocks in the same set of image blocks have the same size, and any two image blocks in different sets of image blocks have different sizes;
[0143] Step 803: using the dual-channel embedding layer to perform embedding operations on the two image block sets respectively, to obtain vector sets corresponding to the two image block sets, wherein the vector set corresponding to any one of the image block sets includes the initial classification vector and the corresponding position vector of each image block in the any one of the image block sets, and the position vector is used to represent the position of the image block in the target image; using the vector sets corresponding to the two image block sets respectively, the two image block sets are updated;
[0144] Step 804: performing feature extraction on the two image block sets respectively using the dual-channel encoder to obtain two classification vectors of the target image, wherein an attention mechanism is added to each encoder in the dual-channel encoder;
[0145] Step 805: performing feature fusion on the two classification vectors to obtain a target classification vector;
[0146] Step 806: normalizing the target classification vector to obtain a normalized target classification vector;
[0147] Step 807: Input the normalized target classification vector into a classifier to obtain the category of the target image.
[0148] Based on the same inventive concept, the present application also provides an image classification device, such as Figure 9 As shown, the device includes:
[0149] Input module 901, used to input the target image to be classified into a pre-trained image classification model; wherein the image classification model includes a dual-channel image segmentation module, a dual-channel encoder and a classifier;
[0150] An image segmentation module 902 is configured to segment the target image into two sets of image blocks using the dual-channel image segmentation module; wherein the image blocks in the same set of image blocks have the same size, and any two image blocks in different sets of image blocks have different sizes;
[0151] a feature extraction module 903 configured to perform feature extraction on the two image block sets respectively using the dual-channel encoder to obtain two classification vectors of the target image, wherein the dual-channel encoder includes an attention mechanism;
[0152] The classification module 904 is configured to perform image classification based on two classification vectors using the classifier to obtain a category of the target image.
[0153] In a possible implementation, the device further includes:
[0154] Embedding module 905, for the image classification model further comprising a dual-channel embedding layer; before using the dual-channel encoder to perform feature extraction on the two image block sets respectively, using the dual-channel embedding layer to perform embedding operations on the two image block sets respectively to obtain vector sets corresponding to the two image block sets respectively, wherein the vector set corresponding to any image block set includes the initial classification vector and the corresponding position vector of each image block in the any image block set, and the position vector is used to represent the position of the image block in the target image;
[0155] The updating module 906 is configured to update the two image block sets respectively by using the vector sets corresponding to the two image block sets.
[0156] In a possible implementation, the encoder of any one channel of the dual-channel encoder includes multiple layers of serially connected transformer encoders;
[0157] The feature extraction module 903 is specifically used to:
[0158] For any one layer of transformer encoders in the multi-layer transformer encoders in the first channel, use the transformer encoder to perform feature extraction on each input vector to obtain each intermediate classification vector, wherein the input vector is each initial classification vector in the image block set corresponding to the first channel or each intermediate classification vector obtained by the transformer encoder in the previous layer of the any one transformer encoder, and the first channel is any one channel of the dual channels;
[0159] Each intermediate classification vector obtained by the transformer encoder located at the last layer in the multi-layer transformer encoder is determined as the classification vector of the target image.
[0160] In one possible implementation, any one layer of the transformer encoder includes a first normalization layer and a multi-head attention mechanism layer; the input of the first normalization layer and the output of the multi-head attention mechanism layer are residually connected;
[0161] The feature extraction module 903 is further configured to:
[0162] Normalizing the input vectors using the first normalization layer to obtain first eigenvectors;
[0163] Utilizing the multi-head attention mechanism layer to calculate each first eigenvector, to obtain each second eigenvector;
[0164] The first eigenvectors and the second eigenvectors are added together to obtain intermediate classification vectors.
[0165] In a possible implementation, if the transformer encoder of any one layer is the transformer encoder of the penultimate layer;
[0166] The feature extraction module 903 is further configured to:
[0167] Receive the second eigenvectors output by the transformer encoder of the previous layer; wherein the number of the second eigenvectors is the same as the number of attention mechanisms in the multi-head attention mechanism layer, each attention mechanism corresponds to one second eigenvector, the number of attention mechanisms corresponding to each transformer encoder layer is the same, and the number of feature parameters in the second eigenvector is the same as the number of image blocks in the image block set corresponding to the transformer encoder;
[0168] Obtain a target weight matrix based on the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the penultimate layer and the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the target layer, where the target layer is the layers before the penultimate layer and the weight matrix includes the weights corresponding to each attention mechanism in the multi-head attention mechanism layer;
[0169] For any attention mechanism in the multi-head attention mechanism layer, use the target weight corresponding to the attention mechanism to process the second target feature vector to obtain the attention weight of the second target feature vector, and determine the feature parameter with the largest value of the attention weight in the second target feature vector as the target feature parameter, wherein the second target feature vector is a second feature vector in which the identifier of the attention mechanism in each second feature vector output by the transformer encoder of the previous layer is the same as the identifier of the any attention mechanism, and the target weight corresponding to the attention mechanism is obtained based on the target weight matrix;
[0170] The eigenvector composed of the target feature parameters corresponding to each attention weight in the target weight matrix is determined as the second eigenvector of the transformer encoder of the penultimate layer.
[0171] In a possible implementation, the device further includes:
[0172] A search module 907 is configured to, before adding the first eigenvectors and the second eigenvectors to obtain the intermediate classification vectors, search, for any first eigenvector, using the position vector corresponding to the first eigenvector, the position vectors corresponding to the second eigenvectors to determine whether there is a target position vector identical to the position vector corresponding to the first eigenvector in the position vectors corresponding to the second eigenvectors;
[0173] A retaining module 908 is configured to retain the first feature vector if it exists;
[0174] The deleting module 909 is configured to delete the first feature vector if it does not exist.
[0175] In a possible implementation, the classification module 904 is specifically configured to:
[0176] Perform feature fusion on the two classification vectors to obtain a target classification vector;
[0177] Normalizing the target classification vector to obtain a normalized target classification vector;
[0178] The normalized target classification vector is input into a classifier to obtain the category of the target image.
[0179] Based on the same inventive concept, an electronic device is also provided in the embodiment of the present application, and the electronic device can realize the function of the aforementioned image classification device, referring to Figure 10 , the electronic device includes:
[0180] At least one processor 1001, and a memory 1002 connected to the at least one processor 1001. The specific connection medium between the processor 1001 and the memory 1002 is not limited in the embodiment of the present application. Figure 10 In the example, the processor 1001 and the memory 1002 are connected via the bus 1000. Figure 10 The bus 1000 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 The diagram is represented by only one thick line, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 1001 may also be referred to as a controller, without limitation to the name.
[0181] In the embodiment of the present application, the memory 1002 stores instructions that can be executed by at least one processor 1001. The at least one processor 1001 can execute the image classification method discussed above by executing the instructions stored in the memory 1002. The processor 1001 can implement Figure 9 The functions of each module in the device shown.
[0182] Among them, the processor 1001 is the control center of the device, which can use various interfaces and lines to connect the various parts of the entire control device, and monitor the device as a whole by running or executing instructions stored in the memory 1002 and calling data stored in the memory 1002, the various functions of the device and processing data.
[0183] In one possible design, processor 1001 may include one or more processing units. Processor 1001 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1001. In some embodiments, processor 1001 and memory 1002 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.
[0184] Processor 1001 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the image classification method disclosed in the embodiments of this application can be directly implemented as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.
[0185] Memory 1002 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. Memory 1002 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. Memory 1002 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 1002 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0186] By designing and programming the processor 1001, the code corresponding to the image classification method described in the above embodiment can be fixed into the chip, so that the chip can execute the image classification method when it is running. Figure 1 The steps of the image classification method of the embodiment shown are as follows: How to design and program the processor 1001 is a technique well known to those skilled in the art and will not be described in detail here.
[0187] An embodiment of the present application also provides a computer-readable storage medium that stores computer-executable instructions required to execute the above-mentioned processor, which includes a program required to execute the above-mentioned processor.
[0188] In some possible embodiments, various aspects of the image classification method provided in the present application can also be implemented in the form of a program product, which includes program code. When the above-mentioned program product is run on an electronic device, the above-mentioned program code is used to enable the above-mentioned electronic device to execute the steps of the image classification method according to various exemplary embodiments of the present application described above in this specification.
[0189] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0190] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0191] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0193] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0194] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A method for classifying an image, characterized in that: The method comprises: Inputting the target image to be classified into a pre-trained image classification model; wherein the image classification model includes a dual-channel image segmentation module, a dual-channel encoder and a classifier; Using the dual-channel image segmentation module, the target image is segmented into two sets of image blocks; wherein the sizes of the image blocks in the same set of image blocks are the same, and the sizes of any two image blocks in different sets of image blocks are different; Using the dual-channel encoder to perform feature extraction on the two image block sets respectively to obtain two classification vectors of the target image, wherein an attention mechanism is added to each encoder in the dual-channel encoder; The classifier is used to perform image classification based on two classification vectors to obtain the category of the target image.
2. The method according to claim 1, characterized in that The image classification model further includes a dual-channel embedding layer; and before performing feature extraction on the two image block sets using the dual-channel encoder, the method further includes: Using the dual-channel embedding layer to respectively embed the two image block sets, obtaining vector sets corresponding to the two image block sets, wherein the vector set corresponding to any one of the image block sets includes an initial classification vector and a corresponding position vector of each image block in the any one of the image block sets, and the position vector is used to represent the position of the image block in the target image; The two image block sets are updated respectively using the vector sets corresponding to the two image block sets.
3. The method according to claim 1, characterized in that The encoder of any one channel of the dual-channel encoder includes multiple layers of transformer encoders connected in series; The step of using the dual-channel encoder to extract features from the two image block sets to obtain a classification vector for the target image includes: For any one layer of transformer encoders in the multi-layer transformer encoders in the first channel, use the transformer encoder to perform feature extraction on each input vector to obtain each intermediate classification vector, wherein the input vector is each initial classification vector in the image block set corresponding to the first channel or each intermediate classification vector obtained by the transformer encoder in the previous layer of the any one transformer encoder, and the first channel is any one channel of the dual channels; Each intermediate classification vector obtained by the transformer encoder located at the last layer in the multi-layer transformer encoder is determined as the classification vector of the target image.
4. The method according to claim 3, characterized in that Any one layer of transformer encoder includes a first normalization layer and a multi-head attention mechanism layer; the input of the first normalization layer and the output of the multi-head attention mechanism layer are residually connected; The transformer encoder is used to extract features from each input vector to obtain each intermediate classification vector, including: Normalizing the input vectors using the first normalization layer to obtain first eigenvectors; Utilizing the multi-head attention mechanism layer to calculate each first eigenvector, to obtain each second eigenvector; The first eigenvectors and the second eigenvectors are added together to obtain intermediate classification vectors.
5. The method according to claim 4, characterized in that If any one layer of transformer encoder is the transformer encoder of the penultimate layer; The multi-head attention mechanism layer is used to calculate each first eigenvector to obtain each second eigenvector, including: Receive the second eigenvectors output by the transformer encoder of the previous layer; wherein the number of the second eigenvectors is the same as the number of attention mechanisms in the multi-head attention mechanism layer, each attention mechanism corresponds to one second eigenvector, the number of attention mechanisms corresponding to each transformer encoder layer is the same, and the number of feature parameters in the second eigenvector is the same as the number of image blocks in the image block set corresponding to the transformer encoder; Obtain a target weight matrix based on the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the penultimate layer and the weight matrix corresponding to the multi-head attention mechanism layer in the transformer encoder of the target layer, where the target layer is the layers before the penultimate layer and the weight matrix includes the weights corresponding to each attention mechanism in the multi-head attention mechanism layer; For any attention mechanism in the multi-head attention mechanism layer, use the target weight corresponding to the attention mechanism to process the second target feature vector to obtain the attention weight of the second target feature vector, and determine the feature parameter with the largest value of the attention weight in the second target feature vector as the target feature parameter, wherein the second target feature vector is a second feature vector in which the identifier of the attention mechanism in each second feature vector output by the transformer encoder of the previous layer is the same as the identifier of the any attention mechanism, and the target weight corresponding to the attention mechanism is obtained based on the target weight matrix; The eigenvector composed of the target feature parameters corresponding to each attention weight in the target weight matrix is determined as the second eigenvector of the transformer encoder of the penultimate layer.
6. The method according to claim 5, characterized in that Before adding the first eigenvectors and the second eigenvectors to obtain the intermediate classification vectors, the method further includes: For any first eigenvector, using the position vector corresponding to the first eigenvector, searching in the position vectors corresponding to the second eigenvectors to determine whether there is a target position vector identical to the position vector corresponding to the first eigenvector in the position vectors corresponding to the second eigenvectors; If it exists, retain the first eigenvector; If it does not exist, the first eigenvector is deleted.
7. The method according to claim 1, characterized in that The classifying the image using the classifier based on the two classification vectors to obtain the category of the target image includes: Perform feature fusion on the two classification vectors to obtain a target classification vector; Normalizing the target classification vector to obtain a normalized target classification vector; The normalized target classification vector is input into a classifier to obtain the category of the target image.
8. An image classification device, characterized in that: The device comprises: An input module, configured to input a target image to be classified into a pre-trained image classification model; wherein the image classification model comprises a dual-channel image segmentation module, a dual-channel encoder, and a classifier; an image segmentation module, configured to segment the target image into two sets of image blocks using the dual-channel image segmentation module; wherein the image blocks in the same set of image blocks have the same size, and any two image blocks in different sets of image blocks have different sizes; a feature extraction module, configured to perform feature extraction on the two sets of image blocks using the dual-channel encoder to obtain two classification vectors of the target image, wherein an attention mechanism is added to each encoder in the dual-channel encoder; The classification module is used to perform image classification based on two classification vectors using the classifier to obtain the category of the target image.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 7 when executing the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.