Lightweight Non-Local Based Image Classification Method and System
By grouping abstract features in lightweight networks and simplifying the acquisition method of Non-Local long connection relationship, the problem of excessive computational complexity in image classification is solved, and the accuracy of image classification and the control of calculation amount is achieved.
Patent Information
- Application Number
- CN202210131553.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-02-14
AI Technical Summary
Although the Non-Local self-attention mechanism can enhance feature expression capabilities in image classification, due to the large amount of calculation, the computational complexity of the image classification method is too high.
By grouping abstract features in lightweight networks, the Non-Local long connection relationship acquisition method is simplified, and the features are extracted using unary long connections and point-to-long connection relationships to reduce the computational complexity.
While maintaining the Non-Local long connection performance, the calculation complexity is significantly reduced, and the image classification accuracy is improved and the calculation amount is controlled.
Smart Images

Figure CN114492659B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image classification, and in particular to an image classification method and system based on lightweight Non-Local. Background Art
[0002] The Non-Local self-attention mechanism was proposed by Xiaolong Wang, Kaiming He, etc. at CVPR2018 (the top conference in computer vision). The Non-Local method utilizes the feature similarity between image pixels to establish a long connection relationship between pixels, which is used to enhance the feature expression ability of each pixel in the image. This relationship does not have the locality limitation of convolution in neural networks. That is to say, no matter how far the pixels are from each other, the Non-Local mechanism can explore the long connection relationship. In other words, the Non-Local method has a field ability that convolution cannot achieve.
[0003] Since the Non-Local self-attention mechanism was proposed, it has been applied in classical deep learning convolutional neural networks and has been verified to be able to well enhance the feature expressiveness of the network. As a means to make up for the too small receptive field of convolution, generally, the Non-Local method combined with classical networks can bring a general performance improvement in different task fields and has universality. The fields where Non-Local is applied include the three basic tasks of computer vision: image classification, object detection, and semantic segmentation. It also includes various specific small fields. For example, the paper by Xiaolong Wang itself uses 3D Non-Local self-attention to process video classification problems. In semantic segmentation, once Non-Local is adopted, the performance far exceeds the form of using pyramids to expand the convolution receptive field in the past.
[0004] However, while Non-Local brings strong feature expression ability, it also has many defects, such as a high computational cost. Therefore, the image classification method using the Non-Local method also has the problem of a large computational amount. Based on this, the present invention provides an image classification method and system based on lightweight Non-Local. This method can not only enhance the feature expression ability of pixels to improve the classification accuracy but also avoid a large increase in the computational amount. Summary of the Invention
[0005] The object of the present invention is to provide an image classification method and system based on lightweight Non-Local, which simplifies the long connection relationship acquisition mechanism of the original Non-Local by using the semantic characteristics of grouped feature implicit clustering in the lightweight network, thereby reducing the computational complexity, so that the image classification method using the Non-Local method can enhance the feature expression ability of the lightweight network and then improve the image classification accuracy while avoiding a large increase in the network computational amount.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] An image classification method based on lightweight Non-Local, the method comprising:
[0008] Inputting a sample image into a lightweight network, and obtaining abstract features after each bottleneck layer of the lightweight network, the lightweight network including ShufflenetV2;
[0009] Dividing the abstract features along the channel dimension into G groups;
[0010] For the abstract features of each group, obtaining a spatial weight through convolution, the spatial weight being global spatial attention information for the current group;
[0011] Using the global spatial attention information of the current group to perform weighted amplitude conversion on the pixel points of the current group to obtain the unary long connection feature of the current group;
[0012] Using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic aggregation weight, and performing global weighted pooling on the current group according to the semantic aggregation weight to obtain the semantic information of the current group;
[0013] Using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic assignment weight for the current group;
[0014] Combining the semantic information of the current group and the semantic assignment weight of the current group to obtain the pairwise long connection feature of the current group;
[0015] Fusing the unary long connection feature of the current group and the pairwise long connection feature of the current group to obtain the complete feature of the current group, and the complete feature of the sample image being the sum of the complete features of G groups;
[0016] Using the complete feature of the sample image to predict the type of the sample image, calculating the cross-entropy loss between the predicted sample image type and the actual sample image type, and training the lightweight network model according to the cross-entropy loss to obtain a trained model;
[0017] Using the trained model to classify an image to obtain the semantic category to which the image belongs.
[0018] The present invention also provides an image classification system based on lightweight Non-Local, the system comprising: an abstract feature acquisition module for inputting a sample image into a lightweight network and obtaining abstract features after each bottleneck layer of the lightweight network, the lightweight network including ShufflenetV2;
[0019] A grouping module for dividing the abstract features into G groups along the channel dimension;
[0020] A spatial weight acquisition module for obtaining a spatial weight for each group of the abstract features by convolution, where the spatial weight is the global spatial attention information for the current group;
[0021] A unary long connection feature acquisition module for performing weighted amplitude conversion on the pixel points of the current group using the global spatial attention information of the current group to obtain the unary long connection feature of the current group;
[0022] A semantic information acquisition module for using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic aggregation weight, and performing global weighted pooling on the current group according to the semantic aggregation weight to obtain the semantic information of the current group;
[0023] A semantic assignment weight acquisition module for using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic assignment weight for the current group;
[0024] A point pair long connection feature acquisition module for combining the semantic information of the current group and the semantic assignment weight of the current group to obtain the point pair long connection feature of the current group;
[0025] A complete feature acquisition module for fusing the unary long connection feature of the current group and the point pair long connection feature of the current group to obtain the complete feature of the current group, where the complete feature of the sample image is the sum of the complete features of G groups;
[0026] A model training module for predicting the type of the sample image using the complete feature of the sample image, calculating the cross-entropy loss between the predicted sample image type and the actual sample image type, and training the lightweight network model according to the cross-entropy loss to obtain a trained model;
[0027] An image classification module for classifying an image using the trained model to obtain the semantic category to which the image belongs.
[0028] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0029] The image classification method and system based on lightweight Non-Local proposed by the present invention utilize the grouped implicit clustering characteristics of the lightweight network to group the abstract features after the bottleneck layer, simplify the method for obtaining the long connection relationship of the original Non-Local, thereby reducing the computational complexity. The features are extracted using the unary long connection relationship and the point pair long connection relationship, retaining the long connection performance originally brought by Non-Local while reducing the computational complexity, and achieving the compatibility between the computational complexity and the Non-Local long connection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 Flowchart of the image classification method based on lightweight Non-Local provided in Embodiment 1 of the present invention;
[0032] Figure 2 Schematic diagram of the traditional Non-Local response sharing phenomenon presented by each group in the lightweight network provided in Embodiment 1 of the present invention under the implicit clustering property;
[0033] Figure 3 Overall flowchart composed of three modules provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0035] The object of the present invention is to provide an image classification method and system based on lightweight Non-Local, which retains the long connection performance originally brought by Non-Local while reducing the computational complexity, and achieves the compatibility between the computational complexity and the Non-Local long connection.
[0036] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0037] Embodiment 1
[0038] This embodiment provides an image classification method based on lightweight Non-Local. Please refer to Figure 1 , the method includes:
[0039] S1. Input the sample image into a lightweight network, and obtain abstract features after each bottleneck layer of the lightweight network. The lightweight network includes ShufflenetV2.
[0040] Set the model processing task as an image classification task in computer vision. According to the general criteria in the image classification task, divide the images into a training set and a test set, such as CIFAR100 and ImageNet. Set the backbone network used as a lightweight network with limited number of channels, such as ShufflenetV2. After the image enters the network as input, abstract features are obtained after each bottleneck layer where C, H, and W are the number of channels, height, and width of the feature X respectively, and the present invention enhances this abstract feature.
[0041] S2. Divide the abstract features into G groups along the channel dimension.
[0042] After each bottleneck layer of the lightweight network, lightweight Non-Local processing is performed on the features. First, feature grouping is completed, and the features after passing through the bottleneck layer are divided into G (abbreviation of Group) groups along the channel dimension. Then the features of each group are where i = 1, 2,..., G. The neural network itself has the characteristic of different groups learning different semantics. Therefore, the grouped features directly have the property of implicit clustering. As Figure 2 shown, as Figure 2 is the traditional Non-Local response sharing phenomenon presented by each group in the lightweight network under the property of implicit clustering. For all pixel points belonging to the semantics of this group, their Non-Local responses are the same, so they can be simplified. While the pixel points (noise pixels) that do not belong to the semantics of this group do not share these responses, such as Figure 2 e in the apple picture and f in the desk lamp picture in
[0043] It should be noted that: The solution of the present invention includes three modules, namely a unary long connection relationship module, a point pair long connection relationship module, and a normalization module. Figure 3 is the overall flowchart composed of the three modules in the present invention, that is, the lightweight Non-Local module. This module can be added after each bottleneck layer of the lightweight network. In the figure, visual feature maps are used to display the pixel point response results of Non-Local long connections.
[0044] The following steps S3 - S4 are the unary long connection relationship module, and the essence of this module is a kind of spatial attention mechanism. Among them, step S3 is the process of obtaining the unary long connection relationship, and S4 is the process of applying the unary long connection relationship.
[0045] S3. For each group of the abstract features, obtain a spatial weight through convolution, and the spatial weight is the global spatial attention information for the current group.
[0046] Specifically: for each group of abstract features, use a lightweight 1x1 convolution to obtain a spatial weight Mask, denoted by the symbol where the superscript u (unary) represents the unary long connection relationship. The formula is as follows:
[0047]
[0048] where σ() represents the Sigmoid activation function, used to control the weight value within the range of 0 - 1. The normalize() function is the application of the normalization module, which will be specifically introduced in the subsequent process. It should be noted that the 1x1 convolution in formula (1) is applied to the entire feature map, that is, considering the semantic information of all groups. represents the similarity between each pixel point in group i and the semantic information of the current group, that is, the unary long connection relationship mentioned above. Although the current group semantic is not directly obtained, this content has been implicitly included in the used 1x1 convolution.
[0049] S4. Use the global spatial attention information of the current group to perform weighted amplitude conversion on the pixel points of this group to obtain the unary long connection feature of the current group.
[0050] According to perform weighted amplitude scaling on the pixel point features of the current group to obtain the unary long connection feature Thereby enhancing the spatial distribution of semantic features within the group while suppressing the interference of noise features. The formula is as follows:
[0051]
[0052] It should also be noted that the subsequent steps S5 - S7 are the point - to - point long connection relationship module. Among them, step S5 is the process of obtaining the point - to - point long connection relationship, and steps S6 - S7 are the process of applying / distributing the point - to - point long connection relationship.
[0053] S5. Use the unary long connection feature of the current group to obtain a semantic aggregation weight through a lightweight convolution, and perform global weighted pooling on the current group according to the semantic aggregation weight to obtain the semantic information of the current group.
[0054] In the traditional Non-Local relationship, the point-to-long connection relationship needs to calculate the similarity between the pixel point Query and all the pixel points Key on the entire feature map, and then use the features of the pixel point Key to enhance the feature representation of the pixel point Query according to the similarity weight. The Key responses obtained for each Query are different. However, in a single group of lightweight networks, we find that due to the implicit clustering characteristics of the group itself, different pixel points share a semantics, so the Key responses of each pixel point Query are the same in a special way. And this shared response in the point-to-long connection relationship actually represents the global semantic information of the current group. Therefore, a lightweight 1x1 convolution is used to obtain a semantic aggregation Mask, denoted by the symbol According to weight, global weighted pooling is performed to obtain the precise semantic information of the current group The formula is as follows:
[0055]
[0056] where σ s () is the softmax function along the entire spatial dimension (H and W).
[0057] S6. Use a lightweight convolution on the unary long connection features of the current group to obtain a semantic assignment weight for the current group.
[0058] S7. Combine the semantic information of the current group and the semantic assignment weight of the current group to obtain the point-to-long connection features of the current group.
[0059] The precise semantic information of the current group obtained in step S5 only needs to be assigned to all pixel points related to the semantics of the current group. If these semantics are deployed to noise pixel points, the noise pixel points will wrongly become part of the semantics of this group, thus interfering with the specific spatial distribution of the features of this group. Therefore, a semantic assignment Mask is also needed to determine the weight of each pixel point's demand for global semantics, denoted by the symbols
[0060] Combined with the semantic information Then the point-to-long connection features can be obtained The formula is as follows:
[0061]
[0062] S8. Fuse the unary long connection features of the current group and the point-to-long connection features of the current group to obtain the complete features of the current group. The complete features of the sample image are the sum of the complete features of G groups.
[0063] Combining the long connection features of both can be achieved by using simple element-wise addition. The calculation formula is as follows:
[0064]
[0065] Where is a learnable weight parameter used to determine the requirement degree of each group for the pair-wise long connection relationship, and is initialized to 0 to support early model learning. is the complete feature of the current group enhanced by lightweight Non-Local. The complete features of each group are combined according to the division method during the previous grouping to obtain the complete feature of the image.
[0066] Considering that the images in the dataset have feature contents with different amplitude deviations, which may have a negative impact on the learning of the attention Mask weights, a normalization module is added to stabilize the learning of the above three Mask weights.
[0067] Specifically, for the weight feature atti, first perform normalization in the spatial dimensions (H and W), and the formula is as follows:
[0068]
[0069] Where μ i and are the mean and variance of the feature atti in the spatial dimensions respectively. ε (default value is 1e-5) is used to ensure the mathematical stability of the division. Then, apply an affine transformation with learnable parameters γ and β to make the normalized value able to restore to the effect of the identity transformation, and the formula is as follows:
[0070]
[0071] S9. Use the complete feature of the sample image to predict the type of the sample image, calculate the cross-entropy loss between the predicted sample image type and the actual sample image type, and train the lightweight network model according to the cross-entropy loss to obtain a trained model.
[0072] Based on the given image classification task, the general training strategy of image classification can be used to normally train the network model with the added lightweight Non-Local module. Generally speaking, according to the true label given in the dataset and the predicted label output by the model, calculate the cross-entropy loss between the two, and then perform backpropagation to correct the network weights through Stochastic Gradient Descent (SGD). Specific training parameters such as the learning rate are specified by the specific classification task.
[0073] S10. Use the trained model to classify the image to obtain the semantic category to which the image belongs.
[0074] Given the image required for the test, it is input into the trained model as input, and finally the maximum score of each predicted class is taken. The class with the maximum value is the predicted classification result given by the model. At this point, the test ends.
[0075] The above scheme of the present invention mainly has several innovative points: 1) The present invention combines the semantic characteristics of implicit clustering of group features in lightweight networks, and finds that the pixel features belonging to this group have the same response on other global pixels. Based on this discovery, the original Non-Local long connection relationship acquisition mechanism is simplified, thereby reducing the computational complexity; 2) The present invention retains the long connection performance originally brought by Non-Local while reducing the computational complexity, and achieves the compatibility of computational complexity and Non-Local long connection; 3) The present invention constructs "lightweight Non-Local" on the basis of lightweight networks, retaining the ultimate realization purpose of lightweight networks, that is, ensuring lightweight so as to be applied to computing edge devices such as smartphones in actual situations.
[0076] Example 2
[0077] This embodiment provides a lightweight Non-Local-based image classification system, the system comprising:
[0078] An abstract feature acquisition module M1, used for inputting a sample image into a lightweight network, and obtaining an abstract feature after each bottleneck layer of the lightweight network, wherein the lightweight network includes ShufflenetV2;
[0079] A grouping module M2, used for dividing the abstract features into G groups along the channel dimension;
[0080] A spatial weight acquisition module M3 is used to obtain a spatial weight for each group of abstract features by convolution, where the spatial weight is the global spatial attention information for the current group;
[0081] A unary long connection feature acquisition module M4 is used to perform weighted amplitude conversion on the pixel points of the group using the global spatial attention information of the current group to obtain a unary long connection feature of the current group;
[0082] The semantic information acquisition module M5 is used to obtain a semantic aggregation weight by using a lightweight convolution for the unary long connection feature of the current group, and perform global weighted pooling on the current group according to the semantic aggregation weight to obtain the semantic information of the current group;
[0083] The semantic allocation weight acquisition module M6 is used to obtain a semantic allocation weight of the current group by using a lightweight convolution to obtain the unary long connection feature of the current group;
[0084] The point - pair long - connection feature acquisition module M7 is used to combine the semantic information of the current group and the semantic assignment weights of the current group to obtain the point - pair long - connection feature of the current group;
[0085] The complete feature acquisition module M8 is used to fuse the unary long - connection feature of the current group and the point - pair long - connection feature of the current group to obtain the complete feature of the current group. The complete feature of the sample image is the sum of the complete features of G groups;
[0086] The model training module M9 is used to predict the type of the sample image by using the complete feature of the sample image, calculate the cross - entropy loss between the predicted sample image type and the actual sample image type, and train the lightweight network model according to the cross - entropy loss to obtain a trained model;
[0087] The image classification module M10 is used to classify an image by using the trained model to obtain the semantic category to which the image belongs.
[0088] For the system disclosed in the embodiments of the present invention, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For related parts, refer to the description in the method section.
[0089] In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A lightweight Non-Local based image classification method, characterized in that, The method includes: Inputting a sample image into a lightweight network, and obtaining abstract features after each bottleneck layer of the lightweight network; the lightweight network includes ShufflenetV2; Dividing the abstract features along the channel dimension into G groups; For the abstract features of each group, obtaining a spatial weight through convolution, and the spatial weight is the global spatial attention information for the current group; Using the global spatial attention information of the current group to perform weighted amplitude conversion on the pixel points of this group, and obtaining the unary long connection feature of the current group; Using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic aggregation weight, and performing global weighted pooling on the current group according to the semantic aggregation weight to obtain the semantic information of the current group; Using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic assignment weight for the current group; Combining the semantic information of the current group and the semantic assignment weight of the current group to obtain the point pair long connection feature of the current group; Fusing the unary long connection feature of the current group and the point pair long connection feature of the current group to obtain the complete feature of the current group, and the complete feature of the sample image is the sum of the complete features of G groups; Using the complete feature of the sample image to predict the type of the sample image, calculating the cross-entropy loss between the predicted sample image type and the actual sample image type, and training the lightweight network model according to the cross-entropy loss to obtain a trained model; Using the trained model to classify an image to obtain the semantic category to which the image belongs.
2. The method according to claim 1, characterized in that, The calculation formula of the spatial weight includes: Among them, is the spatial weight, representing the similarity between each pixel point in the group i and the current group semantic information, The superscript u represents the unary long connection relationship, i = 1, 2,..., G, is the set of real numbers, H and W are the height and width of the feature X respectively; σ() represents the Sigmoid activation function, used to control the weight value within the range of 0-1; The normalize() function is the application of a normalization module, and 1x1 is a 1x1 convolution.
3. The method according to claim 2, characterized in that, The calculation formula of the unary long connection feature includes: In the formula, is a one-dimensional long connection feature, C is the number of channels of feature X, G is G groups, and i = 1, 2,..., G.
4. The method according to claim 3, characterized in that, The process of using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic aggregation weight, and performing global weighted pooling on the current group according to the semantic aggregation weight to obtain the semantic information of the current group specifically includes: The unary long connection features of the current group are used to obtain a semantic aggregation weight through a lightweight 1x1 convolution According to the semantic aggregation weight Global weighted pooling is performed on the current group to obtain the semantic information of the current group The formula is as follows: where σ s () is the softmax function along the entire spatial dimensions H and W.
5. The method according to claim 4, characterized in that, The calculation formula of the semantic assignment weight includes: In the formula, represents the semantic assignment weight, 6. The method according to claim 5, characterized in that, The calculation formula of the point pair long connection feature includes: In the formula, represents the point-to-point long connection feature, 7. The method according to claim 6, characterized in that, The process of fusing the unary long connection feature of the current group and the point pair long connection feature of the current group to obtain the complete feature of the current group specifically includes: Adding the unary long connection feature of the current group and the point pair long connection feature of the current group using element-wise addition, and the calculation formula is as follows: Among them, is a learnable weight parameter used to determine the requirement degree of each group for the point-to-point long connection relationship, and is initialized to a value of 0 to support early model learning. is the complete feature of the current group enhanced by lightweight Non-Local.
8. The method according to claim 1, wherein The method further includes: normalizing the spatial weight, the semantic aggregation weight, and the semantic assignment weight.
9. The method according to claim 8, wherein The normalization of the spatial weight, the semantic aggregation weight, and the semantic assignment weight specifically includes: For the weight features, performing normalization in the spatial dimension, and applying an affine transformation with learnable parameters γ and β to process the normalized values to obtain an identity transformation value, and the weight features include: the spatial weight, the semantic aggregation weight, and the semantic assignment weight; Among them, the formula for performing the normalization includes: where μ i and are the mean and variance of the feature atti in the spatial dimension respectively; ε is used to ensure the mathematical stability of the division; The formula for the affine transformation includes:
10. A lightweight Non-Local based image classification system, wherein The system includes: An abstract feature acquisition module for inputting a sample image into a lightweight network to obtain abstract features after each bottleneck layer of the lightweight network, where the lightweight network includes ShufflenetV2; A grouping module for dividing the abstract features into G groups along the channel dimension; A spatial weight acquisition module for obtaining a spatial weight for each group of the abstract features through convolution, where the spatial weight is the global spatial attention information for the current group; A unary long connection feature acquisition module for performing weighted amplitude conversion on the pixel points of the current group using the global spatial attention information of the current group to obtain the unary long connection feature of the current group; A semantic information acquisition module for using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic aggregation weight, and performing global weighted pooling on the current group according to the semantic aggregation weight to obtain the semantic information of the current group; A semantic assignment weight acquisition module for using a lightweight convolution on the unary long connection feature of the current group to obtain a semantic assignment weight for the current group; A point pair long connection feature acquisition module for combining the semantic information of the current group and the semantic assignment weight of the current group to obtain the point pair long connection feature of the current group; A complete feature acquisition module for fusing the unary long connection feature of the current group and the point pair long connection feature of the current group to obtain the complete feature of the current group, and the complete feature of the sample image is the sum of the complete features of G groups; A model training module for predicting the type of the sample image using the complete feature of the sample image, calculating the cross-entropy loss between the predicted sample image type and the actual sample image type, and training the lightweight network model according to the cross-entropy loss to obtain a trained model; An image classification module for classifying an image using the trained model to obtain the semantic category to which the image belongs.
Citation Information
Patent Citations
Zero-sample learning image classification method based on global and local context awareness
CN112418351A
Lightweight semantic segmentation method based on multi-scale visual feature extraction
CN112634276A