Image Classification Method, Apparatus, Electronic Device, and Storage Medium
By using the instance-aware feature map in the image classification method to acquire the area enlarged image and make predictions, the problem of high resource consumption cost in the prior art is solved, and the effect of reducing resource costs is achieved.
Patent Information
- Application Number
- CN202111227594.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-10-21
AI Technical Summary
In the prior art, the resource consumption cost of the image classification processing method is relatively high, and it is difficult to effectively reduce the overall cost of the image classification system.
By acquiring the image to be classified, a feature map with instance perception is obtained, thereby obtaining an area enlarged image, and using the area enlarged image to predict the attribute category of the image.
This reduces the resource consumption cost of image classification processing, further reduces the resource cost of real-life image classification work needs, and solves the problem of high resource consumption cost.
Smart Images

Figure CN114155395B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image classification method, apparatus, electronic device, and storage medium. Background Art
[0002] In the requirements of large-scale online image recognition work, the same image may contain multiple instance objects from different categories, scales, and positions. Therefore, simply applying a single-label dataset for image recognition, for example, a convolutional neural network model CNN pre-trained based on the visual object recognition research visualization database ImageNet, is still not the ultimate image classification solution.
[0003] In addition, in the image classification scheme based on a strong supervision classification model, due to the serious resource consumption of manually annotating the target position boxes required by the strong supervision classification model, therefore, how to improve the performance of a weakly supervised multi-label image classification model that only depends on image labels and reduce the overall cost of the image classification system has important research value for the actual work requirements based on image recognition.
[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide an image classification method, apparatus, electronic device, and storage medium to at least solve the technical problem of high resource consumption cost in the image classification processing method in the related art.
[0006] According to one aspect of the embodiments of the present invention, an image classification method is provided, including: obtaining an image to be classified; analyzing the image to be classified to obtain a first feature map, where the first feature map is a feature map with instance perception; based on the first feature map, obtaining a region magnified image from the image to be classified, where the region magnified image is a local instance region in the image to be classified; and predicting the belonging category of the image to be classified by using the region magnified image.
[0007] According to another aspect of the embodiments of the present invention, an image classification method is further provided, including: receiving an image to be classified from a client; analyzing the image to be classified to obtain a first feature map, obtaining a region magnified image from the image to be classified based on the first feature map, and predicting the belonging category of the image to be classified by using the region magnified image, where the first feature map is a feature map with instance perception, and the region magnified image is a local instance region in the image to be classified; and feeding back the belonging category of the image to be classified to the client.
[0008] According to another aspect of the embodiments of the present invention, an image classification device is further provided, including: a first acquisition module, configured to acquire an image to be classified; an analysis module, configured to analyze the image to be classified to obtain a first feature map, where the first feature map is a feature map with instance awareness; a second acquisition module, configured to acquire a region magnification image from the image to be classified based on the first feature map, where the region magnification image is a local instance region in the image to be classified; and a classification module, configured to predict the attribution category of the image to be classified by using the region magnification image.
[0009] According to another aspect of the embodiments of the present invention, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored program, where, when the program runs, it controls the device where the non-volatile storage medium is located to execute any one of the above image classification methods.
[0010] According to another aspect of the embodiments of the present invention, a processor is further provided. The processor is configured to run a program, where, when the program runs, it executes any one of the above image classification methods.
[0011] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including: a processor; and a memory, connected to the processor, configured to provide instructions for the processor to perform the following processing steps: Step 1, acquire an image to be classified; Step 2, analyze the image to be classified to obtain a first feature map, where the first feature map is a feature map with instance awareness; Step 3, acquire a region magnification image from the image to be classified based on the first feature map, where the region magnification image is a local instance region in the image to be classified; Step 4, predict the attribution category of the image to be classified by using the region magnification image.
[0012] In the embodiments of the present invention, a method of performing local region recognition based on an instance-aware feature map and predicting the attribution classification of an image to be classified is adopted. By analyzing the acquired image to be classified, a first feature map with instance awareness is obtained; based on the first feature map, a region magnification image is acquired from the image to be classified. Since the region magnification image is a local instance region in the image to be classified; further, the attribution category of the image to be classified can be predicted by using the region magnification image, achieving the purpose of reducing the resource consumption cost of image classification processing, thereby realizing the technical effect of further reducing the resource cost for meeting the actual image classification work requirements, and further solving the technical problem of the relatively high resource consumption cost in the related art of image classification processing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0014] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing an image classification method is shown;
[0015] Figure 2 A flowchart of an image classification method according to an embodiment of the present invention;
[0016] Figure 3 A schematic structural diagram of an image classification system according to an embodiment of the present invention;
[0017] Figure 4 A flowchart of another image classification method according to an embodiment of the present invention;
[0018] Figure 5 A schematic structural diagram of an image classification device according to an embodiment of the present invention;
[0019] Figure 6 A structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed implementation manners
[0020] In order to enable those skilled in the art of the present technology to better understand the present invention solution, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0022] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:
[0023] Transformer: A transformer network structure based on the encoder-decoder, which uses the attention mechanism to capture sequence dependencies in parallel. When used in vision tasks, usually only the encoder is retained.
[0024] Vision Transformer Network (ViT): Used to serialize pictures or images.
[0025] In the related art, the research directions for multi-label image classification can be roughly divided into four categories. First, the first direction is the work based on region proposals. Second, the second direction is the work based on the visual attention mechanism. Third, the third direction is the work of knowledge injection. Finally, other research directions are combinations of some of the above directions or research in other directions.
[0026] For the first direction, usually the method of object detection is used to extract object prediction boxes from the image, and then the multi-label image classification problem is transformed into a multi-class multi-instance learning problem. The disadvantage of this method is that it requires the use of supervision information of the bounding box and the computational cost is too high.
[0027] For the second direction, the attention mechanism needs to be used to capture the potential relationship between labels through convolution. The models in this direction usually have a low accuracy.
[0028] For the third direction, currently it is popular to use GCN to model the potential dependencies between labels, based on the statistics of the co-occurrence relationship of labels in the dataset. However, this type of method may face the problem of unreliable statistical relationships. For other research directions, some solutions need to introduce a large pre-trained object detection model to extract the object categories in the image and the spatial relationships between objects, and combine with the GCN network at the same time. However, the disadvantages of this type of method are also obvious. The effect on the target dataset depends on the performance of the detection model. When the difference between the pre-trained dataset and the target dataset is too large, the model performance may decline.
[0029] Embodiment 1
[0030] According to an embodiment of the present invention, an embodiment of an image processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0031] The method embodiment provided by Embodiment 1 of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 1A hardware structure block diagram of a computer terminal (or mobile device) for implementing an image classification method is shown. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ……, 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0032] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0033] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image classification method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned image classification method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0034] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0035] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0036] Under the above operating environment, the present invention provides an image classification method as Figure 2 shown, Figure 2 which is a flowchart of an image classification method according to an embodiment of the present invention. As Figure 2 shown, the method includes:
[0037] Step S202, obtaining an image to be classified;
[0038] Step S204, analyzing the image to be classified to obtain a first feature map, where the first feature map is a feature map with instance awareness;
[0039] Step S206, based on the first feature map, obtaining a region magnification image from the image to be classified, where the region magnification image is a local instance region in the image to be classified;
[0040] Step S208, using the region magnification image to predict the attribution category of the image to be classified.
[0041] In the embodiment of the present invention, a method of performing local region recognition based on an instance-aware feature map and predicting the attribution classification of an image to be classified is adopted. By analyzing the obtained image to be classified, a first feature map with instance awareness is obtained; based on the first feature map, a region magnification image is obtained from the image to be classified. Since the region magnification image is a local instance region in the image to be classified; furthermore, the region magnification image can be used to predict the attribution category of the image to be classified, achieving the purpose of reducing the resource consumption cost of image classification processing, thereby realizing the technical effect of further reducing the resource cost for meeting the actual image classification work requirements, and further solving the technical problem of the relatively high resource consumption cost of the image classification processing method in the related art.
[0042] It should be noted that the execution subject of the above steps S202 to S206 is the SaaS client. By analyzing and processing the acquired image to be classified, the attribution category of the image to be classified is predicted.
[0043] Optionally, the above image to be classified may be a multi-label image to be classified. For example, the multi-label image may be an image containing multiple objects from different categories, scales, and positions. Optionally, the above instance perception may include, but is not limited to, category semantic perception and spatial relationship perception.
[0044] In the embodiment of the present invention, by analyzing the image to be classified, a first feature map with instance perception is obtained; since the first feature map has instance perception, therefore, based on the first feature map, a region magnified image of the local instance region can be obtained from the above image to be classified. After positioning the region magnified image, the attribution category of the above image to be classified can be predicted using the region magnified image.
[0045] In the embodiment of the present invention, by modeling the operation of image convolution processing into a two-dimensional sequence processing problem, the overall structure of the image classification model can be made more complete and high-quality; the embodiment of the present invention effectively discovers different instance perceptions through the attention map of instance perception, providing an effective perspective to utilize the category semantic information and spatial relationship information extracted by the vision transformer network.
[0046] The above image classification method in the embodiment of the present invention can be understood as a multi-label image classification method based on instance perception of a vision transformer network. The transformer network pre-trained on a large-scale dataset can effectively capture the global dependency relationship of the image to be classified; by obtaining a region magnified image with instance perception ability, that is, an attention image, and then obtaining the prediction box of the local instance region through a positioning method, the local instance region can be determined on the basis of the image to be classified, and then the attribution category of the above image to be classified can be predicted using the region magnified image, achieving the purpose of reducing the resource consumption cost of image classification processing, thereby realizing the technical effect of further reducing the resource cost for meeting the actual image classification work requirements.
[0047] In an optional embodiment, analyzing the above image to be classified to obtain the above first feature map includes:
[0048] Step S302, performing serialization processing on the above image to be classified to obtain a first processing result;
[0049] Step S304, performing transformation processing on the above first processing result to obtain a second processing result and a third processing result;
[0050] Step S306: Obtain the above-mentioned first feature map by using the above-mentioned second processing result and the above-mentioned third processing result.
[0051] Optionally, the image to be classified can be directly serialized into the first processing result. For example, the first processing result can be image patch tokenization (Patch Token). The above transformation processing is performed using the Transformer network structure. The second processing result is the result of taking out the Patch Tokens after the second dimension. The third processing result is the self-attention weight matrix.
[0052] As an optional embodiment, if the dimension of the preprocessed image X to be classified is (B, C, H, W), where B is the batch size, C is the number of channels, H is the height of the feature map, and W is the width of the feature map; the serialization operation level refers to the convolutional dimensionality reduction operation Conv; the feature map after the serialization operation is transformed into a two-dimensional sequence X', with a dimension of (B, N, D), where N is the number of image patches (patch), and D is the number of output channels of Conv.
[0053] Moreover, in the embodiments of the present invention, the position encoding Patch Embedding and the classification token CLS Token can also be initialized simultaneously. Among them, the position encoding Patch Embedding and the Patch Token are added, and the CLS Token and the Patch Token are merged.
[0054] In an optional embodiment, obtaining the above-mentioned first feature map by using the above-mentioned second processing result and the above-mentioned third processing result includes:
[0055] Step S402: Perform convolutional dimensionality reduction processing on the above-mentioned second processing result to obtain a plurality of second feature maps, where the plurality of second feature maps are feature maps with semantic class awareness;
[0056] Step S404: Perform global average pooling processing on the plurality of second feature maps to obtain a first category prediction result, where the first category prediction result is a category prediction result at the global scale;
[0057] Step S406: Select a plurality of third feature maps from the plurality of second feature maps based on the above-mentioned first category prediction result;
[0058] Step S408: Perform repositioning processing on the above-mentioned third processing result and the plurality of third feature maps to obtain the above-mentioned first feature map.
[0059] Optionally, such as Figure 3Schematic structural diagram of the image classification system shown. In an embodiment of the present invention, the image patch tokens obtained by performing serialization processing operations are sent to a transformer layer (Transformer layer) loaded with pre-trained weights for calculation; wherein, the network structure of the Transformer layer may include N stacked encoder layers with the same structure; the above two-dimensional sequence X' can be first processed by a self-attention module and then sent to a feed-forward neural network (Feed Forward Neural Network) module. The core of the self-attention module is the multi-head attention mechanism; the feed-forward neural network module may be composed of two fully connected layers, and the fully connected layers are connected by activation functions.
[0060] In an embodiment of the present invention, for the last hidden layer state value calculated and output by the Transformer layer, the patch tokens after the second dimension are taken out and sent to a 1*1 convolution for dimensionality reduction processing to obtain multiple second feature maps A with class perception, and then a global average pooling operation is performed to obtain a first class prediction result y at the global scale; then according to the value of y, the top N third feature maps A' are taken out from the multiple second feature maps A for subsequent repositioning processing to obtain the above first feature map.
[0061] Among them, the above y is the class confidence, which is equivalent to the probability of each class appearing. Therefore, in an embodiment of the present invention, y can be sorted in reverse order according to size, and the top N third feature maps A' in the second feature map A are selected according to the sorted sequence for the next step of strengthening the spatial relationship perception and dynamic instance positioning.
[0062] In an alternative embodiment, based on the above first feature map, obtaining the above regional enlarged image from the above image to be classified includes:
[0063] Step S502, using the maximum connected region selection method to obtain the regional proposal coordinates of the above first feature map;
[0064] Step S504, performing cropping and enlargement processing on the above image to be classified through the above regional proposal coordinates to obtain the above regional enlarged image.
[0065] As an alternative embodiment, as Figure 3 shown, performing a repositioning operation on the above top N semantic class perception feature maps and self-attention weight matrices to obtain an instance perception feature map; and obtaining the regional proposal coordinates of the first feature map through the maximum connected region selection algorithm, as Figure 3The coordinates of the regions shown, such as Region Proposal 1 and Region Proposal 2, are provided. Meanwhile, the region cropping and magnification of the image to be classified are performed based on these region proposal coordinates at the same resolution as the image to be classified.
[0066] It should be noted that the above-mentioned maximum connected region selection algorithm is integrated in the calculation module of the Sklearn library; the region proposal coordinates are obtained through the maximum connected region selection algorithm, which are the coordinate values of the upper left vertex and the lower right vertex in the maximum connected region; the bilinear interpolation algorithm is used for region cropping and magnification to obtain the above-mentioned magnified region image.
[0067] In an alternative embodiment, predicting the belonging category of the image to be classified using the above-mentioned magnified region image includes:
[0068] Step S602: Serialize the above-mentioned magnified region image to obtain a fourth processing result;
[0069] Step S604: Transform the above-mentioned fourth processing result to obtain a fifth processing result;
[0070] Step S606: Predict the belonging category of the image to be classified using the above-mentioned fifth processing result.
[0071] Due to the problems of high cost of labeled data and insufficient ability to extract regions of interest in the related art, the multi-head self-attention weights output by the calculation of the Transformer itself can be used for region localization in the image feature map; in addition, due to the problems of unreliable statistical relationships and sparse label co-occurrence matrices in the related art, it is difficult to transfer to other datasets.
[0072] In the above alternative embodiment, the fourth processing result is obtained by serializing the magnified region image, and the fifth processing result is obtained by transforming the above-mentioned fourth processing result using the Transformer; then the belonging category of the image to be classified is predicted using the above-mentioned fifth processing result.
[0073] In an alternative embodiment, predicting the belonging category of the image to be classified using the above-mentioned fifth processing result includes:
[0074] Step S702: Perform convolutional dimensionality reduction processing on the above-mentioned fifth processing result to obtain multiple fourth feature maps, where the above-mentioned multiple fourth feature maps are feature maps with semantic category perception;
[0075] Step S704: Perform global average pooling processing on the above-mentioned fourth feature maps to obtain a second category prediction result, where the above-mentioned second category prediction result is a category prediction result at the local scale;
[0076] Step S706: Predict the belonging category of the image to be classified based on the above second-category prediction result.
[0077] Optionally, after obtaining the above fifth processing result, perform convolution dimensionality reduction processing to obtain a fourth feature map with semantic category perception; through global average pooling processing on the above fourth feature map, obtain a second-category prediction result at the local scale; and then, based on the above second-category prediction result, predict the belonging category of the image to be classified.
[0078] In the above optional embodiment, since the fourth feature map of the same size needs to be sent into the Transformer layer for calculation again, combined with the dynamic transformer network structure, it is possible to keep the model performance from degrading during the model training and inference stages while reducing the computational amount, and it is possible to eliminate the disadvantages of using the label co-occurrence matrix in the related art.
[0079] As an optional embodiment, the self-attention weights obtained by each layer of the encoder Encoder can be output and saved simultaneously, and all the saved self-attention weights are calculated to obtain a self-attention weight matrix, that is, the spatial relationship perception matrix required in the embodiments of the present invention. The multi-head self-attention weights obtained through Transformer calculation are first normalized, then multiplied cumulatively starting from the first layer, and finally the weight value of the penultimate layer is taken to obtain the self-attention weight matrix.
[0080] In an optional embodiment, the above image classification method further includes:
[0081] Step S802: Obtain the first loss corresponding to the above first-category prediction result and the second loss corresponding to the above second-category prediction result;
[0082] Step S804: Optimize the belonging category of the image to be classified by backpropagation based on the above first loss and the above second loss.
[0083] In the embodiments of the present invention, by obtaining the first loss corresponding to the first-category prediction result and the second loss corresponding to the second-category prediction result, that is, calculating the losses of the classification prediction outputs at two different scales separately and adding them up, backpropagation optimization is performed.
[0084] In an embodiment of the present invention, an instance-aware multi-label image classification method based on the Vision Transformer network (ViT) is provided. The Vision Transformer network pre-trained on a large-scale dataset can effectively capture the global dependencies of an image. For multi-label images that contain multiple objects from different categories, scales, and positions, not only global information is utilized, but the image patch tokenization and self-attention mechanism of the Vision Transformer network (ViT) are fully utilized to mine rich instances in multi-label images. To this end, the embodiment of the present invention separately proposes a category semantic awareness module and a spatial relationship awareness module, and then combines the two through a repositioning strategy to obtain an attention image with instance awareness ability.
[0085] The present invention also provides an image classification method as Figure 4 shown, Figure 4 which is a flowchart of an image classification method according to an embodiment of the present invention. As Figure 4 shown, the method includes:
[0086] Step S902, receiving an image to be classified from a client;
[0087] Step S904, analyzing the image to be classified to obtain a first feature map, obtaining a region magnified image from the image to be classified based on the first feature map, and predicting the attribution category of the image to be classified by using the region magnified image, where the first feature map is a feature map with instance awareness, and the region magnified image is a local instance region in the image to be classified;
[0088] Step S906, feeding back the attribution category of the image to be classified to the client.
[0089] By adopting the method of performing local region recognition based on an instance-aware feature map and predicting the attribution classification of an image to be classified, through analyzing the obtained image to be classified, a first feature map with instance awareness is obtained; based on the first feature map, a region magnified image is obtained from the image to be classified. Since the region magnified image is a local instance region in the image to be classified, the attribution category of the image to be classified can be predicted by using the region magnified image, achieving the purpose of reducing the resource consumption cost of image classification processing, thereby realizing the technical effect of further reducing the resource cost for meeting the requirements of actual image classification work, and further solving the technical problem of the relatively high resource consumption cost in the image classification processing method in the related art.
[0090] It should be noted that the execution subject of the above steps S902 to S906 is the SaaS server, which is communicatively connected to the SaaS client. The SaaS server analyzes and processes the to-be-classified image obtained, predicts the attribution category of the to-be-classified image, and returns the predicted attribution category to the SaaS client.
[0091] Optionally, the to-be-classified image may be a multi-label image to be classified. For example, the multi-label image may be an image containing multiple objects from different categories, scales, and positions. Optionally, the above instance perception may include, but is not limited to, category semantic perception and spatial relationship perception.
[0092] In the embodiment of the present invention, by analyzing the to-be-classified image, a first feature map with instance perception is obtained; since the first feature map has instance perception, an enlarged image of the local instance region can be obtained from the to-be-classified image based on the first feature map. After locating the enlarged image of the region, the attribution category of the to-be-classified image can be predicted using the enlarged image of the region.
[0093] In the embodiment of the present invention, by modeling the operation of image convolution processing into a two-dimensional sequence processing problem, the overall structure of the image classification model can be made more complete and high-quality; the embodiment of the present invention effectively discovers different instance perceptions through the attention map of instance perception, providing an effective perspective to utilize the category semantic information and spatial relationship information extracted by the vision transformer network.
[0094] The above image classification method in the embodiment of the present invention can be understood as a multi-label image classification method based on instance perception of a vision transformer network. The transformer network pre-trained on a large-scale dataset can effectively capture the global dependency relationship of the to-be-classified image; by obtaining an enlarged image of the local instance region with instance perception ability, that is, the attention image, and then obtaining the prediction box of the local instance region through a localization method, the local instance region can be determined based on the to-be-classified image, and then the attribution category of the to-be-classified image can be predicted using the enlarged image of the region, achieving the purpose of reducing the resource consumption cost of image classification processing, and thus realizing the technical effect of further reducing the resource cost for meeting the requirements of actual image classification work.
[0095] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0097] Embodiment 2
[0098] According to an embodiment of the present invention, there is also provided an apparatus embodiment for implementing the above image classification method. Figure 5 As shown in FIG. 5, a schematic structural diagram of an image classification apparatus according to an embodiment of the present invention, the above apparatus includes: a first acquisition module 500, an analysis module 502, a second acquisition module 504, and a classification module 506, wherein,
[0099] The first acquisition module 500 is configured to acquire an image to be classified; the analysis module 502 is configured to analyze the image to be classified to obtain a first feature map, wherein the first feature map is a feature map with instance awareness; the second acquisition module 504 is configured to, based on the first feature map, acquire a region magnified image from the image to be classified, wherein the region magnified image is a local instance region in the image to be classified; the classification module 506 is configured to use the region magnified image to predict the belonging category of the image to be classified.
[0100] It should be noted here that the above first acquisition module 500, analysis module 502, second acquisition module 504, and classification module 506 correspond to steps S202 to S208 in Embodiment 1. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the apparatus, can run in the computer terminal 10 provided in Embodiment 1.
[0101] In an embodiment of the present invention, a method is adopted to perform local region recognition based on an instance-aware feature map and predict the attribution classification of an image to be classified. By analyzing the obtained image to be classified, a first feature map with instance awareness is obtained; based on the first feature map, a region magnified image is obtained from the image to be classified. Since the region magnified image is a local instance region in the image to be classified, the attribution category of the image to be classified can be predicted by using the region magnified image, achieving the purpose of reducing the resource consumption cost of image classification processing, thereby realizing the technical effect of further reducing the resource cost for meeting the actual image classification work requirements, and further solving the technical problem of high resource consumption cost in the image classification processing method in the related art.
[0102] It should be noted that the preferred implementation manner of this embodiment can refer to the relevant description in Embodiment 1 and will not be elaborated here.
[0103] Embodiment 3
[0104] According to an embodiment of the present invention, an embodiment of an electronic device is further provided. The electronic device can be any computing device in a computing device cluster. The electronic device includes: a processor and a memory, where:
[0105] a processor; and a memory connected to the processor for providing instructions for the processor to perform the following processing steps: Step 1, obtain an image to be classified; Step 2, analyze the image to be classified to obtain a first feature map, where the first feature map is a feature map with instance awareness; Step 3, based on the first feature map, obtain a region magnified image from the image to be classified, where the region magnified image is a local instance region in the image to be classified; Step 4, predict the attribution category of the image to be classified by using the region magnified image.
[0106] In an embodiment of the present invention, a method is adopted to perform local region recognition based on an instance-aware feature map and predict the attribution classification of an image to be classified. By analyzing the obtained image to be classified, a first feature map with instance awareness is obtained; based on the first feature map, a region magnified image is obtained from the image to be classified. Since the region magnified image is a local instance region in the image to be classified, the attribution category of the image to be classified can be predicted by using the region magnified image, achieving the purpose of reducing the resource consumption cost of image classification processing, thereby realizing the technical effect of further reducing the resource cost for meeting the actual image classification work requirements, and further solving the technical problem of high resource consumption cost in the image classification processing method in the related art.
[0107] It should be noted that the preferred implementation manner of this embodiment can refer to the relevant description in Embodiment 1 and will not be elaborated here.
[0108] Example 4
[0109] According to an embodiment of the present invention, an embodiment of a computer terminal can also be provided. The computer terminal can be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the above computer terminal can also be replaced with a terminal device such as a mobile terminal.
[0110] Optionally, in this embodiment, the above computer terminal can be located in at least one of multiple network devices in a computer network.
[0111] In this embodiment, the above computer terminal can execute program code for the following steps in an image classification method: obtaining an image to be classified; analyzing the above image to be classified to obtain a first feature map, where the above first feature map is a feature map with instance awareness; based on the above first feature map, obtaining a region magnified image from the above image to be classified, where the above region magnified image is a local instance region in the above image to be classified; and predicting the belonging category of the above image to be classified by using the above region magnified image.
[0112] Optionally, Figure 6 is a structural block diagram of a computer terminal according to an embodiment of the present invention. As Figure 6 shown, the computer terminal can include: one or more (only one is shown in the figure) processors 602, a memory 604, and a peripheral interface 606.
[0113] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the image classification method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above image classification method. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely disposed relative to the processor, and these remote memories can be connected to the computer terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.
[0114] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtain the image to be classified; analyze the image to be classified to obtain a first feature map, where the first feature map is a feature map with instance awareness; based on the first feature map, obtain a region magnification image from the image to be classified, where the region magnification image is a local instance region in the image to be classified; use the region magnification image to predict the belonging category of the image to be classified.
[0115] Optionally, the processor can also execute the program code of the following steps: perform serialization processing on the image to be classified to obtain a first processing result; perform transformation processing on the first processing result to obtain a second processing result and a third processing result; use the second processing result and the third processing result to obtain the first feature map.
[0116] Optionally, the processor can also execute the program code of the following steps: perform convolutional dimensionality reduction processing on the second processing result to obtain a plurality of second feature maps, where the plurality of second feature maps are feature maps with semantic category awareness; perform global average pooling processing on the plurality of second feature maps to obtain a first category prediction result, where the first category prediction result is a category prediction result at the global scale; select a plurality of third feature maps from the plurality of second feature maps based on the first category prediction result; perform repositioning processing on the third processing result and the plurality of third feature maps to obtain the first feature map.
[0117] Optionally, the processor can also execute the program code of the following steps: use the maximum connected region selection method to obtain the region proposal coordinates of the first feature map; perform cropping and magnification processing on the image to be classified through the region proposal coordinates to obtain the region magnification image.
[0118] Optionally, the processor can also execute the program code of the following steps: perform serialization processing on the region magnification image to obtain a fourth processing result; perform transformation processing on the fourth processing result to obtain a fifth processing result; use the fifth processing result to predict the belonging category of the image to be classified.
[0119] Optionally, the processor can also execute the program code of the following steps: perform convolutional dimensionality reduction processing on the fifth processing result to obtain a plurality of fourth feature maps, where the plurality of fourth feature maps are feature maps with semantic category awareness; perform global average pooling processing on the fourth feature maps to obtain a second category prediction result, where the second category prediction result is a category prediction result at the local scale; based on the second category prediction result, predict the belonging category of the image to be classified.
[0120] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining the first loss corresponding to the above-mentioned first category prediction result, and obtaining the second loss corresponding to the above-mentioned second category prediction result; performing backpropagation optimization on the attribution category of the above-mentioned image to be classified based on the above-mentioned first loss and the above-mentioned second loss.
[0121] Optionally, the above-mentioned processor may also execute the program code of the following steps: receiving the image to be classified from the client; analyzing the above-mentioned image to be classified to obtain a first feature map, obtaining a region magnified image from the above-mentioned image to be classified based on the above-mentioned first feature map, and predicting the attribution category of the above-mentioned image to be classified by using the above-mentioned region magnified image, wherein the above-mentioned first feature map is a feature map with instance awareness, and the above-mentioned region magnified image is a local instance region in the above-mentioned image to be classified; feeding back the attribution category of the above-mentioned image to be classified to the above-mentioned client.
[0122] By adopting the embodiment of the present invention, a solution for an image classification method is provided. By using the method of performing local region recognition based on an instance-aware feature map and predicting the attribution classification of the image to be classified, through analyzing the obtained image to be classified, a first feature map with instance awareness is obtained; based on the above-mentioned first feature map, a region magnified image is obtained from the above-mentioned image to be classified. Since the above-mentioned region magnified image is a local instance region in the above-mentioned image to be classified; thus, the attribution category of the above-mentioned image to be classified can be predicted by using the above-mentioned region magnified image, achieving the purpose of reducing the resource consumption cost of image classification processing, thereby realizing the technical effect of further reducing the resource cost for meeting the requirements of actual image classification work, and further solving the technical problem of high resource consumption cost in the related art of image classification processing methods.
[0123] Those of ordinary skill in the art can understand that Figure 6 the structure shown is only schematic, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 6 It does not limit the structure of the above-mentioned electronic device. For example, the computer terminal may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 6 or have a different configuration from that shown in Figure 6 .
[0124] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the hardware related to the terminal device. This program can be stored in a computer-readable non-volatile storage medium, and the non-volatile storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0125] Embodiment 5
[0126] According to an embodiment of the present invention, an embodiment of a non-volatile storage medium is further provided. Optionally, in this embodiment, the above non-volatile storage medium can be used to store the program code executed by the image classification method provided in the above Embodiment 1.
[0127] Optionally, in this embodiment, the above non-volatile storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0128] Optionally, in this embodiment, the non-volatile storage medium is set to store program code for performing the following steps: obtaining an image to be classified; analyzing the image to be classified to obtain a first feature map, where the first feature map is a feature map with instance awareness; based on the first feature map, obtaining a region magnified image from the image to be classified, where the region magnified image is a local instance region in the image to be classified; using the region magnified image to predict the attribution category of the image to be classified.
[0129] Optionally, in this embodiment, the non-volatile storage medium is set to store program code for performing the following steps: serializing the image to be classified to obtain a first processing result; performing a transformation process on the first processing result to obtain a second processing result and a third processing result; using the second processing result and the third processing result to obtain the first feature map.
[0130] Optionally, in this embodiment, the non-volatile storage medium is set to store program code for performing the following steps: performing a convolutional dimensionality reduction process on the second processing result to obtain a plurality of second feature maps, where the plurality of second feature maps are feature maps with semantic category awareness; performing a global average pooling process on the plurality of second feature maps to obtain a first category prediction result, where the first category prediction result is a category prediction result at the global scale; selecting a plurality of third feature maps from the plurality of second feature maps based on the first category prediction result; performing a repositioning process on the third processing result and the plurality of third feature maps to obtain the first feature map.
[0131] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining the region proposal coordinates of the first feature map by using the maximum connected region selection method; cropping and magnifying the to-be-classified image according to the region proposal coordinates to obtain the region magnified image.
[0132] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: serializing the region magnified image to obtain a fourth processing result; performing a transformation process on the fourth processing result to obtain a fifth processing result; predicting the belonging category of the to-be-classified image by using the fifth processing result.
[0133] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: performing a convolutional dimensionality reduction process on the fifth processing result to obtain a plurality of fourth feature maps, where the plurality of fourth feature maps are feature maps with semantic category awareness; performing a global average pooling process on the fourth feature maps to obtain a second category prediction result, where the second category prediction result is a category prediction result at a local scale; predicting the belonging category of the to-be-classified image based on the second category prediction result.
[0134] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a first loss corresponding to the first category prediction result and obtaining a second loss corresponding to the second category prediction result; performing backpropagation optimization on the belonging category of the to-be-classified image based on the first loss and the second loss.
[0135] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: receiving a to-be-classified image from a client; analyzing the to-be-classified image to obtain a first feature map, obtaining a region magnified image from the to-be-classified image based on the first feature map, and predicting the belonging category of the to-be-classified image by using the region magnified image, where the first feature map is a feature map with instance awareness, and the region magnified image is a local instance region in the to-be-classified image; feeding back the belonging category of the to-be-classified image to the client.
[0136] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0137] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0138] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0139] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0141] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned non-volatile storage media include: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.
[0142] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An image classification method, characterized in that, Including: Obtain the image to be classified; Use a vision transformer network to perform serialization processing on the image to be classified to obtain a first processing result; Use a transformer network structure to perform transformation processing on the first processing result to obtain a second processing result and a third processing result, where the third processing result is a self-attention weight matrix, and the self-attention weight matrix includes: weight values obtained by performing a layer-by-layer cumulative multiplication operation on the normalized self-attention weights; Perform convolutional dimensionality reduction processing on the second processing result to obtain a plurality of second feature maps, where the plurality of second feature maps are feature maps with semantic category perception; perform global average pooling processing on the plurality of second feature maps to obtain a first category prediction result, where the first category prediction result is a category prediction result at the global scale; select a plurality of third feature maps from the plurality of second feature maps based on the first category prediction result; perform repositioning processing on the third processing result and the plurality of third feature maps to obtain a first feature map, where the first feature map is a feature map with category semantic perception and spatial relationship perception; Based on the first feature map, obtain a region magnified image from the image to be classified, where the region magnified image is a local instance region in the image to be classified; Use the region magnified image to predict the belonging category of the image to be classified.
2. The image classification method according to claim 1, characterized in that, Based on the first feature map, obtaining the region magnified image from the image to be classified includes: Use the maximum connected region selection method to obtain the region proposal coordinates of the first feature map; Perform cropping and magnification processing on the image to be classified through the region proposal coordinates to obtain the region magnified image.
3. The image classification method according to claim 1, characterized in that, Using the region magnified image to predict the belonging category of the image to be classified includes: Use the vision transformer network to perform serialization processing on the region magnified image to obtain a fourth processing result; Use the transformer network structure to perform transformation processing on the fourth processing result to obtain a fifth processing result; Use the fifth processing result to predict the belonging category of the image to be classified.
4. The image classification method according to claim 3, characterized in that, Using the fifth processing result to predict the belonging category of the image to be classified includes: Perform convolutional dimensionality reduction processing on the fifth processing result to obtain a plurality of fourth feature maps, where the plurality of fourth feature maps are feature maps with semantic category perception; Perform global average pooling processing on the fourth feature maps to obtain a second category prediction result, where the second category prediction result is a category prediction result at the local scale; Based on the second category prediction result, predict the belonging category of the image to be classified.
5. The image classification method according to claim 4, characterized in that, The image classification method further includes: Obtain a first loss corresponding to the first category prediction result and obtain a second loss corresponding to the second category prediction result; Perform backpropagation optimization on the belonging category of the image to be classified based on the first loss and the second loss.
6. An image classification method, characterized in that, Including: Receive the image to be classified from the client; Use a vision transformer network to perform serialization processing on the image to be classified to obtain a first processing result; The first processing result is processed by a transformer network structure to obtain a second processing result and a third processing result, where the third processing result is a self-attention weight matrix, and the self-attention weight matrix includes: weight values obtained by performing a layer-by-layer multiplication operation on the normalized self-attention weights; the second processing result is subjected to a convolutional dimensionality reduction process to obtain a plurality of second feature maps, where the plurality of second feature maps are feature maps with semantic class awareness; the plurality of second feature maps are subjected to global average pooling to obtain a first class prediction result, where the first class prediction result is a class prediction result at the global scale; a plurality of third feature maps are selected from the plurality of second feature maps based on the first class prediction result; the third processing result and the plurality of third feature maps are subjected to a repositioning process to obtain a first feature map, where the first feature map is a feature map with class semantic awareness and spatial relationship awareness; based on the first feature map, a region magnification image is obtained from the image to be classified, and the attribution class of the image to be classified is predicted using the region magnification image, where the region magnification image is a local instance region in the image to be classified. The attribution class of the image to be classified is fed back to the client.
7. An image classification device, characterized in that, Including: A first acquisition module for acquiring an image to be classified. An analysis module for serially processing the image to be classified using a vision transformer network to obtain a first processing result. The first processing result is processed by a transformer network structure to obtain a second processing result and a third processing result, where the third processing result is a self-attention weight matrix, and the self-attention weight matrix includes: weight values obtained by performing a layer-by-layer multiplication operation on the normalized self-attention weights; the second processing result is subjected to a convolutional dimensionality reduction process to obtain a plurality of second feature maps, where the plurality of second feature maps are feature maps with semantic class awareness; the plurality of second feature maps are subjected to global average pooling to obtain a first class prediction result, where the first class prediction result is a class prediction result at the global scale; a plurality of third feature maps are selected from the plurality of second feature maps based on the first class prediction result; the third processing result and the plurality of third feature maps are subjected to a repositioning process to obtain a first feature map, where the first feature map is a feature map with class semantic awareness and spatial relationship awareness. A second acquisition module for obtaining a region magnification image from the image to be classified based on the first feature map, where the region magnification image is a local instance region in the image to be classified. A classification module for predicting the attribution class of the image to be classified using the region magnification image.
8. A non - volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, where when the program runs, it controls the device where the non-volatile storage medium is located to execute the image classification method according to any one of claims 1 to 6.
9. A processor, characterized in that, The processor is used to run a program, wherein when the program runs, it executes the image classification method according to any one of claims 1 to 6.
10. An electronic device, characterized in that, Comprising: a processor; and a memory, connected to the processor, for providing instructions for the processor to process the following processing steps: Step 1, obtaining an image to be classified; Step 2, performing serialization processing on the image to be classified by using a vision transformer network to obtain a first processing result; Performing transformation processing on the first processing result by using a transformer network structure to obtain a second processing result and a third processing result, wherein the third processing result is a self-attention weight matrix, and the self-attention weight matrix includes: weight values obtained by performing layer-by-layer cumulative multiplication on the normalized self-attention weights; performing convolutional dimensionality reduction processing on the second processing result to obtain a plurality of second feature maps, wherein the plurality of second feature maps are feature maps with semantic category perception; performing global average pooling processing on the plurality of second feature maps to obtain a first category prediction result, wherein the first category prediction result is a category prediction result at the global scale; selecting a plurality of third feature maps from the plurality of second feature maps based on the first category prediction result; performing repositioning processing on the third processing result and the plurality of third feature maps to obtain a first feature map, wherein the first feature map is a feature map with category semantic perception and spatial relationship perception; Step 3, based on the first feature map, obtaining a region magnification image from the image to be classified, wherein the region magnification image is a local instance region in the image to be classified; Step 4, predicting the belonging category of the image to be classified by using the region magnification image.
Citation Information
Patent Citations
Image processing method and computer readable storage medium
CN111428807A