Image classification method, device, electronic device and storage medium

By using multi-layer neural networks and multi-label classification deep learning models for feature extraction and self-attention calculation in multi-label image classification task, the problem of local information loss in convolutional full-value shared networks is solved, and the accuracy of image classification is improved.

CN114120034BActive Publication Date: 2025-05-13BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111350142.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-05-13
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

In the multi-label image classification task, the convolution full-value sharing network may lose local information or the local information is inaccurate after multi-layer convolution, resulting in a reduced classification accuracy.

Method used

By obtaining the images to be classified and inputting them into the preset multi-layer neural network for feature extraction, the output results of each layer of neural network are obtained, and then inputting these output results into the multi-label classification deep learning model for self-attention calculation, obtaining the probability that the image belongs to each category.

Benefits of technology

Image classification is performed by combining the feature extraction results of each layer of neural network to reduce the situation where local details disappear after multi-layer convolution, thereby achieving accurate recognition of small objects in the image and improving the accuracy of multi-label image classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120034B_ABST
    Figure CN114120034B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image classification method, device, electronic device and storage medium, including: obtaining an image to be classified; inputting the image to be classified into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network; inputting the output results of multiple preset layers of neural networks into a multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category. In this way, image classification is performed in combination with the feature extraction results of the image to be classified in each layer of the neural network. Since the feature extraction results of the underlying neural network often contain more local detail information, the image classification results will also refer to these local detail information, reducing the situation where the local detail information disappears after multiple layers of convolution, thereby achieving accurate recognition of small targets in the image to be classified and improving the accuracy of the multi-label image classification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and in particular to an image classification method, device, electronic device and storage medium. Background Art

[0002] The multi-label image classification task aims to classify an image into all the categories it should belong to. Unlike the most common single-label image classification tasks, multi-label image classification tasks usually have multiple labels, which makes multi-label classification networks more difficult to implement and more challenging. The usual classification is based on all entities that have appeared in the image. In actual applications, there is often more than one entity in an image, so the multi-label classification task is a problem with a wider range of application scenarios than the single-label classification network, and it can be found in many downstream tasks.

[0003] In the multi-label image classification task, a convolutional full-value shared network is used to gradually increase the receptive field of each pixel from bottom to top. Its original intention is to gradually obtain the upper-level information with a high semantic level using layer-by-layer convolution kernels. However, in the above process, a lot of local information may be lost or the local information may be incorrect when reaching the upper layer after multiple layers of convolution, which reduces the accuracy of the multi-label image classification task. Summary of the invention

[0004] The present disclosure provides an image classification method, device, electronic device and storage medium to at least solve the problem in the related art that a lot of local information may be lost in the multi-label image classification task or the local information is not correct after passing through multiple layers of convolution to the upper layer, which reduces the accuracy of the multi-label image classification task. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image classification method, comprising:

[0006] Get the image to be classified;

[0007] Inputting the image to be classified into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network;

[0008] The output results of multiple preset layer neural networks are input into a multi-label classification deep learning model for feature extraction to obtain the probability that the image to be classified belongs to each category.

[0009] Optionally, the step of inputting the image to be classified into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network includes:

[0010] Performing image enhancement processing on the image to be classified, and cutting the image to be classified after the image enhancement processing into sub-images of a preset size;

[0011] The sub-image is input into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network.

[0012] Optionally, the step of inputting the sub-image into a preset multi-layer neural network for feature extraction to obtain an output result of each layer of the neural network includes:

[0013] The sub-images are sequentially input into each layer of the neural network for feature extraction, and corresponding pooling operations are performed on the feature extraction results of each layer of the neural network to obtain the output results of each layer of the neural network.

[0014] Optionally, the step of inputting the output result of each layer of the neural network into a multi-label classification deep learning model for feature extraction to obtain the probability that the image to be classified belongs to each category includes:

[0015] Adjust the output results of multiple preset layers of neural networks to a preset size to obtain an intermediate image;

[0016] Inputting the intermediate image into a multi-label classification deep learning model for self-attention calculation to obtain an attention distribution map, wherein the value of each pixel in the attention distribution map is proportional to the correlation between the pixel at the same position in the intermediate image and the intermediate image;

[0017] The output results of the plurality of preset layer neural networks are weighted according to the attention distribution map, and the weighted processing results are input into a multi-layer perceptron to obtain the probability that the image to be classified belongs to each category.

[0018] Optionally, the step of inputting the output results of the neural networks of the plurality of preset layers into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category includes:

[0019] The output results of the last three layers of the neural network are input into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category.

[0020] According to a second aspect of an embodiment of the present disclosure, there is provided an image classification device, including:

[0021] An acquisition unit, configured to acquire an image to be classified;

[0022] A feature extraction unit is configured to perform feature extraction by inputting the image to be classified into a preset multi-layer neural network to obtain an output result of each layer of the neural network;

[0023] The classification unit is configured to input the output results of multiple preset layer neural networks into a multi-label classification deep learning model for feature extraction to obtain the probability that the image to be classified belongs to each category.

[0024] Optionally, the feature extraction unit is configured to perform:

[0025] Performing image enhancement processing on the image to be classified, and cutting the image to be classified after the image enhancement processing into sub-images of a preset size;

[0026] The sub-image is input into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network.

[0027] Optionally, the feature extraction unit is configured to perform:

[0028] The sub-images are sequentially input into each layer of the neural network for feature extraction, and corresponding pooling operations are performed on the feature extraction results of each layer of the neural network to obtain the output results of each layer of the neural network.

[0029] Optionally, the classification unit is configured to execute:

[0030] Adjust the output results of multiple preset layers of neural networks to a preset size to obtain an intermediate image;

[0031] Inputting the intermediate image into a multi-label classification deep learning model for self-attention calculation to obtain an attention distribution map, wherein the value of each pixel in the attention distribution map is proportional to the correlation between the pixel at the same position in the intermediate image and the intermediate image;

[0032] The output results of the plurality of preset layer neural networks are weighted according to the attention distribution map, and the weighted processing results are input into a multi-layer perceptron to obtain the probability that the image to be classified belongs to each category.

[0033] Optionally, the classification unit is configured to execute:

[0034] The output results of the last three layers of the neural network are input into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category.

[0035] According to a third aspect of an embodiment of the present disclosure, there is provided an image classification electronic device, including:

[0036] processor;

[0037] a memory for storing instructions executable by the processor;

[0038] Wherein, the processor is configured to execute the instructions to implement the image classification method described in the first item above.

[0039] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an image classification electronic device, the image classification electronic device is enabled to perform the image classification method described in the first item above.

[0040] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the image classification method described in the first item above is implemented.

[0041] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0042] Obtain an image to be classified; input the image to be classified into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network; input the output result of each layer of the neural network into a multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category.

[0043] In this way, image classification is performed in combination with the feature extraction results of the image to be classified in each layer of the neural network. Since the feature extraction results of the underlying neural network often contain more local detail information, the image classification results will also refer to these local detail information, reducing the situation where local detail information disappears after multiple layers of convolution. This can achieve accurate recognition of small targets in the image to be classified and improve the accuracy of multi-label image classification tasks.

[0044] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0046] Figure 1 The figure is a flowchart of an image classification method according to an exemplary embodiment.

[0047] Figure 2 It is a multi-label classification deep learning model architecture diagram shown according to an exemplary embodiment.

[0048] Figure 3 It is a network overall architecture diagram of an image classification method according to an exemplary embodiment.

[0049] Figure 4 The figure is a block diagram of an image classification device according to an exemplary embodiment.

[0050] Figure 5 The invention is a block diagram of an electronic device for image classification according to an exemplary embodiment.

[0051] Figure 6 The figure is a block diagram of a device for image classification according to an exemplary embodiment. DETAILED DESCRIPTION

[0052] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0054] Figure 1 is a flowchart of an image classification method according to an exemplary embodiment. Figure 1 As shown, the image classification method includes the following steps.

[0055] In step S11, an image to be classified is obtained.

[0056] In some scenarios, images often need to be classified. Since images contain a large amount of information, an image may belong to multiple different categories at the same time. For example, an image may be a face image taken at night. In this case, the image belongs to both a face image and a night scene image. Therefore, multi-label classification is required for the image to be classified to meet the image classification requirements in actual scenarios.

[0057] In the disclosed embodiment, the image to be classified can be an image of any format and any content, and is not specifically limited. The image to be classified contains rich local detail information, which will also affect the classification result of the image to be classified.

[0058] In step S12, the image to be classified is input into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network.

[0059] In this step, the preset multi-layer neural network is used to extract the image features of the image to be classified. It can be understood that in the multi-layer neural network, each layer needs to perform convolution calculations on the image to be classified. Through the convolution kernels layer by layer, the image features of the image to be classified can be extracted, and the upper-level information with a high semantic level can be gradually obtained. However, in the above calculation process, the local detail information of the image to be classified may also be lost or have large deviations.

[0060] In one implementation, before the image to be classified is input into a preset multi-layer neural network for feature extraction, the image to be classified can be first subjected to image enhancement processing, and the image to be classified after the image enhancement processing can be cropped into sub-images of a preset size; then, the sub-images are input into a preset multi-layer neural network for feature extraction to obtain the output results of each layer of the neural network.

[0061] For example, the image enhancement processing may include any one or more processing methods such as Rand-Augment (random enhancement) processing, Mixup (image mixing) processing, and CutMix (cutting and mixing) processing.

[0062] Among them, Rand-Augment is an unsupervised image enhancement method. It can randomly select several of the common data enhancement methods according to the image to be classified and the size of the preset multi-layer neural network to perform image enhancement processing on the image to be classified. Mixup is an algorithm used in computer vision to perform mixed-class enhancement on images, which refers to mixing the image to be classified with other randomly selected data in a preset ratio. CutMix refers to cropping a part of the image to be classified and randomly filling the cropped area with the pixel values ​​of other data.

[0063] Among them, image enhancement processing can purposefully emphasize the overall or local characteristics of the image to be classified, such as improving the color, brightness and contrast of the image to be classified, making the originally unclear image to be classified clear or emphasizing certain interesting features, expanding the difference between the features of different objects in the image to be classified, suppressing uninteresting features, improving the visual effect of the image to be classified, etc. In other words, after image enhancement processing, the feature information of the image to be classified is enhanced, and in the subsequent feature extraction step, more significant and accurate feature extraction results can be obtained.

[0064] The preset size of the sub-image can be determined according to the requirements of the preset multi-layer neural network for the input image, for example, it can be 224*224*3. In this way, the obtained sub-image better meets the requirements of the preset multi-layer neural network, and the obtained feature extraction result is more accurate.

[0065] In one implementation, a sub-image is input into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network, including: inputting the sub-image into each layer of the neural network in turn for feature extraction, performing corresponding pooling operations on the feature extraction results of each layer of the neural network, and obtaining the output result of each layer of the neural network.

[0066] Specifically, firstly, the first layer of the neural network in the preset multi-layer neural network is used as the target neural network, the sub-image is used as the target input image, the target input image is input into the target neural network for feature extraction, and the processing result is subjected to the pooling operation corresponding to the target neural network to obtain the output result of the target neural network;

[0067] Then, the next layer of the target neural network is used as a new target neural network, the output result of the target neural network is used as a new target input image, and the step of inputting the target input image into the target neural network for feature extraction is returned.

[0068] Among them, the pooling operation is to downsample the processing results, and maximum pooling, average pooling, overlapping pooling and pyramid pooling can be adopted, and the present disclosure does not limit this. By performing a pooling operation on the processing results of each layer of the neural network, the processing results can be reduced in dimension, so that the feature dimension of the output result is reduced, thereby effectively reducing the network parameters of the neural network and preventing the overfitting phenomenon of the neural network.

[0069] In step S13, the output results of the neural networks of multiple preset layers are input into a multi-label classification deep learning model for feature extraction to obtain the probability that the image to be classified belongs to each category.

[0070] After obtaining the feature extraction results of the image to be classified, the image to be classified can be subjected to multi-label classification according to the feature extraction results. The multi-label classification deep learning model can be a Transformer model, or a neural network model or an LSTM (Long Short-Term Memory) model, etc.

[0071] In one implementation, the output results of multiple preset layer neural networks are input into a multi-label classification deep learning model for feature extraction, which may specifically include the following steps:

[0072] First, the output results of multiple preset layers of neural networks are adjusted to a preset size to obtain an intermediate image;

[0073] Then, the intermediate image is input into the multi-label classification deep learning model for self-attention calculation to obtain an attention distribution map. The value of each pixel in the attention distribution map is proportional to the correlation between the pixel at the same position in the intermediate image and the intermediate image.

[0074] Furthermore, the output results of multiple preset layers of neural networks are weighted according to the attention distribution map, and the weighted processing results are input into the multi-layer perceptron to obtain the probability that the image to be classified belongs to each category.

[0075] It can be understood that when extracting features from the image to be classified, the output results of each layer of the neural network have different sizes. Therefore, it is necessary to adjust the output results of multiple preset layers of neural networks to a uniform size before they can be input into the label classification deep learning model for self-attention calculation. Self-attention calculation can filter out a small amount of important information from a large amount of feature information of the image to be classified, and through weighted processing, the multi-layer perceptron focuses on this important information, thereby reducing the amount of calculation and obtaining more accurate image classification results.

[0076] like Figure 2 The figure shows the architecture of a multi-label classification deep learning model. The multi-label classification deep learning model specifically uses the Transformer model, which mainly includes self-attention calculation and multi-layer perceptron (MLP).

[0077] Among them, a1, a2, a3 and a4 are input information, and the input information is a one-dimensional vector with a dimension of (d, 1). The Transformer model performs matrix multiplication operations with each input information through three fully connected (FC) layers Wq, Wk and Wv to obtain the feature information corresponding to each input information in the query, key and value feature spaces, respectively. Among them, the dimension of the feature information is (d, 1). For example, the feature information corresponding to a1 in the query, key and value feature spaces is q 1 , k 1 and v 1 .

[0078] Then, for q 1 ,q 2 ,q 3 ,q 4 and k 1 , k 2 , k 3 , k 4 Perform the scaled inner product separately, for example, q 1 and k 1 Perform the scaled inner product to get α 1,1 , q 1 and k 2Perform the scaled inner product to get α 1,2 , q 1 and k 3 Perform the scaled inner product to get α 1,3 , q 1 and k 4 Perform the scaled inner product to get α 1,4 , and so on, and then input the calculation results into the activation layer (softmax), so that the probability weights of other pixels for each pixel are obtained, that is, the attention distribution map. For example, the probability weight of a1 about a1 is expressed as The probability weight of a2 with respect to a1 is expressed as The probability weight of a3 with respect to a1 is expressed as The probability weight of a4 with respect to a1 is expressed as Etc. Then, the weighted processing result is input into the multi-layer perceptron to obtain the probability that the image to be classified belongs to each category.

[0079] In this step, the multiple preset layer neural networks refer to certain layers in the preset multi-layer neural network, which can be every layer or certain specified layers, and there is no specific limitation.

[0080] For example, it can be the last three layers in a preset multi-layer neural network, that is, the output results of the last three layers of the neural network are input into a multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category.

[0081] It can be understood that the higher the number of layers of the neural network, the more abstract the extracted image features are, the higher the semantic level, and the less underlying detail information it contains. Although the low-level neural network contains more underlying detail information, the extracted feature information is not representative. Therefore, by selecting the output results of the last three layers, on the one hand, the underlying detail information can be retained, and on the other hand, the accuracy of the feature extraction results can be guaranteed.

[0082] like Figure 3 The figure shows the overall network architecture of the image classification method proposed in the present disclosure. The dimension of the image to be classified is H×W×3, and the preset multi-layer neural network has 4 layers. First, the image to be classified passes through the first layer of the neural network with a 7×7 convolution kernel, and the dimension of the output result is Then, each neural network layer is input layer by layer, and the dimensions of the output results are and

[0083] Each neural network layer consists of several layers of neurons (blocks). Specifically, if the dimension of the image to be classified is 224×224×3, the number of neurons from the first layer to the fourth layer is 3, 4, 23, and 3 respectively. Each neuron is composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in series. There is a pooling layer between each layer of the neural network for pooling operations. For example, the average of each four adjacent points can be used to obtain a result, and finally the output result is obtained.

[0084] Since the dimensions of the output results of the last three layers are different, namely 28×28×512, 14×14×1024 and 7×7×2048, they cannot be directly used for self-attention calculation using the Transformer model. Instead, they need to be adjusted to the same size of 7×7×512 using a convolutional network. Then, in the output results of the last three layers after the adjustment, each pixel is used as a feature point, and the internal value of the feature point is determined according to the value of the feature point on 512 channels. The feature points are then concatenated together to obtain 147×512 input features, which are input into the Transformer model for self-attention calculation. Furthermore, the output of the Transformer model is connected to the multi-layer perceptron to extract the number and dimension of categories. After the activation layer, the predicted score of the image to be classified in each category is obtained, that is, the probability that the image to be classified belongs to each category, and the classification result of the image to be classified is obtained.

[0085] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure performs image classification in combination with the feature extraction results of the image to be classified in each layer of the neural network. Since the feature extraction results of the underlying neural network often contain more local detail information, the image classification results will also refer to these local detail information, thereby reducing the disappearance of local detail information after multiple layers of convolution, thereby achieving accurate recognition of small targets in the image to be classified and improving the accuracy of multi-label image classification tasks.

[0086] Figure 4 is a block diagram of an image classification device according to an exemplary embodiment, the device comprising:

[0087] An acquisition unit 401 is configured to acquire an image to be classified;

[0088] The feature extraction unit 402 is configured to perform feature extraction by inputting the image to be classified into a preset multi-layer neural network to obtain an output result of each layer of the neural network;

[0089] The classification unit 403 is configured to execute the input of the output results of multiple preset layer neural networks into a multi-label classification deep learning model for feature extraction to obtain the probability that the image to be classified belongs to each category.

[0090] In one implementation, the feature extraction unit is configured to perform:

[0091] Performing image enhancement processing on the image to be classified, and cutting the image to be classified after the image enhancement processing into sub-images of a preset size;

[0092] The sub-image is input into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network.

[0093] In one implementation, the feature extraction unit is configured to perform:

[0094] The sub-images are sequentially input into each layer of the neural network for feature extraction, and corresponding pooling operations are performed on the feature extraction results of each layer of the neural network to obtain the output results of each layer of the neural network.

[0095] In one implementation, the classification unit is configured to execute:

[0096] Adjust the output results of multiple preset layers of neural networks to a preset size to obtain an intermediate image;

[0097] Inputting the intermediate image into a multi-label classification deep learning model for self-attention calculation to obtain an attention distribution map, wherein the value of each pixel in the attention distribution map is proportional to the correlation between the pixel at the same position in the intermediate image and the intermediate image;

[0098] The output results of multiple preset layers of neural networks are weighted according to the attention distribution map, and the weighted processing results are input into a multi-layer perceptron to obtain the probability that the image to be classified belongs to each category.

[0099] In one implementation, the classification unit is configured to execute:

[0100] The output results of the last three layers of the neural network are input into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category.

[0101] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure performs image classification in combination with the feature extraction results of the image to be classified in each layer of the neural network. Since the feature extraction results of the underlying neural network often contain more local detail information, the image classification results will also refer to these local detail information, thereby reducing the disappearance of local detail information after multiple layers of convolution, thereby achieving accurate recognition of small targets in the image to be classified and improving the accuracy of multi-label image classification tasks.

[0102] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0103] Figure 5 The invention is a block diagram of an electronic device for image classification according to an exemplary embodiment.

[0104] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor of an electronic device to perform the above method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0105] In an exemplary embodiment, a computer program product is also provided. When the computer program product is executed on a computer, the computer implements the above-mentioned image classification method.

[0106] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure performs image classification in combination with the feature extraction results of the image to be classified in each layer of the neural network. Since the feature extraction results of the underlying neural network often contain more local detail information, the image classification results will also refer to these local detail information, thereby reducing the disappearance of local detail information after multiple layers of convolution, thereby achieving accurate recognition of small targets in the image to be classified and improving the accuracy of multi-label image classification tasks.

[0107] Figure 6 is a block diagram of a device 800 for image classification according to an exemplary embodiment.

[0108] For example, the apparatus 800 may be a mobile phone, a computer, a digital broadcast electronic device, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0109] Reference Figure 6 , the device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0110] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0111] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0112] The power supply component 807 provides power to the various components of the device 800. The power supply component 807 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.

[0113] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0114] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the device 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0115] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0116] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, and the sensor assembly 814 can also detect the position change of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0117] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0118] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the methods described in the first and second aspects.

[0119] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the device 800 to complete the above method. Optionally, for example, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0120] In an exemplary embodiment, a computer program product containing instructions is also provided. When the computer program product is run on a computer, the computer is enabled to execute the image classification method described in the first embodiment above.

[0121] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure performs image classification in combination with the feature extraction results of the image to be classified in each layer of the neural network. Since the feature extraction results of the underlying neural network often contain more local detail information, the image classification results will also refer to these local detail information, thereby reducing the disappearance of local detail information after multiple layers of convolution, thereby achieving accurate recognition of small targets in the image to be classified and improving the accuracy of multi-label image classification tasks.

[0122] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0123] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image classification method, characterized in that: include: Get the image to be classified; Inputting the image to be classified into a preset multi-layer neural network for feature extraction processing to obtain the output result of each layer of the neural network; Inputting the output results of the neural networks of the multiple preset layers into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category; The output results of the plurality of preset layer neural networks are input into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category, including: Adjusting the output results of the neural networks of the plurality of preset layers to a second preset size to obtain an intermediate image; Inputting the intermediate image into a multi-label classification deep learning model for self-attention calculation to obtain an attention distribution map, wherein the value of each pixel in the attention distribution map is proportional to the correlation between the pixel at the same position in the intermediate image and the intermediate image; The output results of the plurality of preset layer neural networks are weighted according to the attention distribution map, and the weighted processing results are input into a multi-layer perceptron to obtain the probability that the image to be classified belongs to each category.

2. The image classification method according to claim 1, characterized in that: The step of inputting the image to be classified into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network includes: Performing image enhancement processing on the image to be classified, and cutting the image to be classified after the image enhancement processing into sub-images of a first preset size; The sub-image is input into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network.

3. The image classification method according to claim 2, characterized in that: The step of inputting the sub-image into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network includes: The sub-images are sequentially input into each layer of the neural network for feature extraction, and corresponding pooling operations are performed on the feature extraction results of each layer of the neural network to obtain the output results of each layer of the neural network.

4. The image classification method according to claim 1, characterized in that: The output results of the plurality of preset layer neural networks are input into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category, including: The output results of the last three layers of the neural network are input into the multi-label classification deep learning model for self-attention calculation to obtain the probability that the image to be classified belongs to each category.

5. An image classification device, characterized in that: include: An acquisition unit, configured to acquire an image to be classified; A feature extraction unit is configured to perform feature extraction by inputting the image to be classified into a preset multi-layer neural network to obtain an output result of each layer of the neural network; A classification unit is configured to input the output results of the plurality of preset layers of neural networks into a multi-label classification deep learning model for feature extraction, and obtain the probability that the image to be classified belongs to each category; The classification unit is configured to perform: Adjust the output results of multiple preset layers of neural networks to a preset size to obtain an intermediate image; Inputting the intermediate image into a multi-label classification deep learning model for self-attention calculation to obtain an attention distribution map, wherein the value of each pixel in the attention distribution map is proportional to the correlation between the pixel at the same position in the intermediate image and the intermediate image; The output results of the plurality of preset layer neural networks are weighted according to the attention distribution map, and the weighted processing results are input into a multi-layer perceptron to obtain the probability that the image to be classified belongs to each category.

6. The image classification device according to claim 5, characterized in that: The feature extraction unit is configured to perform: Performing image enhancement processing on the image to be classified, and cutting the image to be classified after the image enhancement processing into sub-images of a preset size; The sub-image is input into a preset multi-layer neural network for feature extraction to obtain the output result of each layer of the neural network.

7. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image classification method as described in the first item of claims 1 to 4.

8. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor, the processor is enabled to perform the image classification method as described in the first of claims 1 to 4.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the image classification method described in the first item of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Picture feature extraction method and device, target re-identification method and device and electronic equipment

    CN111310518A

  • Transform-based thermal imaging method for detecting internal defects of pipeline

    CN113034469A