Object Recognition Method and Device, Storage Medium, and Electronic Device
By training a neural network to focus on attention distribution at the feature level through attention map extraction, the method improves multi-label classification accuracy in neural network models.
Patent Information
- Application Number
- CN202011074278.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-10-09
AI Technical Summary
In the prior art, the neural network model does not pay attention to the attention distribution of the classifier at the feature layer during multi-label classification, resulting in a low accuracy of classification results.
By acquiring the target image, the attention map of each target object is extracted using the target recognition neural network, and the classification labels and locations of the objects are identified in the target recognition neural network, and the initial classification neural network and the initial recognition neural network are trained until the first convergence condition is reached to correct the output results of the recognition neural network.
The accuracy of the results of the neural network model for multi-label classification is improved, ensuring that the attention distribution of the classifier at the feature layer is paid attention to, and improving the accuracy of multi-label classification.
Smart Images

Figure CN112132231B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular, to a method and apparatus for identifying an object, a storage medium, and an electronic device. Background Art
[0002] With the development of artificial intelligence, using a neural network model for pattern recognition is the current trend of development. Some achievements have been made in using a convolutional neural network model for image recognition. For example, multi-label classification is performed through a convolutional neural network model. That is, an image containing multiple objects to be recognized is input into the convolutional neural network, and then the convolutional neural network outputs the recognition result.
[0003] Currently, when performing multi-label classification through a neural network model, each label is regarded as a binary classification task during the training process of the model. When using a convolutional neural network for feature extraction, for a given picture, multiple features can be extracted through an attention mechanism, and then these features are used for multi-label classification. However, since the multi-labels in the training dataset are associated with each other. The attention distribution of the classifier in the feature layer is not emphasized during the training process of the neural network model. It may occur that the attention region corresponding to it on the image is not interpretable. As a result, the accuracy of the multi-label classification result is relatively low.
[0004] Aiming at the problem in the related art that due to the fact that the attention distribution of the classifier in the feature layer is not emphasized during the training process of the neural network model in the prior art, the accuracy of the multi-label classification result of the neural network model is relatively low, there is currently no effective solution. Summary of the Invention
[0005] Embodiments of the present invention provide a method and apparatus for identifying an object, a storage medium, and an electronic device, so as to at least solve the technical problem that the accuracy of the multi-label classification result of the neural network model is relatively low because the attention distribution of the classifier in the feature layer is not emphasized during the training process of the neural network model in the prior art.
[0006] According to one aspect of an embodiment of the present invention, there is provided a method for identifying an object, including: obtaining a target image, where the target image includes at least one target object to be identified; extracting, in a target recognition neural network, an attention map corresponding to each of the target objects in the target image, where the target recognition neural network is obtained by training based on an initial classification neural network and an initial recognition neural network until a first convergence condition is reached, and the first convergence condition indicates that an output value of a loss function between a first output result of the classification neural network in training and a second output result of the recognition neural network is within a first preset range, the classification neural network is used to identify an image including a single-label object, and the recognition neural network is used to identify an image including a multi-label object; in the target recognition neural network, identifying, based on the attention map, a classification label corresponding to each of the target objects and a position of each of the target objects in the target image.
[0007] According to another aspect of an embodiment of the present invention, there is further provided an apparatus for identifying an object, including: an obtaining module, configured to obtain a target image, where the target image includes at least one target object to be identified; an extracting module, configured to extract, in a target recognition neural network, an attention map corresponding to each of the target objects in the target image, where the target recognition neural network is obtained by training based on an initial classification neural network and an initial recognition neural network until a first convergence condition is reached, and the first convergence condition indicates that an output value of a loss function between a first output result of the classification neural network in training and a second output result of the recognition neural network is within a first preset range, the classification neural network is used to identify an image including a single-label object, and the recognition neural network is used to identify an image including a multi-label object; a recognition module, configured to identify, in the target recognition neural network, a classification label corresponding to each of the target objects and a position of each of the target objects in the target image based on the attention map.
[0008] Optionally, the above device is further configured to obtain a plurality of sample images before obtaining the target image, where the plurality of sample images include: a first sample image set and a second sample image set, the first sample image set includes: sample images containing single-label objects and sample images without labels, and the second sample image set includes: sample images containing single-label objects, sample images containing multi-label objects, and sample images without labels; training the initial classification neural network using the first sample image set to obtain a first attention map corresponding to the first sample image set, and the first output result includes the first attention map; training the initial recognition neural network using the second sample image set to obtain a second attention map corresponding to the second sample image set, and the second output result includes the second attention map; determining that the first convergence condition is met when the output value of the loss function between the attention map corresponding to the single-label object in the first attention map and the attention map corresponding to the single-label object in the second attention map is within the first preset range.
[0009] Optionally, the above device is further configured to implement training the initial classification neural network using the first sample image set to obtain a first attention map corresponding to the first sample image set in the following manner: obtaining N groups of first sample image subsets in the first sample image set, where each image in each group of the first sample image subsets includes a single-label object and a plurality of the sample images without labels, and N is an integer greater than or equal to 1; training N classification sub-networks in the initial classification neural network respectively using the N groups of first sample image subsets to obtain N groups of object attention maps, where the first attention map includes the N groups of object attention maps, and each group of the object attention maps includes the object attention map corresponding to the single-label object.
[0010] Optionally, the above device is further configured to implement training N classification sub-networks in the initial classification neural network respectively using the N groups of first sample image subsets to obtain N groups of object attention maps in the following manner: obtaining the i-th group of the first sample image subsets, where i is an integer greater than or equal to 1 and less than or equal to N; inputting the i-th group of the first sample image subsets into the i-th classification sub-network to obtain the i-th group of the object attention maps; saving the i-th group of the object attention maps.
[0011] Optionally, the above device is further configured to implement the training of the above initial recognition neural network using the above second sample image set to obtain the second attention map corresponding to the above second sample image set in the following manner: training the above initial recognition neural network with the above second sample image set for j rounds to obtain the attention map of the j-th round, where j is greater than or equal to 1; when it is determined that the output value of the loss function between the attention maps corresponding to the M single-label objects in the attention map of the j-th round and the attention maps corresponding to the M single-label objects included in the above N groups of object attention maps is within the above first preset range, it is determined that the above first convergence condition is reached, where M is greater than or equal to 1 and less than N, and the above M single-label objects include the above target object.
[0012] Optionally, the above device is further configured to implement the training of the above initial classification neural network using the above first sample image set to obtain the first attention map corresponding to the above first sample image set in the following manner: training the above initial recognition neural network with the above N groups of first sample image subsets to obtain N groups of initial object attention maps; performing a first process on the above N groups of initial object attention maps to obtain a first processing result; performing binarization processing and Gaussian filtering processing on the above first processing result to obtain the above first attention map.
[0013] Optionally, the above device is further configured to implement the above first process on the above N groups of initial object attention maps to obtain a first processing result in the following manner: calculating the product of the mean of the above N groups of initial object attention maps and a preset value to obtain a first adjustment value; determining the difference between the above N groups of initial object attention maps and the above first adjustment value to obtain the above first processing result.
[0014] Optionally, the above device is further configured to implement the training of the above initial recognition neural network using the above second sample image set to obtain the second attention map corresponding to the above second sample image set in the following manner: training the above initial recognition neural network with the above second sample image set to obtain the initial attention map corresponding to the above second sample image set; calculating the difference between the initial attention map corresponding to the above second sample image set and the mean of the initial attention map corresponding to the above second sample image set to obtain a second processing result; performing normalization processing on the above second processing result to obtain the above second attention map.
[0015] Optionally, the above device is further configured such that the estimated classification label of the labeled object output by the above target recognition neural network and the known classification label of the above labeled object satisfy a second convergence condition, where the second convergence condition is used to indicate that the output value of the loss function between the above estimated classification label and the above known classification label is within a second preset range; the estimated position of the labeled object output by the above target recognition neural network and the known position of the above labeled object in the sample image satisfy a third convergence condition, where the third convergence condition is used to indicate that the output value of the loss function between the above estimated position and the above known position is within a third preset range; where the above labeled object includes the above single-label object and the above multi-label object, and the above sample image includes the sample image of the above single-label object and the sample image of the above multi-label object.
[0016] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the above object recognition method when running.
[0017] According to another aspect of the embodiments of the present invention, there is also provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to execute the above object recognition method through the above computer program.
[0018] In the embodiments of the present invention, the first output result of the classification neural network is used to correct the second output result of the recognition neural network. When training based on the initial classification neural network and the initial recognition neural network until the target recognition neural network obtained when reaching the first convergence condition, the target neural network is used to classify and recognize at least one target object in the target image. The first convergence condition indicates that the output value of the loss function between the first output result of the classification neural network in training and the second output result of the recognition neural network is within a first preset range, achieving the purpose of paying attention to the attention distribution of the classifier in the feature layer during the training process of the neural network model, thereby realizing the technical effect of improving the result accuracy of the multi-label classification of the neural network model, and further solving the technical problem that the result accuracy of the multi-label classification of the neural network model is relatively low due to the lack of attention to the attention distribution of the classifier in the feature layer in the training process of the existing neural network model. Description of the Drawings
[0019] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0020] Figure 1Schematic diagram of an application environment of an optional object recognition method according to an embodiment of the present invention;
[0021] Figure 2 Schematic flowchart of an optional object recognition method according to an embodiment of the present invention;
[0022] Figure 3 Schematic diagram of training an optional classification sub-network model according to an embodiment of the present invention Figure 1 ;
[0023] Figure 4 Schematic diagram of training an optional classification sub-network model according to an embodiment of the present invention Figure 2 ;
[0024] Figure 5 Schematic diagram of an optional target recognition neural network model according to an embodiment of the present invention;
[0025] Figure 6 Schematic diagram of an optional image acquisition according to an embodiment of the present invention;
[0026] Figure 7 Schematic diagram of training an initial classification neural network CNN0 according to an embodiment of the present invention;
[0027] Figure 8 Schematic diagram of training an initial classification neural network CNN1 according to an embodiment of the present invention;
[0028] Figure 9 Schematic diagram of training an initial recognition neural network CNN2 according to an embodiment of the present invention;
[0029] Figure 10 Schematic diagram of the structure of an optional object recognition device according to an embodiment of the present invention;
[0030] Figure 11 Schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed implementation manners
[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] According to one aspect of the embodiments of the present invention, a method for identifying an object is provided. Optionally, as an alternative implementation, the above-mentioned method for identifying an object can be, but is not limited to, applied to an environment as Figure 1 shown.
[0034] Optionally, in this embodiment, the above-mentioned user device 102 can be a terminal device configured with a client, and can include, but is not limited to, at least one of the following: mobile phone (such as Android mobile phone, iOS mobile phone, etc.), laptop computer, tablet computer, handheld computer, MID (Mobile Internet Devices), PAD, desktop computer, smart TV, etc. The client can be a video client, instant messaging client, browser client, education client, etc. The target image can be obtained through the client installed in the user device. The user device 102 can be configured with a display 108, which can be used to display the target image and the recognition result of the target image. The user device can also be configured with a processor 106 and a memory 104. Among them, the processor 106 is used to process the obtained target image, and the memory 104 is used to store the target image. The above-mentioned network 110 can include, but is not limited to: wired network, wireless network. Among them, the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that implement wireless communication. The above-mentioned server can be a single server, or a server cluster composed of multiple servers, or a cloud server. The above-mentioned server 112 is provided with a database 114 and a processing engine 116. Among them, the database 114 is used to store data, for example, the data of the obtained target image, and the model parameters obtained during the training of the initial classification neural network and the initial recognition neural network, etc. The processing engine 116 can be used to train the initial classification neural network and the initial recognition neural network to obtain a target recognition neural network. It can also be used to identify at least one target object to be recognized included in the target image using the target neural network model.
[0035] The above is only an example, and no limitation is made in this embodiment.
[0036] Specifically, the following steps will be implemented through the above settings of virtual keys:
[0037] As in step S102, obtain a target image, where the target image includes at least one target object to be recognized; as in step S104, extract the attention map corresponding to each of the target objects in the target image in a target recognition neural network, where the target recognition neural network is obtained by training based on an initial classification neural network and an initial recognition neural network until a first convergence condition is reached. The first convergence condition indicates that the output value of the loss function between the first output result of the classification neural network in training and the second output result of the recognition neural network is within a first preset range. The classification neural network is used to recognize images containing single-label objects, and the recognition neural network is used to recognize images containing multi-label objects; as in step S106, in the target recognition neural network, based on the attention map, recognize the classification label corresponding to each of the target objects and the position of each of the target objects in the target image.
[0038] Optionally, as an alternative implementation, as Figure 2 shown, the object recognition method includes:
[0039] Step S202, obtain a target image, where the target image includes at least one target object to be recognized;
[0040] Step S204, extract the attention map corresponding to each of the target objects in the target image in a target recognition neural network, where the target recognition neural network is obtained by training based on an initial classification neural network and an initial recognition neural network until a first convergence condition is reached. The first convergence condition indicates that the output value of the loss function between the first output result of the classification neural network in training and the second output result of the recognition neural network is within a first preset range. The classification neural network is used to recognize images containing single-label objects, and the recognition neural network is used to recognize images containing multi-label objects;
[0041] Step S206, in the target recognition neural network, based on the attention map, recognize the classification label corresponding to each of the target objects and the position of each of the target objects in the target image.
[0042] Through the above steps, the second output result of the recognition neural network is corrected using the first output result of the classification neural network. When training based on the initial classification neural network and the initial recognition neural network until the target recognition neural network is obtained when the first convergence condition is reached, the target neural network is used to classify and recognize at least one target object in the target image. The first convergence condition indicates that the output value of the loss function between the first output result of the classification neural network and the second output result of the recognition neural network during training is within the first preset range, achieving the purpose of paying attention to the attention distribution of the classifier in the feature layer during the training process of the neural network model. Thus, the technical effect of improving the accuracy of the multi-label classification result of the neural network model is achieved, and furthermore, the technical problem that the accuracy of the multi-label classification result of the neural network model is relatively low due to the lack of attention to the attention distribution of the classifier in the feature layer during the training process of the neural network model in the prior art is solved.
[0043] As an optional implementation manner, the above target image may include multiple target objects to be recognized. For example, an image including a mouse and a keyboard at the same time, or an image including target objects such as a mouse, a keyboard, a monitor, a host, a speaker, and a headset at the same time. The types and quantities of objects included in the image are not limited and can be determined according to actual situations.
[0044] As an optional implementation manner, the above initial classification neural network may be a convolutional neural network, and the number of initial classification neural networks may be determined according to the number of objects to be recognized. For example, if the objects to be recognized include a mouse and a keyboard, the initial classification neural networks may be two initial convolutional neural networks, CNN0 and CNN1. The initial classification neural network CNN0 is trained using a training image set only including a mouse. During the process of training CNN0 using the training image set only including a mouse, the first output result output by CNN0 can be obtained, and the first output result includes the attention map of the mouse. Similarly, the initial neural network CNN1 is trained using a training image set only including a keyboard. During the process of training CNN1, the first output result output by CNN1 can be obtained, and the first output result includes the attention map of the keyboard. The above image including a single-label object is an image including only one object to be recognized, such as an image only including a keyboard or an image only including a mouse.
[0045] As an alternative embodiment, the above initial recognition neural network may also be a convolutional neural network, such as CNN2. The images of the multiple objects to be recognized, such as an image containing both a mouse and a keyboard. The training set of the initial neural network includes images containing only single-label objects, as well as images containing multi-label objects, such as images containing only a mouse, images containing only a keyboard, and images containing both a mouse and a keyboard. During the training process of the initial recognition neural network CNN2, a second output result output by CNN2 can be obtained, and the second output result includes an attention map of the keyboard and an attention map of the mouse.
[0046] As an alternative embodiment, since the above initial classification neural network CNN0 is trained using a training image set containing only mice, it can learn the attention distribution of mouse features. The initial classification neural network CNN1 is trained using a training image set containing only keyboards, and it can learn the attention distribution of keyboard features. The attention distribution of the mouse learned by the initial classification neural network CNN0 and the attention distribution of the mouse learned by the keyboard are used to correct the feature maps of the mouse and keyboard learned in the initial recognition neural network CNN2. Specifically, it can be achieved by making the feature map of the mouse output by CNN0 satisfy the first convergence condition with the feature map of the mouse output by CNN2, and the feature map of the keyboard output by CNN1 satisfy the first convergence condition with the feature map of the keyboard output by CNN1. The first convergence condition can be determined according to the actual situation. The above target neural network model may be a CNN2 network model that satisfies the first convergence condition. In this embodiment, only the target neural network model can be used to recognize the mouse in an image containing only a mouse, or only the target neural network model can be used to recognize the keyboard in an image containing only a keyboard, or only the target neural network model can be used to recognize the mouse and keyboard in an image containing both a mouse and a keyboard. The specific recognition results may include the label of the mouse, the label of the keyboard, and the positions of the mouse and keyboard in the target image respectively. In this embodiment, the mouse and keyboard are only used to explain this embodiment, and the objects to be recognized in a specific image can be determined according to the actual situation.
[0047] In the prior art, a trained neural network model is only used to identify samples to be recognized that are the same as the training samples. For example, the CNN0 obtained using mouse training samples is only used to identify mice and cannot identify keyboards, while the CNN1 obtained using keyboard training samples is only used to identify keyboards and cannot identify mice. If one wants to identify both mice and keyboards simultaneously, it is necessary to use training samples that contain both mice and keyboards to train the neural network model. At this time, the training samples of mice and keyboards are associated together, and the training process of the neural network model does not pay attention to the attention distribution between different training samples. In this application, by making the output value of the loss function between the first output result of the classification neural network and the second output result of the recognition neural network within the first preset range, the second output result of the recognition neural network is corrected using the first output result of the classification neural network. The resulting target neural network model can accurately identify mice and can also accurately identify keyboards. For an image that contains both a mouse and a keyboard, it can also accurately identify the positions of the mouse and the keyboard in the image. This can achieve an improvement in the accuracy of multi-label object recognition.
[0048] Optionally, before obtaining the target image, the method further includes: obtaining a plurality of sample images, where the plurality of sample images include: a first sample image set and a second sample image set. The first sample image set includes: sample images containing single-label objects and sample images without labels. The second sample image set includes: sample images containing single-label objects, sample images containing multi-label objects, and sample images without labels; training the initial classification neural network using the first sample image set to obtain a first attention map corresponding to the first sample image set, and the first output result includes the first attention map; training the initial recognition neural network using the second sample image set to obtain a second attention map corresponding to the second sample image set, and the second output result includes the second attention map; determining that the first convergence condition is met when the output value of the loss function between the attention map corresponding to the single-label object in the first attention map and the attention map corresponding to the single-label object in the second attention map is within the first preset range.
[0049] As an optional implementation manner, the first sample image set is an image containing a single-label object. For example, the image only containing a mouse or the image only containing a keyboard. The object without a label can be other objects, such as the background of the image or an object that does not need to be concerned about. The second sample image set includes: images containing single-label objects, such as the image only containing a mouse and the image only containing a keyboard, and images containing multi-label objects, such as the image containing both a mouse and a keyboard.
[0050] As an optional implementation, the initial classification neural network CNN0 is trained using an image set that only contains a mouse, and the initial classification neural network CNN1 is trained using an image set that only contains a keyboard. During the training of CNN0, an attention map of the mouse can be obtained. During the training of CNN1, an attention map of the keyboard can be obtained. The first attention map includes the attention map of the mouse output by the above-mentioned CNN0 and the attention map of the keyboard output by CNN1.
[0051] As an optional implementation, the initial recognition neural network CNN2 is trained using a second sample image set. During the training of CNN2, an attention map of the mouse and an attention map of the keyboard output by CNN2 are obtained, corresponding to the second attention map. In this embodiment, the attention map of the mouse output by CNN0 and the attention map of the mouse output by CNN2 satisfy the first convergence condition, and the attention map of the keyboard output by CNN1 and the attention map of the keyboard output by CNN2 satisfy the first convergence condition. The above-mentioned first convergence condition may be that the output value of the loss function is within a predetermined range. The loss function can be selected according to the actual situation. For example, it can be a cross-entropy function. The first preset range can be determined according to the actual situation. For example, it can be 0.5, 0.3, 0.1, etc.
[0052] Optionally, training the above-mentioned initial classification neural network using the above-mentioned first sample image set to obtain the first attention map corresponding to the above-mentioned first sample image set includes: obtaining N groups of first sample image subsets in the above-mentioned first sample image set, where each image in each group of the above-mentioned first sample image subsets includes a single-label object and multiple above-mentioned unlabeled objects, and N is an integer greater than or equal to 1; using the above-mentioned N groups of first sample image subsets to train N classification sub-networks in the above-mentioned initial classification neural network respectively to obtain N groups of object attention maps, where the above-mentioned first attention map includes the above-mentioned N groups of object attention maps, and each group of the above-mentioned object attention maps includes the object attention map corresponding to the above-mentioned single-label object.
[0053] As an optional implementation, each image included in each group of the above-mentioned N groups of first sample image subsets contains a single-label object and multiple unlabeled objects. For example, it can be the above-mentioned image that only contains a mouse and the image that only contains a keyboard. Each group of the above-mentioned first sample image subsets corresponds to an initial classification neural network, and each group of the above-mentioned first sample image subsets is used to train the corresponding initial classification neural network respectively. For example, using the first sample image subset that only contains a mouse to train CNN0, an attention map of the mouse can be obtained, and using the first sample image subset that only contains a keyboard to train CNN1, an attention map of the keyboard can be obtained.
[0054] Optionally, use the above N groups of first sample image subsets to train the N classification sub-networks in the above initial classification neural network respectively to obtain N groups of object attention maps, including: obtaining the i-th group of the above first sample image subsets, where i is an integer greater than or equal to 1 and less than or equal to N; inputting the i-th group of the above first sample image subsets into the i-th above classification sub-network to obtain the i-th group of the above object attention maps; saving the i-th group of the above object attention maps.
[0055] As an alternative implementation, the i-th group of first sample image subsets may be an image set including single-label objects. For example, it may be an image set containing a mouse. The i-th group of first sample image subsets can be used to train the corresponding classification sub-network to obtain the attention map of the single-label object included in the image subset.
[0056] As a preferred implementation, taking the i-th group of first sample image subsets as an image set containing a mouse and the classification sub-network as CNN0 as an example for illustration, as Figure 3 is a schematic diagram of an alternative classification sub-network model training according to an embodiment of the present invention Figure 1 and specifically may include the following steps:
[0057] Step S1, input the image set A containing a mouse into the classification sub-network CNN0;
[0058] Step S2, during the training of CNN0, the last convolutional layer F is a feature map of size CxHxW. The feature vector w0 for identifying the mouse is convolved with F to obtain the attention map M0;
[0059] Step S3, use the attention map M0 as the weight F for spatial domain weighted summation to obtain the corresponding feature w1, and then calculate the similarity between w0 and w1. When the similarity between w0 and w1 meets the preset condition, stop the training of CNN0. After the training is completed, for all the images in the training set A, extract their attention maps and save them.
[0060] As a preferred implementation, taking the i-th group of first sample image subsets as an image set containing a keyboard and the classification sub-network as CNN1 as an example for illustration, as Figure 4 is a schematic diagram of an alternative classification sub-network model training according to an embodiment of the present invention Figure 2 and specifically may include the following steps:
[0061] Step S1, input the image set B containing a keyboard into the classification sub-network CNN1;
[0062] Step S2, during the training of CNN1, the last convolutional layer F is a feature map of size CxHxW, and the feature vector w3 for identifying the mouse is used. The feature vector w3 is convolved with F to obtain the attention map M1;
[0063] Step S3, use the attention map M1 as the weight F for spatial domain weighted summation to obtain the corresponding feature w4, and then calculate the similarity between w3 and w4. When the similarity between w3 and w4 meets the preset condition, stop the training of CNN1. After the training is completed, for all the images in the training set B, extract their attention maps and save them.
[0064] Optionally, training the above initial recognition neural network with the above second sample image set to obtain the second attention map corresponding to the above second sample image set, including: using the above second sample image set to train the above initial recognition neural network for j rounds to obtain the j-th round attention map, where j is greater than or equal to 1; when it is determined that the output value of the loss function between the attention maps corresponding to the M single-label objects in the j-th round attention map and the attention maps corresponding to the M single-label objects included in the above N groups of object attention maps is within the above first preset range, it is determined that the first convergence condition is reached, where M is greater than or equal to 1 and less than N, and the above M single-label objects include the above target object.
[0065] As an optional implementation manner, the training process of the neural network model is to repeatedly adjust the parameters of the neural network model until the training stops when the preset convergence condition is reached, and a trained neural network model is obtained. In this embodiment, it is assumed that the model obtained by training the initial recognition neural network for j rounds is the j-th round recognition neural network, and the value of j can be an integer greater than or equal to 1. For example, the first round of training is to input the training data set into the initial recognition neural network for the first training, and the second round of training is to adjust the model parameters based on the model parameters obtained in the first round of training to obtain the j-th round recognition neural network. Through continuous iteration until the output result meets the convergence condition, the target neural network model is obtained.
[0066] As an optional implementation manner, when training the initial recognition neural network for j rounds, if the single-label objects output by the obtained j-th round recognition neural network satisfy the first convergence condition with some or all of the object attention maps in the N groups of object attention maps obtained during the training of the initial classification neural network, stop the j-th round recognition neural network, and use the obtained j-th round recognition neural network as the target neural network model.
[0067] As a preferred embodiment, it is assumed that the CNN0 is trained using only the images containing a mouse to obtain the attention map of the mouse, and the CNN1 is trained using only the images containing a keyboard to obtain the attention map of the keyboard. The CNN2 is trained using only the images containing a mouse and only the images containing a keyboard. If, in the j-th round of training of the CNN2, the j-th round of CNN2 network model is obtained, and if the attention map of the mouse output by the j-th round of CNN2 network model satisfies the first convergence condition with the attention map of the mouse output by the CNN0, the training of the j-th round of CNN2 network model can be stopped, and the j-th round of CNN2 network model is used as the target neural network model. This target neural network model can identify images containing a mouse.
[0068] Alternatively, if, in the j-th round of training of the CNN2, the j-th round of CNN2 network model is obtained, and if the attention map of the keyboard output by the j-th round of CNN2 network model satisfies the first convergence condition with the attention map of the keyboard output by the CNN0, the training of the j-th round of CNN2 network model can be stopped, and the j-th round of CNN2 network model is used as the target neural network model. This target neural network model can identify images containing a keyboard.
[0069] Alternatively, if, in the j-th round of training of the CNN2, the j-th round of CNN2 network model is obtained, and if the attention map of the keyboard output by the j-th round of CNN2 network model satisfies the first convergence condition with the attention map of the keyboard output by the CNN0, and the attention map of the mouse output by the j-th round of CNN2 network model satisfies the first convergence condition with the attention map of the mouse output by the CNN0. Stop the training of the j-th round of CNN2 network model, and use the j-th round of CNN2 network model as the target neural network model. This target neural network model can not only identify images containing a keyboard, but also identify images containing a mouse, and can also identify images containing both a mouse and a keyboard.
[0070] In this embodiment, the above-mentioned images containing only a mouse are only used to illustrate that the keyboard is not included in the image, but may include other unlabeled objects, such as the image background, or other image elements. Similarly, the images containing only a keyboard are only used to illustrate that the mouse is not included in the image, and may also include other image elements.
[0071] Optionally, training the above-mentioned initial classification neural network using the above-mentioned first sample image set to obtain the first attention map corresponding to the above-mentioned first sample image set includes: training the above-mentioned initial recognition neural network using the above-mentioned N groups of first sample image subsets to obtain N groups of initial object attention maps; performing a first process on the above-mentioned N groups of initial object attention maps to obtain a first processing result; performing binarization processing and Gaussian filtering processing on the above-mentioned first processing result to obtain the above-mentioned first attention map.
[0072] As an alternative implementation, in order to avoid the initial attention maps obtained by training the initial classification neural network and the initial attention maps obtained by training the initial recognition neural network from having different sizes and shapes, it is necessary to adjust the initial attention maps output by the above models. For example, the size of the mouse obtained when training CNN0 may not be the same as the size of the mouse obtained when training CNN2. Therefore, it is necessary to adjust the sizes of the mice output by CNN0 and CNN2 so that the sizes of the mice output by the two models are the same.
[0073] As a preferred implementation, what is obtained by training the initial recognition neural network with N groups of first sample image subsets is N groups of initial object attention maps. After performing first processing, binary processing, and Gaussian processing on the N groups of initial object attention maps, a first attention map is obtained. In this embodiment, the above processing is performed on the initial attention maps obtained by training the initial classification neural network so that the obtained first attention map is the same size as the second attention map obtained when training the initial recognition neural network.
[0074] Optionally, the above first processing of the above N groups of initial object attention maps to obtain a first processing result includes: calculating the product of the mean of the above N groups of initial object attention maps and a preset value to obtain a first adjustment value; determining the difference between the above N groups of initial object attention maps and the above first adjustment value to obtain the above first processing result.
[0075] As an alternative implementation, taking the mouse attention map M0 as an example of the initial object attention map, the processing of M0 may include the following steps:
[0076] Step S1, calculate the mean of the mouse attention map M0 to obtain M0m;
[0077] Step S2, calculate the product of M0m and a preset value α to obtain a first adjustment value αM0m;
[0078] Step S3, calculate the difference between the mouse attention map M0 and the first adjustment value αM0m, M0 - M0m*α;
[0079] Step S4, perform binarization according to whether M0 - M0m*α is greater than 0 to obtain M0b;
[0080] M0b = Bin(M0 - M0m*α)
[0081] where Bin represents the binarization operation;
[0082] Step S5, perform Gaussian filtering on the binarized M0b, and take the part with a value greater than β as the attention criterion to obtain the first attention map of the mouse.
[0083] The specific Gaussian filtering formula can be:
[0084] M0s = Gaussian(M0b) x (Gaussian(M0b) > β),
[0085] where Gaussian represents Gaussian filtering. In the above description, both α and β are adjustable coefficients. The value of α can be near 1, and the value of β is between 0 and 1.
[0086] Optionally, training the initial recognition neural network using the above second sample image set to obtain the second attention map corresponding to the second sample image set includes: training the initial recognition neural network using the above second sample image set to obtain the initial attention map corresponding to the second sample image set; calculating the difference between the initial attention map corresponding to the second sample image set and the mean of the initial attention map corresponding to the second sample image set to obtain a second processing result; performing normalization processing on the second processing result to obtain the second attention map.
[0087] As an optional implementation, in order to make the first attention consistent with the second attention map, it is necessary to adjust the initial attention map obtained when training the initial recognition neural network. In this embodiment, the initial attention map obtained when training the initial recognition neural network can be normalized.
[0088] Specifically, taking the initial object attention map as the mouse attention map M0' as an example for illustration, the processing of M0' can include the following steps:
[0089] Step S1, calculate the mean of the mouse attention map M0' to obtain M0'm;
[0090] Step S2, calculate the difference between the attention map M0' and the mean M0'm to obtain the second processing result M0' - M0'm;
[0091] Step S3, perform normalization processing on the above second processing result M0' - M0'm. Specifically, normalization processing can be performed through the following formula to obtain the second attention map M0c':
[0092] M0c' = sigmoid(M0' - M0m')
[0093] In this embodiment, the cross-entropy loss function can be used to establish the error Lattn between M0c' and M0s. When the error Lattn value, where the error Lattn can be used to represent the output value of the cross-entropy loss function, is within the first preset range, it is determined that the training of the initial recognition neural network is completed to obtain the target neural network model.
[0094] As an alternative implementation, such as Figure 5 FIG. is a schematic diagram of an alternative target recognition neural network model according to an embodiment of the present invention. Taking the case where the input target image includes both a mouse and a keyboard as an example, since the trained target neural network model can learn the attention distributions of the mouse and the keyboard respectively, an attention map M0' of the mouse and an attention map M1' of the keyboard can be obtained. Therefore, the target neural network model can accurately identify the mouse, can also identify the keyboard, and can also simultaneously identify the mouse and the keyboard, and the recognition result also includes the positions of the mouse and the keyboard in the target image respectively.
[0095] Optionally, the above method further includes at least one of the following: the predicted classification label of the labeled object output by the above target recognition neural network satisfies a second convergence condition with the known classification label of the labeled object, where the second convergence condition is used to indicate that the output value of the loss function between the predicted classification label and the known classification label is within a second preset range; the predicted position of the labeled object output by the above target recognition neural network in the sample image satisfies a third convergence condition with the known position of the labeled object in the sample image, where the third convergence condition is used to indicate that the output value of the loss function between the predicted position and the known position is within a third preset range; where the above labeled object includes the above single-label object and the above multi-label object, and the above sample image includes the sample image of the above single-label object and the sample image of the above multi-label object.
[0096] As an alternative implementation, when training the initial recognition neural network, the classification loss function Lclf1 can also be minimized. The labels of the labeled objects included in the training image set are known, that is, the classification results are known. When training the initial recognition neural network, the predicted classification results of the labeled objects can be obtained. When the error value output by the loss function Lclf1 between the predicted classification results and the known classification results is within a second preset range, it is determined that the second convergence condition is satisfied.
[0097] As an alternative implementation, when training the initial recognition neural network, the classification loss function Lclf2 can also be minimized. The positions of the labeled objects included in the training image set in the image are known. When training the initial recognition neural network, the predicted results of the positions of the labeled objects in the target image are obtained, and the predicted positions can be obtained. When the error value output by the loss function Lclf2 between the predicted positions of the labeled objects in the target image and the known positions is within a third preset range, it is determined that the third convergence condition is satisfied.
[0098] As an alternative embodiment, the obtained target neural network model after training may only meet the first convergence condition, or meet the first convergence condition and the second convergence condition simultaneously, or meet the first convergence condition and the third convergence condition simultaneously, or meet the first convergence condition, the second convergence condition, and the third convergence condition simultaneously. The specific convergence conditions to be met can be determined according to the actual situation.
[0099] As an alternative embodiment, the acquisition method of the above target image can include various ways. It can be an image obtained by shooting vehicles and pedestrians driving on the road with a camera installed on the roadside, or an image taken by a user with a mobile terminal, or an image in an application installed on the mobile terminal.
[0100] The following uses specific embodiments to illustrate the scenarios to which the present application can be applied.
[0101] Taking the recognition of vehicles and pedestrians driving on the road as an example, images on the road can be collected by a camera installed on the roadside, such as Figure 6 is an alternative schematic diagram of image acquisition according to an embodiment of the present invention. The figure includes a pedestrian 602, a vehicle 604, and a camera 606. Images of pedestrians and vehicles on the current road are collected by the camera in the figure. The collected images are used as a training image set.
[0102] The pedestrians in the image are labeled with person labels, and the vehicles in the image are labeled with vehicle labels. The image may also include other unlabeled objects, such as bicycles, electric vehicles, and other unlabeled objects. The training image set includes sample image set 1, sample image set 2, and sample image set 3. Sample image set 1 includes pedestrians labeled with person labels, sample image set 2 includes vehicles labeled with vehicle labels, and sample image set 3 includes pedestrians labeled with person labels and vehicles labeled with vehicle labels.
[0103] The initial classification neural network CNN0 is trained using sample image set 1. During the training process, a pedestrian attention map M0 can be obtained. Figure 7 is an alternative schematic diagram of training the initial classification neural network CNN0 according to an embodiment of the present invention. The initial classification neural network CNN1 is trained using sample image set 2. During the training process, a vehicle attention map M1 can be obtained, such as Figure 8 is an alternative schematic diagram of training the initial classification neural network CNN1 according to an embodiment of the present invention. The initial recognition neural network CNN2 is trained using sample image set 1, sample image set 2, and sample image set 3. During the training process, a pedestrian attention map M0' and a vehicle attention map M1' can be obtained. Such as Figure 9 is an alternative schematic diagram of training the initial recognition neural network CNN2 according to an embodiment of the present invention.
[0104] During the training of the initial recognition neural network, if the loss functions between the pedestrian attention map M0 and the pedestrian attention map M0', and between the vehicle attention map M1 and the vehicle attention map M1' both satisfy the first convergence condition, stop training the initial recognition neural network to obtain the target recognition neural network. The trained target neural network can identify the vehicles and pedestrians driving on the road and can identify the positions of the pedestrians and vehicles on the road respectively.
[0105] Through this embodiment, the pedestrians and vehicles on the road can be identified to achieve the purpose of monitoring the traffic state.
[0106] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0107] According to another aspect of the embodiments of the present invention, there is also provided an object recognition device for implementing the above object recognition method. As Figure 10 shown, the device includes: an acquisition module 1002 for acquiring a target image, where the target image includes at least one target object to be recognized; an extraction module 1004 for extracting the attention map corresponding to each target object in the target image in the target recognition neural network, where the target recognition neural network is obtained until the first convergence condition is reached when training based on the initial classification neural network and the initial recognition neural network, and the first convergence condition indicates that the output value of the loss function between the first output result of the classification neural network in training and the second output result of the recognition neural network is within a first preset range, the classification neural network is used to recognize images containing single-label objects, and the recognition neural network is used to recognize images containing multi-label objects; a recognition module 1006 for recognizing the classification label corresponding to each target object and the position of each target object in the target image in the target recognition neural network based on the attention map.
[0108] Optionally, the above device is further configured to obtain a plurality of sample images before obtaining the target image, where the plurality of sample images include: a first sample image set and a second sample image set, the first sample image set includes: sample images containing single-label objects and sample images without labels, and the second sample image set includes: sample images containing single-label objects, sample images containing multi-label objects, and sample images without labels; training the initial classification neural network with the first sample image set to obtain a first attention map corresponding to the first sample image set, and the first output result includes the first attention map; training the initial recognition neural network with the second sample image set to obtain a second attention map corresponding to the second sample image set, and the second output result includes the second attention map; determining that the first convergence condition is met when the output value of the loss function between the attention map corresponding to the single-label object in the first attention map and the attention map corresponding to the single-label object in the second attention map is within the first preset range.
[0109] Optionally, the above device is further configured to implement training the initial classification neural network with the first sample image set to obtain a first attention map corresponding to the first sample image set in the following manner: obtaining N groups of first sample image subsets in the first sample image set, where each image in each group of the first sample image subsets includes a single-label object and a plurality of the sample images without labels, and N is an integer greater than or equal to 1; training N classification sub-networks in the initial classification neural network respectively with the N groups of first sample image subsets to obtain N groups of object attention maps, where the first attention map includes the N groups of object attention maps, and each group of the object attention maps includes the object attention map corresponding to the single-label object.
[0110] Optionally, the above device is further configured to implement training N classification sub-networks in the initial classification neural network respectively with the N groups of first sample image subsets to obtain N groups of object attention maps in the following manner: obtaining the i-th group of the first sample image subsets, where i is an integer greater than or equal to 1 and less than or equal to N; inputting the i-th group of the first sample image subsets into the i-th classification sub-network to obtain the i-th group of the object attention maps; and saving the i-th group of the object attention maps.
[0111] Optionally, the above device is further configured to implement training the initial recognition neural network with the second sample image set to obtain the second attention map corresponding to the second sample image set in the following manner: training the initial recognition neural network with the second sample image set for j rounds to obtain the attention map of the j-th round, where j is greater than or equal to 1; when determining that the output value of the loss function between the attention maps corresponding to the M single-label objects in the attention map of the j-th round and the attention maps corresponding to the M single-label objects included in the N groups of object attention maps is within the first preset range, it is determined that the first convergence condition is reached, where M is greater than or equal to 1 and less than N, and the target object is included in the M single-label objects.
[0112] Optionally, the above device is further configured to implement training the initial classification neural network with the first sample image set to obtain the first attention map corresponding to the first sample image set in the following manner: training the initial recognition neural network with the N groups of first sample image subsets to obtain N groups of initial object attention maps; performing a first process on the N groups of initial object attention maps to obtain a first processing result; performing binarization processing and Gaussian filtering processing on the first processing result to obtain the first attention map.
[0113] Optionally, the above device is further configured to implement performing a first process on the N groups of initial object attention maps to obtain a first processing result in the following manner: calculating the product of the mean of the N groups of initial object attention maps and a preset value to obtain a first adjustment value; determining the difference between the N groups of initial object attention maps and the first adjustment value to obtain the first processing result.
[0114] Optionally, the above device is further configured to implement training the initial recognition neural network with the second sample image set to obtain the second attention map corresponding to the second sample image set in the following manner: training the initial recognition neural network with the second sample image set to obtain the initial attention map corresponding to the second sample image set; calculating the difference between the initial attention map corresponding to the second sample image set and the mean of the initial attention map corresponding to the second sample image set to obtain a second processing result; performing normalization processing on the second processing result to obtain the second attention map.
[0115] Optionally, the above-mentioned device is further used when the estimated classification label of the labeled object output by the above-mentioned target recognition neural network and the known classification label of the above-mentioned labeled object satisfy a second convergence condition, where the second convergence condition is used to indicate that the output value of the loss function between the estimated classification label and the known classification label is within a second preset range; the estimated position of the labeled object output by the above-mentioned target recognition neural network and the known position of the above-mentioned labeled object in the sample image satisfy a third convergence condition, where the third convergence condition is used to indicate that the output value of the loss function between the estimated position and the known position is within a third preset range; where the above-mentioned labeled object includes the above-mentioned single-label object and the above-mentioned multi-label object, and the above-mentioned sample image includes the sample image of the above-mentioned single-label object and the sample image of the above-mentioned multi-label object.
[0116] According to another aspect of the embodiments of the present invention, there is also provided an electronic device for implementing the above-mentioned object recognition method, and this electronic device can be Figure 1 the terminal device or server shown in the figure. This embodiment will be described by taking this electronic device as an example. As Figure 11 shown in the figure, the electronic device includes a memory 1102 and a processor 1104. A computer program is stored in the memory 1102, and the processor 1104 is configured to execute the steps in any one of the above method embodiments through the computer program.
[0117] Optionally, in this embodiment, the above-mentioned electronic device may be at least one of multiple network devices in a computer network.
[0118] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through the computer program:
[0119] S1. Obtain a target image, where at least one target object to be recognized is included in the above-mentioned target image;
[0120] S2. Extract the respective attention maps corresponding to each of the above-mentioned target objects in the above-mentioned target image in a target recognition neural network, where the target recognition neural network is obtained by training based on an initial classification neural network and an initial recognition neural network until a first convergence condition is reached. The first convergence condition indicates that the output value of the loss function between the first output result of the classification neural network in training and the second output result of the recognition neural network is within a first preset range. The classification neural network is used to recognize images containing single-label objects, and the recognition neural network is used to recognize images containing multi-label objects;
[0121] S3. In the target recognition neural network, identify the classification labels corresponding to each of the above-mentioned target objects and the positions of each of the above-mentioned target objects in the above-mentioned target image based on the above-mentioned attention maps.
[0122] Optionally, those of ordinary skill in the art can understand that Figure 11 the structure shown is only illustrative, and the electronic device or electronic equipment can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 11 It does not limit the structure of the above-mentioned electronic device or electronic equipment. For example, the electronic device or electronic equipment may further include more or fewer components (such as a network interface, etc.) than those shown Figure 11 in the figure, or have a different configuration from that shown Figure 11 in the figure.
[0123] Among them, the memory 1102 can be used to store software programs and modules, such as the program instructions / modules corresponding to the object recognition method and device in the embodiments of the present invention. The processor 1104 executes various functional applications and data processing by running the software programs and modules stored in the memory 1102, that is, implements the above-mentioned object recognition method. The memory 1102 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1102 may further include a memory remotely disposed relative to the processor 1104, and these remote memories can be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 1102 can specifically but not limitedly be used to store information such as image data, a neural network model structure, and model parameters. As an example, as Figure 11 shown, the above-mentioned memory 1102 may include but is not limited to the acquisition module 1002, the extraction module 1004, and the recognition module 1006 in the above-mentioned object recognition device. In addition, it may also include but is not limited to other module units in the above-mentioned object recognition device, which will not be elaborated in this example.
[0124] Optionally, the above-mentioned transmission device 1106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one instance, the transmission device 1106 includes a Network Interface Controller (NIC), which can be connected to other network devices and routers through a network cable so as to communicate with the Internet or a local area network. In one instance, the transmission device 1106 is a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0125] In addition, the above-mentioned electronic device further includes: a display 1108 for displaying the order information to be processed; and a connection bus 1110 for connecting each module component in the above-mentioned electronic device.
[0126] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes in a network communication manner. Among them, the nodes may form a peer-to-peer (P2P) network, and any form of computing device, such as an electronic device like a server or a terminal, can become a node in the blockchain system by joining the peer-to-peer network.
[0127] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various optional implementation manners. Among them, the computer program is set to execute the steps in any one of the above method embodiments when running.
[0128] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store a computer program for executing the following steps:
[0129] S1. Obtain a target image, where at least one target object to be recognized is included in the above-mentioned target image;
[0130] S2. Extract the attention map corresponding to each of the above-mentioned target objects in the above-mentioned target image in a target recognition neural network. Among them, the target recognition neural network is obtained until a first convergence condition is reached when training is based on an initial classification neural network and an initial recognition neural network. The first convergence condition indicates that the output value of the loss function between the first output result of the classification neural network in the training and the second output result of the recognition neural network is within a first preset range. The classification neural network is used to recognize an image containing a single-label object, and the recognition neural network is used to recognize an image containing a multi-label object;
[0131] S3. In the target recognition neural network, based on the above-mentioned attention map, recognize the classification label corresponding to each of the above-mentioned target objects and the position of each of the above-mentioned target objects in the above-mentioned target image.
[0132] Optionally, in this embodiment, those of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. This program can be stored in a computer-readable storage medium, which can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0133] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0134] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0135] In the above embodiments of the present invention, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0136] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0137] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, in each embodiment of the present invention, each functional unit may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0139] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for identifying an object, characterized in that, Including: Obtain a target image, where at least one target object to be recognized is included in the target image; Extract, in a target recognition neural network, an attention map corresponding to each of the target objects in the target image, where the target recognition neural network is obtained by training an initial classification neural network and an initial recognition neural network based on a plurality of sample images until a first convergence condition is reached. The first convergence condition indicates that the output value of a loss function between a first attention map in a first output result and a second attention map in a second output result is within a first preset range. The first attention map is obtained by training the initial classification neural network using the plurality of sample images, the second attention map is obtained by training the initial recognition neural network using the plurality of sample images, and the single-label object corresponding to the first attention map is the same as the single-label object corresponding to the second attention map. The classification neural network is used to recognize an image containing a single-label object, and the recognition neural network is used to recognize an image containing a multi-label object; In the target recognition neural network, based on the attention map, recognize the classification label corresponding to each of the target objects and the position of each of the target objects in the target image.
2. The method according to claim 1, characterized in that, Before the obtaining the target image, the method further includes: Obtain the plurality of sample images, where the plurality of sample images include a first sample image set and a second sample image set. The first sample image set includes sample images containing single-label objects and sample images without labeled objects. The second sample image set includes sample images containing single-label objects, sample images containing multi-label objects, and sample images without labeled objects; Use the first sample image set to train the initial classification neural network to obtain the first attention map corresponding to the first sample image set; Use the second sample image set to train the initial recognition neural network to obtain the second attention map corresponding to the second sample image set; When it is determined that the output value of the loss function between the attention map corresponding to the single-label object in the first attention map and the attention map corresponding to the single-label object in the second attention map is within the first preset range, it is determined that the first convergence condition is reached.
3. The method according to claim 2, wherein The using the first sample image set to train the initial classification neural network to obtain the first attention map corresponding to the first sample image set includes: Obtain N groups of first sample image subsets in the first sample image set, where each image in each group of the first sample image subsets includes a single-label object and a plurality of the unlabeled objects, and N is an integer greater than or equal to 1; Use the N groups of first sample image subsets to train N classification sub-networks in the initial classification neural network respectively to obtain N groups of object attention maps, where the first attention map includes the N groups of object attention maps, and each group of the object attention maps includes the object attention map corresponding to the single-label object.
4. The method according to claim 3, wherein Training the N classification sub-networks in the initial classification neural network respectively with the N groups of first sample image subsets to obtain N groups of object attention maps, including: Obtaining the i-th group of the first sample image subsets, where i is an integer greater than or equal to 1 and less than or equal to N; Inputting the i-th group of the first sample image subsets into the i-th classification sub-network to obtain the i-th group of object attention maps; Saving the i-th group of object attention maps.
5. The method according to claim 4, wherein Training the initial recognition neural network with the second sample image set to obtain the second attention map corresponding to the second sample image set, including: Training the initial recognition neural network with the second sample image set for j rounds to obtain the attention map of the j-th round, where j is greater than or equal to 1; When it is determined that the output value of the loss function between the attention maps corresponding to the M single-label objects in the attention map of the j-th round and the attention maps corresponding to the M single-label objects included in the N groups of object attention maps is within the first preset range, it is determined that the first convergence condition is reached, where M is greater than or equal to 1 and less than N, and the M single-label objects include the target object.
6. The method according to claim 3, wherein The training of the initial classification neural network with the first sample image set to obtain the first attention map corresponding to the first sample image set includes: Training the initial recognition neural network with the N groups of first sample image subsets to obtain N groups of initial object attention maps; Performing a first processing on the N groups of initial object attention maps to obtain a first processing result; Performing binarization processing and Gaussian filtering processing on the first processing result to obtain the first attention map.
7. The method according to claim 6, wherein The performing a first processing on the N groups of initial object attention maps to obtain a first processing result includes: Calculating the product of the mean of the N groups of initial object attention maps and a preset value to obtain a first adjustment value; Determining the difference between the N groups of initial object attention maps and the first adjustment value to obtain the first processing result.
8. According to the method described in claim 2, the training of the initial recognition neural network with the second sample image set to obtain the second attention map corresponding to the second sample image set includes: Training the initial recognition neural network with the second sample image set to obtain the initial attention map corresponding to the second sample image set; Calculating the difference between the initial attention map corresponding to the second sample image set and the mean of the initial attention map corresponding to the second sample image set to obtain a second processing result; Performing normalization processing on the second processing result to obtain the second attention map.
9. The method according to claim 2, wherein The method further includes at least one of the following: The estimated classification label of the labeled object output by the target recognition neural network satisfies the second convergence condition with the known classification label of the labeled object, where the second convergence condition is used to indicate that the output value of the loss function between the estimated classification label and the known classification label is within the second preset range; The estimated position of the labeled object output by the target recognition neural network in the sample image satisfies a third convergence condition with the known position of the labeled object in the sample image, where the third convergence condition is used to indicate that the output value of the loss function between the estimated position and the known position is within a third preset range; wherein, the labeled object includes the single-label object and the multi-label object, and the sample image includes the sample image of the single-label object and the sample image of the multi-label object.
10. An identification device for an object, characterized in that, Comprising: An acquisition module, configured to acquire a target image, wherein the target image includes at least one target object to be recognized; An extraction module, configured to extract an attention map corresponding to each target object in the target image in a target recognition neural network, wherein the target recognition neural network is obtained by training an initial classification neural network and an initial recognition neural network based on a plurality of sample images until a first convergence condition is reached. The first convergence condition indicates that the output value of the loss function between the first attention map in the first output result and the second attention map in the second output result is within a first preset range. The first attention map is obtained by training the initial classification neural network with the plurality of sample images, and the second attention map is obtained by training the initial recognition neural network with the plurality of sample images. And the single-label object corresponding to the first attention map is the same as the single-label object corresponding to the second attention map. The classification neural network is used to recognize images containing single-label objects, and the recognition neural network is used to recognize images containing multi-label objects; A recognition module, configured to recognize the classification label corresponding to each target object and the position of each target object in the target image in the target recognition neural network based on the attention map.
11. A computer-readable storage medium, the computer-readable storage medium including a stored program, wherein, When the program runs, it executes the method described in any one of claims 1 to 9.
12. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 9 through the computer program.
Citation Information
Patent Citations
Multi-label image classification method, apparatus and device, and storage medium
CN109165666A
A multi-label classification method and system
CN109376757A