Training method of multi-task recognition model, and multi-task recognition method and device
By calculating the second gradient related to the loss value in the multi-task recognition model to adjust the update direction of network parameters, the overfitting and underfitting problems caused by uneven loss value are solved, and the generalization of the model is improved.
Patent Information
- Application Number
- CN202510223094.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
When training neural network models, uneven loss values of each task will cause the model to over-learn the features of high-loss value tasks, leading to over-fitting problems, and at the same time, it is impossible to fully learn task features with small loss values, resulting in under-fitting problems, which will affect the generalization of the model.
By calculating the second gradient of each network parameter for each preset attribute, which is positively correlated with the loss value, limiting the size of the target gradient, reducing the impact of high-loss-value tasks on model learning, and improving the learning possibility of tasks with lower loss value.
This method effectively reduces the impact of high-loss value tasks on model learning, avoids overfitting and underfitting problems, and improves the generalization of the model.
Smart Images

Figure CN120147778A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a training method for a multi-task recognition model, a multi-task recognition method, and an apparatus. Background Art
[0002] A neural network model can process multiple different tasks simultaneously. Specifically, when an image containing a vehicle is input into the neural network model, the neural network model can identify different attribute categories of the vehicle. The identification of each attribute category is a task, where the different attribute categories can be vehicle color, vehicle model, and license plate color.
[0003] During the process of training the neural network model, a sample image can be input into the neural network model to obtain the prediction results of each task; then, according to the difference between the prediction result of each task and the true label, the loss value corresponding to this task is calculated; according to the sum of the loss values corresponding to each task, the network parameters of the model are adjusted.
[0004] However, when the loss value of a certain task is much larger than that of other tasks, during the process of adjusting the network parameters, it will be more inclined to reduce the error of this task with a high loss value, resulting in the neural network model overlearning the features of this task with a high loss value, and thus there will be an overfitting problem for the task with a high loss value; correspondingly, for other tasks with smaller loss values, the neural network model will not be able to fully learn the features of these tasks, and thus there will be an underfitting problem for the tasks with smaller loss values. That is to say, the problem of uneven loss values of each task will seriously affect the comprehensive performance of the neural network model for each task, and the adaptability of the trained model to each task with uneven loss values is poor, and the generalization of the model is not high. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a training method for a multi-task recognition model, a multi-task recognition method, and an apparatus to improve the generalization of the model. The specific technical solutions are as follows:
[0006] In a first aspect, the embodiments of the present invention provide a training method for a multi-task recognition model, and the method includes:
[0007] Obtain a sample image containing a sample object, and obtain the true attribute information of the sample object for each preset attribute;
[0008] Process the sample image by using the current multi-task recognition model to obtain the predicted attribute information of the sample object for each preset attribute;
[0009] Based on the difference between the predicted attribute information and the true attribute information of each preset attribute for the sample object, obtain the loss value corresponding to the preset attribute;
[0010] For each network parameter in the current multi-task recognition model, calculate the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameter, and obtain the first gradient of the network parameter for the preset attribute;
[0011] Calculate the second gradient of the network parameter for the preset attribute; wherein, the second gradient of the network parameter for the preset attribute is positively correlated with the loss value of the sample object for the preset attribute;
[0012] Calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute, and obtain the target gradient of the network parameter for the preset attribute;
[0013] Based on the sum of the target gradients of the network parameter for each preset attribute, adjust the network parameter to obtain a trained multi-task recognition model.
[0014] Optionally, calculating the second gradient of the network parameter for the preset attribute includes:
[0015] Perform norm constraint on the first gradient of the network parameter for the preset attribute to obtain the second gradient of the network parameter for the preset attribute.
[0016] Optionally, the obtaining the sample image including the sample object includes:
[0017] Obtain a plurality of initial images including the sample object;
[0018] Perform random image processing and / or image fusion processing on the plurality of initial images to obtain a sample image; wherein, the random image processing includes: random horizontal flipping, random image rotation, and random image blurring.
[0019] Optionally, before obtaining the true attribute information of the sample object for each preset attribute, the method further includes:
[0020] Obtain the true attribute information of the sample object for each preset attribute in each initial image;
[0021] The obtaining the true attribute information of the sample object for each preset attribute includes:
[0022] In the case where each sample image is obtained by processing the corresponding initial image based on the random image processing, use the true attribute information of the sample object for each preset attribute in the initial image corresponding to the sample image as the true attribute information of the sample object for each preset attribute;
[0023] When the sample image is obtained by processing the corresponding initial image based on the image fusion processing, the true attribute information of the sample object in the initial image corresponding to the sample image for each preset attribute is fused according to a preset weight to obtain the true attribute information of the sample object for each preset attribute.
[0024] Optionally, the sample image includes images of multiple batches;
[0025] Processing the sample image by using the current multi-task recognition model to obtain the predicted attribute information of the sample object for each preset attribute, including:
[0026] Obtain the sample image of the first batch as the current sample image to be processed;
[0027] Input the current sample image to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample image to be processed for each preset attribute;
[0028] Based on the difference between the predicted attribute information and the true attribute information of the sample object for each preset attribute, obtain the loss value corresponding to the preset attribute, including:
[0029] For each preset attribute, based on the difference between the predicted attribute information and the true attribute information of the current sample image to be processed for the preset attribute, obtain the loss value corresponding to the preset attribute;
[0030] For each network parameter in the current multi-task recognition model, calculate the change rate of the loss value of the sample object for each preset attribute in the direction of the network parameter to obtain the first gradient of the network parameter for the preset attribute, including:
[0031] For each network parameter in the current multi-task recognition model, calculate the current change rate of the loss value of the sample object for each preset attribute in the direction of the network parameter to obtain the current first gradient of the network parameter for the preset attribute;
[0032] Calculate the second gradient of the network parameter for the preset attribute, including:
[0033] Calculate the current second gradient of the network parameter for the preset attribute;
[0034] Calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute to obtain the target gradient of the network parameter for the preset attribute, including:
[0035] Calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute currently, to obtain the target gradient of the network parameter for the preset attribute currently;
[0036] Based on the sum of the target gradients of the network parameter for each preset attribute, adjust the network parameter to obtain a trained multi-task recognition model, including:
[0037] Obtain the sum of the target gradients of the network parameter for each preset attribute currently, adjust the network parameter to obtain the adjustment result of this batch, so as to update the current multi-task recognition model;
[0038] Obtain the sample images of the next batch as the current sample images to be processed, and return to execute the step of inputting the current sample images to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample images to be processed for each preset attribute;
[0039] Calculate the average level of the adjustment results of multiple batches to obtain the network parameter of the trained multi-task recognition model.
[0040] Optionally, the calculating the average level of the adjustment results of multiple batches to obtain the network parameter of the trained multi-task recognition model includes:
[0041] Calculate the weighted sum of the adjustment results of each batch to obtain the network parameter of the trained multi-task recognition model; wherein, the weight of the adjustment result of each batch is negatively correlated with the generated duration of the adjustment result of this batch; the generated duration of the adjustment result of this batch represents: the interval duration between the generation moment of the adjustment result of this batch and the current moment.
[0042] Optionally, the multi-task recognition model includes: an input network, a feature extraction network, and an output network;
[0043] Inputting the current sample images to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample images to be processed for each preset attribute includes:
[0044] Input the sample image into the input network to obtain a first feature vector;
[0045] Input the first feature vector into the feature extraction network to obtain a second feature vector;
[0046] Based on the processing of the second feature vector by the output network, obtain the predicted attribute information of the current sample images to be processed for each preset attribute.
[0047] Optionally, the feature extraction network includes a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fourth feature extraction sub-network;
[0048] Inputting the first feature vector into the feature extraction network to obtain a second feature vector includes:
[0049] Inputting the first feature vector into the first feature extraction sub-network to obtain a first sub-feature vector;
[0050] Inputting the first sub-feature vector and the first feature vector into the second feature extraction sub-network to obtain a second sub-feature vector;
[0051] Inputting the second sub-feature vector and the first sub-feature vector into the third feature extraction sub-network to obtain a third sub-feature vector;
[0052] Inputting the third sub-feature vector and the second sub-feature vector into the fourth feature extraction sub-network to obtain a second feature vector;
[0053] Processing the second feature vector based on the output network to obtain prediction attribute information of the current to-be-processed sample image for each preset attribute includes:
[0054] Inputting the second feature vector and the third sub-feature vector into the output network to obtain prediction attribute information of the current to-be-processed sample image for each preset attribute.
[0055] Optionally, the sample object is a vehicle, and each preset attribute includes any combination of the following: vehicle color, vehicle type, and license plate color.
[0056] In a second aspect, an embodiment of the present invention provides a multi-task recognition method, and the method includes:
[0057] Obtaining a to-be-recognized image including a to-be-recognized object;
[0058] Inputting the to-be-recognized image into a pre-trained multi-task recognition model to obtain attribute information of the to-be-recognized object for each preset attribute; wherein, the multi-task recognition model is trained based on the training method of the above multi-task recognition model.
[0059] Optionally, the to-be-recognized object is a vehicle, and each preset attribute includes any combination of the following: vehicle color, vehicle type, and license plate color.
[0060] In a third aspect, an embodiment of the present invention provides a training device for a multi-task recognition model, and the device includes:
[0061] A first acquisition module, configured to acquire a sample image including a sample object, and acquire true attribute information of the sample object for each preset attribute;
[0062] A first processing module, configured to process the sample image by using a current multi-task recognition model to obtain predicted attribute information of the sample object for each preset attribute;
[0063] A second acquisition module, configured to obtain a loss value corresponding to the preset attribute based on a difference between the predicted attribute information and the true attribute information of the sample object for each preset attribute;
[0064] A first calculation module, configured to calculate, for each network parameter in the current multi-task recognition model, a change rate of the loss value of the sample object for each preset attribute in the direction of the network parameter, to obtain a first gradient of the network parameter for the preset attribute;
[0065] A second calculation module, configured to calculate a second gradient of the network parameter for the preset attribute; wherein, the second gradient of the network parameter for the preset attribute is positively correlated with the loss value of the sample object for the preset attribute;
[0066] A third calculation module, configured to calculate a difference between the first gradient and the second gradient of the network parameter for the preset attribute, to obtain a target gradient of the network parameter for the preset attribute;
[0067] An adjustment module, configured to adjust the network parameter based on a sum of the target gradients of the network parameter for each preset attribute, to obtain a trained multi-task recognition model.
[0068] In a fourth aspect, an embodiment of the present invention provides a multi-task recognition device, where the device includes:
[0069] A third acquisition module, configured to acquire a to-be-recognized image including a to-be-recognized object;
[0070] An input module, configured to input the to-be-recognized image into a pre-trained multi-task recognition model to obtain attribute information of the to-be-recognized object for each preset attribute; wherein, the multi-task recognition model is trained by using the training method of the multi-task recognition model as described above.
[0071] An embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0072] The memory is configured to store a computer program;
[0073] A processor, when executing a program stored in a memory, implements the training method and the multi-task recognition method of the above multi-task recognition model.
[0074] An embodiment of the present invention further provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the training method and the multi-task recognition method of the above multi-task recognition model are implemented.
[0075] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the training method and the multi-task recognition method of the above multi-task recognition model are implemented.
[0076] Beneficial effects of the embodiments of the present invention:
[0077] In an embodiment of the present invention, for each network parameter in the current multi-task recognition model, since the second gradient of the network parameter for each preset attribute is positively correlated with the loss value of the sample object for the preset attribute, the greater the loss value, the greater the second gradient; the target gradient of the network parameter for the preset attribute is the difference between the first gradient and the second gradient of the network parameter for the preset attribute. The target gradient is constrained by the second gradient, and the second gradient is positively correlated with the loss value. That is to say, if the loss value of the sample object for a preset attribute is greater, the constraint on the target gradient of a network parameter for the preset attribute is greater. Through the second gradient, the target gradient of the network parameter for the preset attribute can be restricted from being too large. That is, through the second gradient, the target gradient of the preset attribute with a high loss value can be restricted from being too large, reducing the proportion of the target gradient of the preset attribute with a high loss value in the sum of the target gradients of each preset attribute, and increasing the proportion of the target gradients of other preset attributes with lower loss values in the sum of the target gradients of each preset attribute. Therefore, the influence of the preset attribute with a high loss value on the sum of the target gradients of each preset attribute can be reduced, that is, the influence of the task with a high loss value on the model learning is reduced, avoiding the model from overlearning the features of the task with a high loss value, increasing the possibility of the model learning the features of other tasks, reducing the possibility of overfitting problems for the task with a high loss value and underfitting problems for the task with a small loss value, and improving the generalization of the model.
[0078] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. Description of the Drawings
[0079] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0080] Figure 1 Schematic flow chart of the first training method for a multi-task recognition model provided by an embodiment of the present invention;
[0081] Figure 2 Schematic flow chart of the second training method for a multi-task recognition model provided by an embodiment of the present invention;
[0082] Figure 3 Schematic flow chart of the processing steps based on the network structure of the multi-task recognition model in a training method for a multi-task recognition model provided by an embodiment of the present invention;
[0083] Figure 4 Schematic flow chart of a multi-task recognition method provided by an embodiment of the present invention;
[0084] Figure 5 Schematic structural diagram of a multi-task recognition model provided by an embodiment of the present invention;
[0085] Figure 6 Schematic structural diagram of a training device for a multi-task recognition model provided by an embodiment of the present invention;
[0086] Figure 7 Schematic structural diagram of a multi-task recognition device provided by an embodiment of the present invention;
[0087] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0088] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on the present invention belong to the scope of protection of the present invention.
[0089] To better understand the embodiments of the present invention, the prior art will be introduced in detail below.
[0090] Object structuring is a relatively common technical direction in product applications. Its technical principle is to use technical means such as object detection and object classification to obtain various attribute information of items as structured information to complete structured tasks. This process can help users accurately obtain the knowledge they are concerned about, thereby efficiently completing different downstream tasks. Among them, the downstream tasks can be further data analysis of the structured information. For example, the downstream tasks can be data statistics and trend prediction, etc.
[0091] Taking common traffic vehicles as an example, processing their structured tasks can effectively meet the needs for vehicle attribute information in various scenarios. In parking lot management, vehicle structuring technology can quickly identify information such as vehicle models and license plate colors, realizing automated vehicle access management and billing; in the scenario of road violation capture, this technology helps to accurately identify the attributes of illegal vehicles, providing strong data support for traffic law enforcement and playing an important role in the field of intelligent transportation.
[0092] For such structured tasks, most of the existing solutions obtain vehicle information in a traditional way. For example, the attribute information of vehicles is obtained by comparing vehicle images with image databases or vehicle images with license plate databases.
[0093] However, in order to achieve the comparison of the database, the image acquisition device needs to interact with the server where the database is located, which takes a lot of time and cannot respond in time to the requirements of scenarios such as road violation capture and fast access management in intelligent parking lots. The real-time performance is poor, seriously affecting the actual application effect.
[0094] In the prior art, a neural network model that has been trained processes multiple different tasks at the same time. However, during the process of training the neural network model, when the loss value of a certain task is much larger than that of other tasks, in the process of adjusting the network parameters, it will be more inclined to reduce the error of the task with the high loss value, resulting in the neural network model overlearning the characteristics of the task with the high loss value, and thus there will be an overfitting problem for the task with the high loss value; correspondingly, for other tasks with smaller loss values, the neural network model will not be able to fully learn the characteristics of these tasks, and thus there will be an underfitting problem for the tasks with smaller loss values. That is to say, the problem of uneven loss values of each task will seriously affect the comprehensive performance of the neural network model for each task, and the generalization of the neural network model is not high.
[0095] In order to improve the generalization of the trained model, the embodiments of the present invention provide a training method for a multi-task recognition model, a multi-task recognition method and a device.
[0096] First, a training method for a multi-task recognition model provided by the embodiments of the present invention will be introduced below.
[0097] Among them, a training method for a multi-task recognition model provided by an embodiment of the present invention can be applied to an electronic device. Exemplarily, the electronic device can be a remote server or a local computer.
[0098] A training method for a multi-task recognition model provided by an embodiment of the present invention may include the following steps:
[0099] Obtain a sample image containing a sample object, and obtain the true attribute information of the sample object for each preset attribute;
[0100] Use the current multi-task recognition model to process the sample image to obtain the predicted attribute information of the sample object for each preset attribute;
[0101] Based on the difference between the predicted attribute information and the true attribute information of the sample object for each preset attribute, obtain the loss value corresponding to the preset attribute;
[0102] For each network parameter in the current multi-task recognition model, calculate the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameter to obtain the first gradient of the network parameter for the preset attribute;
[0103] Calculate the second gradient of the network parameter for the preset attribute; wherein, the second gradient of the network parameter for the preset attribute is positively correlated with the loss value of the sample object for the preset attribute;
[0104] Calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute to obtain the target gradient of the network parameter for the preset attribute;
[0105] Based on the sum of the target gradients of the network parameter for each preset attribute, adjust the network parameter to obtain a trained multi-task recognition model.
[0106] In an embodiment of the present invention, for each network parameter in the current multi-task recognition model, since the second gradient of the network parameter with respect to each preset attribute is positively correlated with the loss value of the sample object with respect to the preset attribute, the greater the loss value, the greater the second gradient; the target gradient of the network parameter with respect to the preset attribute is the difference between the first gradient and the second gradient of the network parameter with respect to the preset attribute, and the target gradient is constrained by the second gradient. Since the second gradient is positively correlated with the loss value, that is, if the loss value of the sample object with respect to a preset attribute is greater, the constraint on the target gradient of a network parameter with respect to the preset attribute is greater. The second gradient can limit the target gradient of the network parameter with respect to the preset attribute from being too large, that is, the second gradient can limit the target gradient of the preset attribute with a high loss value from being too large, reduce the proportion of the target gradient of the preset attribute with a high loss value in the sum of the target gradients of all preset attributes, and increase the proportion of the target gradients of other preset attributes with lower loss values in the sum of the target gradients of all preset attributes. Therefore, the influence of the preset attributes with high loss values on the sum of the target gradients of all preset attributes can be reduced, that is, the influence of the tasks with high loss values on model learning is reduced, the model is prevented from overlearning the features of the tasks with high loss values, the possibility of the model learning the features of other tasks is increased, and the possibilities of overfitting problems occurring for the tasks with high loss values and underfitting problems occurring for the tasks with small loss values are reduced, and the generalization of the trained model can be improved.
[0107] The following combines Figure 1 to introduce a training method for a multi-task recognition model provided by an embodiment of the present invention. As Figure 1 shown, the method may include steps S101 - S107.
[0108] S101, obtain a sample image including a sample object, and obtain the true attribute information of the sample object for each preset attribute.
[0109] It can be understood that in order to train a multi-task recognition model capable of recognizing multiple preset attributes, an electronic device may obtain a sample image including a sample object and the true attribute information of the sample object for each preset attribute.
[0110] The electronic device may identify the sample object in the image to obtain an image including the sample object as the sample image including the sample object. Exemplarily, the electronic device may identify the image through a pre-trained object recognition model to determine whether the image includes a sample object, and then use the image including the sample object as the sample image including the sample object. Among them, the sample object recognition model may be a You Only Look Once (YOLO) model.
[0111] The true attribute information of the sample object in the sample image for each preset attribute can be labeled by the staff according to the true label of the sample object for the preset attribute, or can be the true attribute information of the sample object in the existing dataset. The preset attribute can be the attribute category of the sample object to be recognized that is preset in advance.
[0112] Optionally, in one implementation, the sample object is a vehicle, and each preset attribute includes any combination of the following: vehicle color, vehicle type, and license plate color.
[0113] It can be understood that in order to train a multi-task recognition model capable of recognizing vehicles, the sample object can be a vehicle, and each preset attribute can include any combination of the following: vehicle color, vehicle type, and license plate color. Among them, the vehicle color can be white, black, or red, etc., the vehicle type can be a sedan, a sports car, or a sport utility vehicle (SUV), etc., and the license plate color is green or blue. Exemplarily, sample images containing vehicles and the true attribute information of the vehicles for each preset attribute can be obtained from the Chinese City Parking Dataset (CCPD). Among them, the combinations of each preset attribute can be: vehicle color, vehicle type, and license plate color; the combinations of each preset attribute can also be: license plate color and vehicle type; the combinations of each preset attribute can also be vehicle color and license plate color; the combinations of each preset attribute can also be vehicle color and vehicle type.
[0114] The multi-task recognition model is used to recognize each preset attribute. That is, each task that the multi-task recognition model needs to process is to recognize each preset attribute. For example, if each preset attribute of the sample object includes: vehicle color, vehicle type, and license plate color, then the true attribute information of the sample object for each preset attribute includes: the vehicle color is white, the vehicle type is a sedan, and the license plate color is blue; among them, the recognition of the vehicle color can be regarded as the first task, the recognition of the vehicle type can be regarded as the second task, and the recognition of the license plate color can be regarded as the third task, and the multi-task recognition model can process the above three tasks simultaneously.
[0115] In the embodiments of the present invention, by obtaining sample images containing vehicles and the true attribute information of the vehicles for each preset attribute, a multi-task recognition model capable of recognizing vehicles can be trained, so as to obtain structured information in application scenarios such as parking lot automatic entrance and exit gates or road speeding capture.
[0116] Optionally, in another implementation, the sample object is a pedestrian, and each preset attribute includes: pedestrian posture and pedestrian position. For example, the pedestrian posture can be standing or falling, etc., and the pedestrian position can be on the sidewalk or on the roadway, etc. In the embodiments of the present invention, by obtaining a sample image containing a pedestrian and the true attribute information of the pedestrian for each preset attribute, a multi-task recognition model capable of recognizing pedestrians can be trained, so as to obtain structured information in application scenarios such as traffic roads or parks.
[0117] Optionally, in one implementation, the specific steps of obtaining a sample image containing a sample object may include steps A1-A2.
[0118] A1. Obtain a plurality of initial images containing the sample object.
[0119] A2. Perform random image processing and / or image fusion processing on the plurality of initial images to obtain a sample image. Among them, the random image processing includes: random horizontal flipping, random image rotation, and random image blurring.
[0120] It can be understood that, in order to increase the diversity and randomness of the sample image, the electronic device can obtain a plurality of initial images containing the sample object, and the initial image itself can also be used as a sample image for training. The electronic device can perform random image processing and / or image fusion processing on the initial image to obtain a sample image.
[0121] Among them, the random image processing includes: random horizontal flipping, random image rotation, and random image blurring. Exemplarily, in the random horizontal flipping process, the initial image can be flipped according to a preset axis of symmetry, where the axis of symmetry can be the vertical midline of the image; in the random image rotation process, the initial image can be rotated according to a preset rotation center and rotation angle; in the random image blurring process, the initial image can be blurred by a preset blurring algorithm, where the preset blurring algorithm can be a Gaussian blurring algorithm or a mean blurring algorithm, etc. Among them, the image fusion processing can be to fuse the pixel values at the same pixel position of a plurality of initial images according to a preset weight.
[0122] It should be noted that the random image processing is the processing performed on one initial image, and the image fusion processing is the processing performed on multiple images. The multiple images fused by the image fusion processing can include images that have undergone random image processing and / or images that have not undergone random image processing.
[0123] Optionally, only random image processing can be performed on the initial image, or only image fusion processing can be performed on the initial image, or random image processing can be performed on the initial image first and then image fusion processing, or image fusion processing can be performed on the initial image first and then random image processing.
[0124] In an embodiment of the present invention, an electronic device can perform random image processing and / or image fusion processing on an initial image to increase the diversity of training samples, enabling a multi-task recognition model to learn the features of more new data and improving the generalization ability of the model.
[0125] Optionally, in one implementation, before obtaining the true attribute information of the sample object for each preset attribute, the method may further include step B1, and the specific steps of obtaining the true attribute information of the sample object for each preset attribute may include steps B2 - B3.
[0126] B1, obtain the true attribute information of the sample object for each preset attribute in each initial image.
[0127] It can be understood that, in order to train a multi-task recognition model, an electronic device can obtain the true attribute information of the sample object for each preset attribute in the initial image.
[0128] B2, when each sample image is obtained by performing random image processing on the corresponding initial image, use the true attribute information of the sample object for each preset attribute in the initial image corresponding to the sample image as the true attribute information of the sample object for each preset attribute.
[0129] It can be understood that for an initial image, if random image processing is performed on the initial image to obtain the sample image corresponding to the initial image, the preset attributes of the sample object in the initial image will not change, and the true attribute information of the sample image corresponding to the initial image is the true attribute information of the initial image. Therefore, when each sample image is obtained by performing random image processing on the corresponding initial image, the true attribute information of the sample object for each preset attribute in the initial image corresponding to the sample image can be used as the true attribute information of the sample object for each preset attribute.
[0130] B3, when the sample image is obtained by performing image fusion processing on the corresponding initial image, fuse the true attribute information of the sample object for each preset attribute in the initial image corresponding to the sample image according to a preset weight to obtain the true attribute information of the sample object for each preset attribute.
[0131] It can be understood that image fusion processing can be to perform image fusion on multiple initial images according to a preset weight to obtain the corresponding sample image. When obtaining the true attribute information of the sample image obtained by image fusion processing, the true attribute information of the sample object for the same preset attribute in each initial image fused can be fused according to the preset weight used during image fusion.
[0132] Exemplarily, the image fusion process can be a mixup process. For two initial images among multiple initial images, these two initial images can be fused according to a preset weight, and the true attribute information of the sample object in these two initial images for the same preset attribute can be fused.
[0133] The sample image obtained through the mixup process can be expressed by the following formula: Where, represents the sample image obtained by fusing the i-th initial image and the j-th initial image, x i represents the i-th initial image, x j represents the j-th initial image, m represents the preset weight, represents the true attribute information of the sample object in the fused sample image for each preset attribute, y i represents the true attribute information of the sample object in the i-th initial image for each preset attribute, y j represents the true attribute information of the sample object in the j-th initial image for each preset attribute. Among them, m can conform to the beta distribution, and the beta distribution can be expressed as, m ~ Beta(β, β), β ∈ (0, ∞).
[0134] Optionally, if the sample image is obtained based on random image processing and image fusion processing, since random image processing does not affect the true attribute information of the sample object in the initial image for each preset attribute, the true attribute information of the sample object in the sample image for each preset attribute can be obtained in the way of image fusion processing. Specifically, the true attribute information of the sample object in the initial image corresponding to the sample image for each preset attribute is fused to obtain the true attribute information of the sample object for each preset attribute.
[0135] In the embodiments of the present invention, according to the acquisition method of the sample image, for the sample image obtained by random image processing, the true attribute information of the sample object in the initial image corresponding to the sample image for each preset attribute is used as the true attribute information of the sample object for each preset attribute; for the sample image obtained by image fusion processing, according to the preset weight, the true attribute information of the sample object in the initial image corresponding to the sample image for each preset attribute is fused to obtain the true attribute information of the sample object for each preset attribute, which can ensure accurate true attribute information is obtained.
[0136] S102. Use the current multi-task recognition model to process the sample image to obtain the predicted attribute information of the sample object for each preset attribute.
[0137] It can be understood that the electronic device can process the sample image using the current multi-task recognition model. Specifically, a sample image can be input into the current multi-task recognition model to obtain the predicted attribute information of the sample object for each preset attribute.
[0138] Among them, the multi-task recognition model can be a neural network model with a preset structure. Exemplarily, the multi-task recognition model can be a Residual Network (ResNet) model, such as the resnet14 model.
[0139] It should be noted that each preset attribute recognized by the multi-task recognition model will have multiple classification results. A classification result can be represented by a one-dimensional vector. That is to say, the output result of the multi-task recognition model can be a multi-dimensional vector, and this multi-dimensional vector contains vectors representing the predicted attribute information of each preset attribute. The predicted attribute information of a preset attribute can be represented by a vector with the same number of dimensions as the number of classification results of this preset attribute. For example, if the preset attribute is the license plate color, the classification results of the license plate color include two classification results: blue and green. The predicted attribute information for the license plate color can be represented by a two-dimensional vector. Therefore, the total dimension of the multi-dimensional vector is the sum of the number of classification results of each preset attribute. For example, the multi-task recognition model can recognize two preset attributes, and each preset attribute has 6 classification results. The multi-task recognition model can output a 12-dimensional multi-dimensional vector to represent the predicted attribute information for each preset attribute.
[0140] It should be noted that the true attribute information of the sample object for each preset attribute can be obtained by encoding the true label of this preset attribute according to the OneHot encoding technique, and each classification result of this preset attribute has a one-dimensional vector representation. Exemplarily, if the preset attribute is the vehicle color, the classification results of the vehicle color include red, black, and white. The true attribute information for the vehicle color being red can be represented by the vector (1, 0, 0), the true attribute information for the vehicle color being black can be represented by the vector (0, 1, 0), and the true attribute information for the vehicle color being white can be represented by the vector (0, 0, 1). When calculating the loss value of this preset attribute subsequently, the difference between the vector representing the true attribute information and the vector representing the predicted attribute information can be compared. Specifically, since both the true attribute information and the predicted attribute information of this preset attribute are represented by vectors, the cross-entropy loss value or binary cross-entropy loss value of the vector representing the true attribute information and the vector representing the predicted attribute information can be calculated to obtain an accurate loss value.
[0141] Optionally, in one implementation, the electronic device can pre-train the multi-task recognition model. Specifically, a large-scale data set can be used to initially train the multi-task recognition model so that the multi-task recognition model has the ability to understand and recognize images. Subsequently, a data set related to the preset attributes can be used to further train the pre-trained multi-task recognition model. That is, steps S101 - S107 can be performed subsequently to train the pre-trained multi-task recognition model. Among them, the large-scale data set can be an existing image data set. For example, the ImageNet1K data set containing images of 1000 categories; the data set related to the preset attributes can be a data set containing the true attribute information of the preset attributes. For example, if the preset attribute is the license plate color, the data set related to the preset attributes can be a data set containing the true attribute information of the license plate color.
[0142] In the embodiments of the present invention, the multi-task recognition model can be pre-trained so that the multi-task recognition model has the ability to understand and recognize images. The subsequent training can be to use a data set with a smaller data volume to adjust the pre-trained multi-task recognition model. Compared with the method without pre-training, it can make the multi-task recognition model more accurate in recognizing the preset attributes and improve the efficiency of training the multi-task recognition model to recognize the preset attributes.
[0143] S103, based on the difference between the predicted attribute information and the true attribute information of the sample object for each preset attribute, obtain the loss value corresponding to the preset attribute.
[0144] It can be understood that the electronic device can use a preset loss function to calculate the loss value between the predicted attribute information and the true attribute information of the sample object for each preset attribute. Among them, the preset loss function can be a Binary Cross Entropy Loss (BCE) function or a Cross Entropy Loss (CE) function.
[0145] S104, for each network parameter in the current multi-task recognition model, calculate the change rate of the loss value of the sample object for each preset attribute in the direction of the network parameter, and obtain the first gradient of the network parameter for the preset attribute.
[0146] It can be understood that in order to adjust the network parameters of the multi-task recognition model, the gradient descent method can be used to calculate the rate of change of the loss function in the direction of the current network parameters as the gradient, and then update the parameters along the opposite direction of the gradient until the minimum value of the loss function is found. Among them, the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameters can also be referred to as the partial derivative of the loss value of the sample object for each preset attribute with respect to the network parameters. Exemplarily, the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameters can be calculated according to the following formula: Where G W represents the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameters, L i represents the loss value of the i-th preset attribute, and W represents the network parameters.
[0147] S105, calculate the second gradient of the network parameter for the preset attribute.
[0148] Among them, the second gradient of the network parameter for the preset attribute is positively correlated with the loss value of the sample object for the preset attribute.
[0149] It can be understood that among the loss values of each task, there may be a task whose loss value is much larger than that of other tasks, resulting in low generalization of the trained multi-task recognition model. That is to say, among the loss values of each preset attribute, there may be a preset attribute with a large loss value. To reduce the impact of the preset attribute with a high loss value on the training effect, the electronic device can set a penalty term for the first gradient of each network parameter for the preset attribute.
[0150] Specifically, for each network parameter, the electronic device can calculate the second gradient of the network parameter for the preset attribute based on the loss value of the sample object for the preset attribute, and use the second gradient as the penalty term. The calculated second gradient is positively correlated with the loss value of the sample object for the preset attribute.
[0151] Optionally, in one implementation, the electronic device can use the loss value of the sample object for the preset attribute as the second gradient of the network parameter for the preset attribute. Optionally, in another implementation, the electronic device can multiply the loss value of the sample object for the preset attribute by a preset coefficient to obtain the second gradient of the network parameter for the preset attribute.
[0152] S106, calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute to obtain the target gradient of the network parameter for the preset attribute.
[0153] It can be understood that the second gradient of the network parameter with respect to the preset attribute is positively correlated with the loss value of the sample object with respect to the preset attribute. The larger the loss value, the larger the second gradient, and the smaller the loss value, the smaller the second gradient. Taking the second gradient as the minuend to weaken the first gradient can achieve the purpose of restricting the target gradient, reduce the target gradient of the preset attribute with a high loss value, and do not weaken or weaken less the target gradient with a low loss value, so as to obtain a relatively stable target gradient.
[0154] In the prior art, usually the rate of change of the loss value of the sample object with respect to each preset attribute in the direction of the network parameter is directly used as the gradient for adjusting the network parameter, and the gradient is not weakened. The gradient is positively correlated with the loss value, and the gradient of the task with a high loss value will also be very high. When adjusting the network parameter, it will be more inclined to reduce the error of the task with a high loss value. In the embodiment of the present invention, the first gradient of the preset attribute with a high loss value can be weakened through the second gradient, restricting the magnitude of the target gradient for adjusting the network parameter with respect to the preset attribute, reducing the proportion of the target gradient of the preset attribute with a high loss value in the sum of the target gradients of each preset attribute, increasing the proportion of the target gradients of other preset attributes with lower loss values in the sum of the target gradients of each preset attribute, and reducing the influence of the preset attribute with a high loss value on the sum of the target gradients of each preset attribute.
[0155] Optionally, in one implementation, step S106 includes performing a norm constraint on the first gradient of the network parameter with respect to the preset attribute to obtain the second gradient of the network parameter with respect to the preset attribute.
[0156] It can be understood that the electronic device can use a preset norm regularization method to perform a norm constraint on the first gradient of the network parameter with respect to the preset attribute to obtain the second gradient of the network parameter with respect to the preset attribute. Among them, the preset norm regularization method can be L2 norm regularization or L1 norm regularization. Since a norm constraint is performed on the first gradient, and the first gradient is the partial derivative of the loss value with respect to the network parameter, and the first gradient and the loss value are in a positive correlation relationship, performing a norm constraint on the first gradient does not affect the positive or negative nature of the first gradient. Therefore, the second gradient obtained after the norm constraint is positively correlated with the loss value.
[0157] Among them, the second gradient can be expressed by the following formula: W represents a network parameter, i represents the i-th preset attribute, L i represents the loss value corresponding to the i-th preset attribute, represents the second gradient of the network parameter with respect to the preset attribute; E represents the mathematical expectation of the mean value; represents the first gradient of the network parameter with respect to the preset attribute, It represents the L2-norm regularization of the first gradient of the network parameter with respect to the preset attribute. Since the calculation of the gradient has a linear relationship, when calculating the norm regularization of the gradient, for example, calculating the product of the L1-norm regularization and the L2-norm regularization of the gradient, the complex calculation process can be decomposed into a simple linear combination calculation process. For example, the accumulation of the loss value can be disassembled into the accumulation of the gradient, which is more convenient for calculating the norm regularization. Such norm regularization is also easy to calculate and obtain in practical applications, which can improve the calculation efficiency.
[0158] In the embodiment of the present invention, the second gradient is obtained by performing norm constraint on the first gradient, and the second gradient is used as the penalty term, which can accurately control the size of the penalty term of the gradient and improve the accuracy of the target gradient. And generally speaking, directly using the loss value as the penalty term will be too large and cause the optimization objective to become optimizing the loss plus a penalty term, which may lead to the situation of underfitting, resulting in a loss of accuracy when the model balances the performance of various tasks. However, gradient penalty will not cause such a situation. The loss value can be adjusted using a preset coefficient, and the adjusted loss value is used as the penalty term. However, the setting of the preset coefficient requires multiple attempts. The embodiment of the present invention can avoid the setting of the preset coefficient, reduce the training difficulty, and better balance the parameter update logic when optimizing each task.
[0159] S107. Based on the sum of the target gradients of the network parameter with respect to each preset attribute, adjust the network parameter to obtain a trained multi-task recognition model.
[0160] It can be understood that for each network parameter, the sum of the target gradients of the network parameter with respect to each preset attribute can be used to adjust the network parameter. Exemplarily, the difference between the network parameter and the sum of the target gradients of the network parameter with respect to each preset attribute can be calculated as the adjusted network parameter to obtain a trained multi-task recognition model.
[0161] In the embodiments of the present invention, for each network parameter in the current multi-task recognition model, since the second gradient of the network parameter with respect to each preset attribute is positively correlated with the loss value of the sample object with respect to the preset attribute, the larger the loss value, the larger the second gradient; the target gradient of the network parameter with respect to the preset attribute is the difference between the first gradient and the second gradient of the network parameter with respect to the preset attribute, and the target gradient is constrained by the second gradient. Since the second gradient is positively correlated with the loss value, that is, if the loss value of the sample object with respect to a preset attribute is larger, the constraint on the target gradient of a network parameter with respect to the preset attribute is greater. Through the second gradient, the target gradient of the network parameter with respect to the preset attribute can be restricted from being too large. That is, through the second gradient, the target gradient of the preset attribute with a high loss value can be restricted from being too large, reducing the proportion of the target gradient of the preset attribute with a high loss value in the sum of the target gradients of all preset attributes, and increasing the proportion of the target gradients of other preset attributes with lower loss values in the sum of the target gradients of all preset attributes. Therefore, the influence of the preset attributes with high loss values on the sum of the target gradients of all preset attributes can be reduced, that is, the influence of the tasks with high loss values on model learning is reduced, avoiding the model from overlearning the features of the tasks with high loss values, increasing the possibility of the model learning the features of other tasks, reducing the possibility of overfitting problems for tasks with high loss values and underfitting problems for tasks with small loss values, and improving the generalization of the trained model.
[0162] Optionally, in one embodiment, based on Figure 1 the training method of the multi-task recognition model shown, the sample image includes multiple batches of images. As Figure 2 shown, step S102 includes steps S1021 - S1022, step S103 includes step S1031, step S104 includes step S1041, step S105 includes step S1051, step S106 includes step S1061, and step S107 includes steps S1071 - S1073.
[0163] S1021, Obtain the first batch of sample images as the current sample image to be processed.
[0164] S1022, Input the current sample image to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample image to be processed for each preset attribute.
[0165] S1031, For each preset attribute, based on the difference between the predicted attribute information and the true attribute information of the current sample image to be processed for the preset attribute, obtain the loss value corresponding to the preset attribute.
[0166] S1041. For each network parameter in the current multi-task recognition model, calculate the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameter currently, to obtain the first gradient of the network parameter currently for the preset attribute.
[0167] S1051. Calculate the second gradient of the network parameter currently for the preset attribute.
[0168] S1061. Calculate the difference between the first gradient and the second gradient of the network parameter currently for the preset attribute, to obtain the target gradient of the network parameter currently for the preset attribute.
[0169] S1071. Obtain the sum of the target gradients of the network parameter currently for each preset attribute, adjust the network parameter, to obtain the adjustment result of this batch, so as to update the current multi-task recognition model;
[0170] S1072. Obtain the next batch of sample images as the current sample image to be processed, and return to execute the step of inputting the current sample image to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample image to be processed for each preset attribute;
[0171] It can be understood that the sample image can include multiple batches of images, each batch of images can include multiple images. Using one batch of images can perform one round of training on the multi-task recognition model to obtain the network parameters adjusted after this round of training. The electronic device can input each batch of sample images as the current sample image to be processed into the current multi-task recognition model, process the current sample image to be processed using the current network parameters, to obtain the current adjusted network parameters as the adjustment result of this batch. After processing the sample images of each batch, multiple batches of adjustment results can be obtained.
[0172] It should be noted that when processing the sample images of each batch, the adjustment result of the previous batch of this batch is used as the network parameter of the multi-task recognition model for processing the sample images of this batch. Therefore, the processing effect of this batch is usually better than that of the previous batch. Thus, the accuracy of the adjustment result of this batch is higher than the accuracy of the adjustment result of the previous batch, that is, the later the batch to which the adjustment result belongs, the higher the accuracy of the adjustment result.
[0173] S1073. Calculate the average level of the adjustment results of multiple batches to obtain the network parameters of the trained multi-task recognition model.
[0174] It can be understood that the electronic device can calculate the average level of the adjustment results of multiple batches. Optionally, the mean value of the adjustment results of multiple batches can be calculated to obtain the network parameters of the trained multi-task recognition model; or, a specified number of adjustment results can be selected from the adjustment results of multiple batches, and the mean value of the specified number of adjustment results can be calculated to obtain the network parameters of the trained multi-task recognition model.
[0175] In the embodiments of the present invention, using the average level of the adjustment results of each batch to obtain the trained network parameters can avoid the overfitting problem caused by only focusing on the adjustment results of a certain batch, improve the accuracy of the results output by the trained multi-task recognition model, and at the same time improve the generalization of the trained multi-task recognition model.
[0176] Optionally, in one implementation manner, step S1073 includes calculating the weighted sum of the adjustment results of each batch to obtain the network parameters of the trained multi-task recognition model.
[0177] Among them, the weight of the adjustment result of each batch is negatively correlated with the generated duration of the adjustment result of this batch; the generated duration of the adjustment result of this batch indicates the interval duration between the generation moment of the adjustment result of this batch and the current moment.
[0178] It can be understood that the electronic device can calculate the interval duration between the generation moment of the adjustment result of each batch and the current moment to obtain the generated duration of the adjustment result of this batch; then, according to the relationship that the generated duration is negatively correlated with the weight, set the weight of the adjustment result of each batch; finally, calculate the weighted sum of the adjustment results of each batch according to the set weights of the adjustment results of each batch to obtain the network parameters of the trained multi-task recognition model. Specifically, it can be processed in the way of Exponential Moving Average (EMA). In the process of setting the weight of the adjustment result of each batch according to the relationship that the generated duration is negatively correlated with the weight, the weight of the adjustment result of the first batch can be set to an initial value, and then, according to the relationship that the generated duration is negatively correlated with the weight, on the basis of the initial value, increase the weight of the adjustment result of each batch by a specified multiple to obtain the weight of the adjustment result of each batch.
[0179] Optionally, in another implementation manner, the Stochastic Weight Averaging (SWA) method can be used to calculate the average level of the adjustment results of multiple batches. Specifically, the adjustment results of the later batches are more accurate. A specified number of adjustment results can be selected from the adjustment results of the last half of the batches, and the mean value of the specified number of adjustment results can be calculated to obtain the network parameters of the trained multi-task recognition model.
[0180] In the embodiments of the present invention, the average level of the adjustment results of each batch is used to obtain the trained network parameters, which can avoid the overfitting problem caused by only focusing on the adjustment results of a certain batch, and improve the generalization of the trained multi-task recognition model while improving the accuracy of the results output by the trained multi-task recognition model.
[0181] Optionally, in one embodiment, the multi-task recognition model includes: an input network, a feature extraction network, and an output network. Based on Figure 2 the training method of the multi-task recognition model shown above, Figure 3 FIG. is a schematic flowchart of the processing steps performed based on the network structure of the multi-task recognition model in a training method of a multi-task recognition model. As Figure 3 shown, step S1022 includes S1022a-S1022c.
[0182] S1022a: Input the sample image into the input network to obtain a first feature vector.
[0183] It can be understood that the input network may include two convolutional layers with a size of 3. Through the input network, preliminary feature extraction can be performed on the sample image to obtain a first feature vector.
[0184] S1022b: Input the first feature vector into the feature extraction network to obtain a second feature vector.
[0185] It can be understood that the feature extraction network may include multiple convolutional layers. Through the feature extraction network, further feature extraction can be performed on the first feature vector to obtain a second feature vector.
[0186] Optionally, the feature extraction network may be a network structure set based on a residual network, and the feature extraction network includes multiple feature extraction sub-networks. During the process of the feature extraction network processing the first feature vector, the output data of each feature extraction sub-network except the first feature extraction sub-network in the feature extraction network and the output data of the previous feature extraction sub-network of this feature extraction sub-network may be superimposed as the input data of the next feature extraction sub-network of this feature extraction sub-network, so that the next feature extraction sub-network can perform feature extraction based on more information; for the first feature extraction sub-network, the output data of the input network may be used as the input data of the first feature extraction sub-network.
[0187] Exemplarily, in one implementation, the feature extraction network includes four feature extraction sub-networks, and the feature extraction network includes a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fourth feature extraction sub-network. Step S1022b includes steps C1-C4.
[0188] C1. Input the first eigenvector into the first feature extraction sub-network to obtain the first sub-eigenvector.
[0189] C2. Input the first sub-eigenvector and the first eigenvector into the second feature extraction sub-network to obtain the second sub-eigenvector.
[0190] C3. Input the second sub-eigenvector and the first sub-eigenvector into the third feature extraction sub-network to obtain the third sub-eigenvector.
[0191] C4. Input the third sub-eigenvector and the second sub-eigenvector into the fourth feature extraction sub-network to obtain the second eigenvector.
[0192] It can be understood that the first feature extraction sub-network, the second feature extraction sub-network, the third feature extraction sub-network, and the fourth feature extraction sub-network all include two convolutional layers with a size of 1 and one convolutional layer with a size of 3. The first feature extraction sub-network, the second feature extraction sub-network, the third feature extraction sub-network, and the fourth feature extraction sub-network can be connected through a residual network. Specifically, the first eigenvector can be input into the first feature extraction sub-network to obtain the first sub-eigenvector; then the first sub-eigenvector and the first eigenvector are input into the second feature extraction sub-network to obtain the second sub-eigenvector; then the second sub-eigenvector and the first sub-eigenvector are input into the third feature extraction sub-network to obtain the third sub-eigenvector, and finally the third sub-eigenvector and the second sub-eigenvector are input into the fourth feature extraction sub-network to obtain the second eigenvector.
[0193] S1022c. Process the second eigenvector based on the output network to obtain the predicted attribute information of the current to-be-processed sample image for each preset attribute.
[0194] It can be understood that the output network can include a pooling layer. Through the output network, the second eigenvector can be further feature-extracted and the dimension can be reduced to obtain the predicted attribute information of the current to-be-processed sample image for each preset attribute.
[0195] Optionally, in one implementation, the second eigenvector and the third sub-eigenvector can be input into the output network to obtain the predicted attribute information of the current to-be-processed sample image for each preset attribute.
[0196] In the embodiments of the present invention, the image can be processed through the network structure of the multi-task recognition model to obtain the predicted attribute information of the current to-be-processed sample image for each preset attribute. The network structure of the multi-task recognition model is relatively simple and can be deployed on a camera to directly process the collected image to obtain the attribute information, with a wide range of application scenarios.
[0197] The following combines with Figure 4 to introduce a multi-task recognition method provided by an embodiment of the present invention. As Figure 4 shown, the method may include steps S401-S402.
[0198] S401, obtain a to-be-recognized image including a to-be-recognized object.
[0199] It can be understood that the multi-task recognition method can be applied to an electronic device that can obtain videos collected by a camera and can recognize image frames in the collected videos. Exemplarily, the electronic device can be a camera or a server. Optionally, when the server can have both the functions of training a network and recognizing images, the execution subject of the multi-task recognition method and the multi-task recognition model training method can be the same execution subject; optionally, when the server can only have the function of training a network or recognizing images, the execution subject of the multi-task recognition method and the multi-task recognition model training method can not be the same execution subject.
[0200] Optionally, in one implementation, the to-be-recognized object is a vehicle, and each preset attribute includes any combination of the following: vehicle color, vehicle type, and license plate color.
[0201] It can be understood that when the execution subject is a camera, the camera can collect road information at the set angular position. The camera can select image frames at specified intervals from the image frames collected in a period of time as to-be-recognized images for processing; the camera can also use each image frame collected in real time as a to-be-recognized image.
[0202] S402, input the to-be-recognized image into a pre-trained multi-task recognition model to obtain attribute information of the to-be-recognized object for each preset attribute.
[0203] Among them, the multi-task recognition model is trained based on the multi-task recognition model training method provided in the foregoing embodiment.
[0204] It can be understood that the pre-trained multi-task recognition model can accurately recognize the attribute information of each preset attribute. Inputting the to-be-recognized image into the pre-trained multi-task recognition model can obtain the attribute information of the to-be-recognized object for each preset attribute.
[0205] In the embodiment of the present invention, since the generalization of the multi-task recognition model trained by using the multi-task recognition model training method provided in the foregoing embodiment is improved, the accuracy of the attribute information obtained by applying the multi-task recognition model is also improved.
[0206] Optionally, in one implementation, the multi-task recognition method further includes: if the attribute information of the object to be recognized for each preset attribute reaches the preset threshold of the preset attribute, then the attribute information of the object to be recognized for each preset attribute is used as the output result.
[0207] It can be understood that a threshold filtering method can be used to reject incorrect classification results. Specifically, a preset threshold can be set for each preset attribute in advance. If the attribute information of the object to be recognized for each preset attribute reaches the preset threshold of the preset attribute, then the attribute information of the object to be recognized for each preset attribute is used as the output result.
[0208] Optionally, in one implementation, a preset detection model can be used to determine the positions of the same object to be recognized in multiple frames of images to be recognized, and obtain the attribute information of the same object to be recognized in different images to be recognized. If the attribute information is the same, then the attribute information is used as the output result. Among them, the preset detection model can be a model set based on Kalman filtering.
[0209] To more clearly understand the effects of the embodiments of the present invention, the multi-task recognition model trained by the embodiments of the present invention is tested using the CCPD dataset. It can be found that the recognition accuracy for vehicle types reaches 92.03%, the recognition accuracy for vehicle colors reaches 91.27%, and the recognition accuracy for license plate types reaches 97.38%. Each task has a relatively high recognition accuracy. Moreover, in actual chip tests, the multi-task recognition model trained by the embodiments of the present invention can reach a processing frame rate of 10 frames per second, enabling the model to well handle the application requirements of various scenarios. For example, cameras installed in parking lots can directly use the multi-task recognition model for real-time recognition locally without relying on cloud services and avoiding interactions with databases.
[0210] Figure 5 The following is a schematic structural diagram of a multi-task recognition model provided by an embodiment of the present invention. As Figure 5 shown, the multi-task recognition model includes: an input network, a feature extraction network, and an output network. The feature extraction network includes a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fourth feature extraction sub-network. The input network may include two convolutional layers with a size of 3×3. The first feature extraction sub-network, the second feature extraction sub-network, the third feature extraction sub-network, and the fourth feature extraction sub-network all include two convolutional layers with a size of 1×1 and one convolutional layer with a size of 3×3. The output network may include a pooling layer.
[0211] Specifically, the sample image can be used as input data. Through the input network, preliminary feature extraction can be performed on the sample image to obtain the first feature vector. Then, the first feature vector is input into the first feature extraction sub-network to obtain the first sub-feature vector. Then, the first sub-feature vector and the first feature vector are input into the second feature extraction sub-network to obtain the second sub-feature vector. Then, the second sub-feature vector and the first sub-feature vector are input into the third feature extraction sub-network to obtain the third sub-feature vector. Finally, the third sub-feature vector and the second sub-feature vector are input into the fourth feature extraction sub-network to obtain the second feature vector. Through the output network, further feature extraction can be performed on the second feature vector and the third sub-feature vector, and the dimension can be reduced to obtain the predicted attribute information of the current sample image to be processed for each preset attribute. The obtained predicted attribute information can be used as output data.
[0212] The embodiment of the present invention also provides a training device for a multi-task recognition model, as Figure 6 shown. The training device for the multi-task recognition model includes:
[0213] The first acquisition module 610 is configured to acquire a sample image including a sample object, and acquire the true attribute information of the sample object for each preset attribute;
[0214] The first processing module 620 is configured to use the current multi-task recognition model to process the sample image to obtain the predicted attribute information of the sample object for each preset attribute;
[0215] The second acquisition module 630 is configured to obtain the loss value corresponding to the preset attribute based on the difference between the predicted attribute information and the true attribute information of the sample object for each preset attribute;
[0216] The first calculation module 640 is configured to calculate, for each network parameter in the current multi-task recognition model, the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameter, to obtain the first gradient of the network parameter for the preset attribute;
[0217] The second calculation module 650 is configured to calculate the second gradient of the network parameter for the preset attribute; wherein, the second gradient of the network parameter for the preset attribute is positively correlated with the loss value of the sample object for the preset attribute;
[0218] The third calculation module 660 is configured to calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute, to obtain the target gradient of the network parameter for the preset attribute;
[0219] An adjustment module 670, configured to adjust the network parameters based on the sum of the target gradients of each preset attribute with respect to the network parameters, so as to obtain a trained multi-task recognition model.
[0220] Optionally, the second calculation module 650 is specifically configured to perform a norm constraint on the first gradient of the network parameters with respect to the preset attribute to obtain a second gradient of the network parameters with respect to the preset attribute.
[0221] Optionally, the first acquisition module 610 is specifically configured to acquire a plurality of initial images each containing a sample object;
[0222] Perform random image processing and / or image fusion processing on the plurality of initial images to obtain sample images; wherein, the random image processing includes: random horizontal flipping, random image rotation, and random image blurring.
[0223] Optionally, the device further includes:
[0224] A fourth acquisition module, configured to acquire the true attribute information of the sample object in each initial image for each preset attribute;
[0225] The first acquisition module 610 is specifically configured to, when each sample image is obtained by processing the corresponding initial image based on the random image processing, use the true attribute information of the sample object in the initial image corresponding to the sample image for each preset attribute as the true attribute information of the sample object for each preset attribute;
[0226] When the sample image is obtained by processing the corresponding initial image based on the image fusion processing, fuse the true attribute information of the sample object in the initial image corresponding to the sample image for each preset attribute according to a preset weight to obtain the true attribute information of the sample object for each preset attribute.
[0227] Optionally, the sample image includes images of multiple batches;
[0228] The first processing module 620 includes:
[0229] A first acquisition unit, configured to acquire the sample images of the first batch as the current sample images to be processed;
[0230] An input unit, configured to input the current sample images to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample images to be processed for each preset attribute;
[0231] The second acquisition module 630 is specifically configured to, for each preset attribute, obtain the loss value corresponding to the preset attribute based on the difference between the predicted attribute information and the true attribute information of the preset attribute in the current sample image to be processed;
[0232] The first calculation module 640 is specifically configured to, for each network parameter in the current multi-task recognition model, calculate the change rate of the loss value of the sample object for each preset attribute in the direction of the network parameter currently, and obtain the first gradient of the network parameter currently for the preset attribute;
[0233] The second calculation module 650 is specifically configured to calculate the second gradient of the network parameter currently for the preset attribute;
[0234] The third calculation module 660 is specifically configured to calculate the difference between the first gradient and the second gradient of the network parameter currently for the preset attribute, and obtain the target gradient of the network parameter currently for the preset attribute;
[0235] The adjustment module 670 includes:
[0236] The second acquisition unit is configured to obtain the sum of the target gradients of the network parameter currently for each preset attribute, adjust the network parameter, and obtain the adjustment result of this batch to update the current multi-task recognition model;
[0237] The third acquisition unit is configured to obtain the sample image of the next batch as the current sample image to be processed, and return to execute the step of inputting the current sample image to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample image to be processed for each preset attribute;
[0238] The calculation unit is configured to calculate the average level of the adjustment results of multiple batches to obtain the network parameters of the trained multi-task recognition model.
[0239] Optionally, the calculation unit is specifically configured to calculate the weighted sum of the adjustment results of each batch to obtain the network parameters of the trained multi-task recognition model; wherein, the weight of the adjustment result of each batch is negatively correlated with the generated duration of the adjustment result of this batch; the generated duration of the adjustment result of this batch means: the time interval between the generation moment of the adjustment result of this batch and the current moment.
[0240] Optionally, the multi-task recognition model includes: an input network, a feature extraction network, and an output network;
[0241] The input unit includes:
[0242] The first input sub-unit is configured to input the sample image into the input network to obtain a first feature vector;
[0243] A second input subunit, configured to input the first feature vector into the feature extraction network to obtain a second feature vector;
[0244] A processing subunit, configured to process the second feature vector based on the output network to obtain prediction attribute information of the current sample image to be processed for each preset attribute.
[0245] Optionally, the feature extraction network includes a first feature extraction sub-network, a second feature extraction sub-network, a third feature extraction sub-network, and a fourth feature extraction sub-network;
[0246] The second input subunit is specifically configured to input the first feature vector into the first feature extraction sub-network to obtain a first sub-feature vector;
[0247] Input the first sub-feature vector and the first feature vector into the second feature extraction sub-network to obtain a second sub-feature vector;
[0248] Input the second sub-feature vector and the first sub-feature vector into the third feature extraction sub-network to obtain a third sub-feature vector;
[0249] Input the third sub-feature vector and the second sub-feature vector into the fourth feature extraction sub-network to obtain a second feature vector;
[0250] The processing subunit is specifically configured to input the second feature vector and the third sub-feature vector into the output network to obtain prediction attribute information of the current sample image to be processed for each preset attribute.
[0251] Optionally, the sample object is a vehicle, and each preset attribute includes any combination of the following: vehicle color, vehicle type, and license plate color.
[0252] In an embodiment of the present invention, for each network parameter in the current multi-task recognition model, since the second gradient of the network parameter with respect to each preset attribute is positively correlated with the loss value of the sample object with respect to the preset attribute, the greater the loss value, the greater the second gradient; the target gradient of the network parameter with respect to the preset attribute is the difference between the first gradient and the second gradient of the network parameter with respect to the preset attribute, and the target gradient is constrained by the second gradient. Since the second gradient is positively correlated with the loss value, that is, if the loss value of the sample object with respect to a preset attribute is greater, the constraint on the target gradient of a network parameter with respect to the preset attribute is greater. Through the second gradient, the target gradient of the network parameter with respect to the preset attribute can be restricted from being too large. That is, through the second gradient, the target gradient of the preset attribute with a high loss value can be restricted from being too large, reducing the proportion of the target gradient of the preset attribute with a high loss value in the sum of the target gradients of all preset attributes, and increasing the proportion of the target gradients of other preset attributes with lower loss values in the sum of the target gradients of all preset attributes. Therefore, the influence of the preset attributes with high loss values on the sum of the target gradients of all preset attributes can be reduced, that is, the influence of the tasks with high loss values on model learning is reduced, avoiding the model from overlearning the features of the tasks with high loss values, increasing the possibility of the model learning the features of other tasks, reducing the possibility of overfitting problems for the tasks with high loss values and underfitting problems for the tasks with small loss values, and enabling a trained multi-task recognition model to be obtained.
[0253] An embodiment of the present invention further provides a multi-task recognition device, as Figure 7 shown. The multi-task recognition device includes:
[0254] A third acquisition module 710, configured to acquire a to-be-recognized image including a to-be-recognized object;
[0255] An input module 720, configured to input the to-be-recognized image into a pre-trained multi-task recognition model to obtain attribute information of the to-be-recognized object for each preset attribute; wherein, the multi-task recognition model is trained by using the training method of the multi-task recognition model provided in the foregoing embodiment.
[0256] In an embodiment of the present invention, since the generalization of the multi-task recognition model trained by using the training method of the multi-task recognition model provided in the foregoing embodiment is improved, the accuracy of the attribute information obtained by applying the multi-task recognition model is also improved.
[0257] An embodiment of the present invention further provides an electronic device, as Figure 8 shown, including a processor 801, a communication interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communication interface 802, and the memory 803 complete communication with each other through the communication bus 804.
[0258] A memory 803 for storing computer programs;
[0259] A processor 801, when executing the programs stored on the memory 803, implements the training method and multi-task recognition method of the multi-task recognition model provided in the above embodiments.
[0260] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0261] The communication interface is used for communication between the above electronic device and other devices.
[0262] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0263] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0264] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the training method and multi-task recognition method of any of the above multi-task recognition models are implemented.
[0265] In another embodiment provided by the present invention, there is also provided a computer program product including instructions. When it runs on a computer, it causes the computer to execute the training method and the multi-task recognition method of any of the multi-task recognition models in the above embodiments.
[0266] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0267] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0268] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0269] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A training method for a multi-task recognition model, characterized in that: The method comprises: Acquire a sample image containing a sample object, and acquire real attribute information of the sample object for each preset attribute; Processing the sample image using the current multi-task recognition model to obtain predicted attribute information of the sample object for each preset attribute; Based on the difference between the predicted attribute information and the actual attribute information of the sample object for each preset attribute, obtaining a loss value corresponding to the preset attribute; For each network parameter in the current multi-task recognition model, calculate the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameter to obtain a first gradient of the network parameter for the preset attribute; Calculating a second gradient of the network parameter for the preset attribute; wherein the second gradient of the network parameter for the preset attribute is positively correlated with the loss value of the sample object for the preset attribute; Calculating a difference between a first gradient and a second gradient of the network parameter for the preset attribute to obtain a target gradient of the network parameter for the preset attribute; Based on the sum of the target gradients of the network parameters for each preset attribute, the network parameters are adjusted to obtain a trained multi-task recognition model.
2. The method according to claim 1, characterized in that Calculating a second gradient of the network parameter for the preset attribute includes: A norm constraint is performed on the first gradient of the network parameter with respect to the preset attribute to obtain a second gradient of the network parameter with respect to the preset attribute.
3. The method according to claim 1, characterized in that The step of obtaining a sample image containing a sample object comprises: Acquire multiple initial images containing sample objects; Perform random image processing and / or image fusion processing on multiple initial images to obtain sample images; wherein the random image processing includes: random horizontal flipping, random image rotation and random image blurring.
4. The method according to claim 3, characterized in that Before obtaining the real attribute information of the sample object for each preset attribute, the method further includes: Obtaining real attribute information of the sample object in each initial image for each preset attribute; The obtaining of the real attribute information of the sample object for each preset attribute includes: When each sample image is obtained by processing the corresponding initial image based on the random image processing, the real attribute information of the sample object for each preset attribute in the initial image corresponding to the sample image is used as the real attribute information of the sample object for each preset attribute; In the case where the sample image is obtained by processing the corresponding initial image based on the image fusion processing, the real attribute information of the sample object for each preset attribute in the initial image corresponding to the sample image is fused according to the preset weight to obtain the real attribute information of the sample object for each preset attribute.
5. The method according to claim 1, characterized in that: The sample images include multiple batches of images; The sample image is processed using the current multi-task recognition model to obtain predicted attribute information of the sample object for each preset attribute, including: Get the sample images of the first batch as the current sample images to be processed; Inputting the current sample image to be processed into the current multi-task recognition model to obtain predicted attribute information of the current sample image to be processed for each preset attribute; Based on the difference between the predicted attribute information and the actual attribute information of the sample object for each preset attribute, a loss value corresponding to the preset attribute is obtained, including: For each preset attribute, based on the difference between the predicted attribute information of the current sample image to be processed and the actual attribute information for the preset attribute, a loss value corresponding to the preset attribute is obtained; For each network parameter in the current multi-task recognition model, calculating the rate of change of the loss value of the sample object for each preset attribute in the direction of the network parameter, and obtaining a first gradient of the network parameter for the preset attribute, including: For each network parameter in the current multi-task recognition model, calculate the current change rate of the loss value of the sample object for each preset attribute in the direction of the network parameter, and obtain the first gradient of the network parameter for the preset attribute; Calculating a second gradient of the network parameter for the preset attribute includes: Calculating a second gradient of the network parameter for the preset attribute; Calculating the difference between the first gradient and the second gradient of the network parameter for the preset attribute to obtain a target gradient of the network parameter for the preset attribute includes: Calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute to obtain the target gradient of the network parameter for the preset attribute; Based on the sum of the target gradients of the network parameters for each preset attribute, the network parameters are adjusted to obtain a trained multi-task recognition model, including: Obtain the target gradient sum of the network parameter for each preset attribute, adjust the network parameter, and obtain the adjustment result of the batch to update the current multi-task recognition model; Obtaining the next batch of sample images as the current sample images to be processed, and returning to execute the step of inputting the current sample images to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample images to be processed for each preset attribute; Calculate the average level of adjustment results of multiple batches to obtain the network parameters of the trained multi-task recognition model.
6. The method according to claim 5, characterized in that The average level of the adjustment results of the plurality of batches is calculated to obtain the network parameters of the trained multi-task recognition model, including: The weighted sum of the adjustment results of each batch is calculated to obtain the network parameters of the trained multi-task recognition model; wherein the weight of each batch of adjustment results is negatively correlated with the generation time of the adjustment results of the batch; the generation time of the adjustment results of the batch represents: the interval between the generation time of the adjustment results of the batch and the current time.
7. The method according to claim 5, characterized in that The multi-task recognition model includes: an input network, a feature extraction network and an output network; Input the current sample image to be processed into the current multi-task recognition model to obtain the predicted attribute information of the current sample image to be processed for each preset attribute, including: Inputting the sample image into the input network to obtain a first feature vector; Inputting the first feature vector into the feature extraction network to obtain a second feature vector; The second feature vector is processed based on the output network to obtain predicted attribute information of the current sample image to be processed for each preset attribute.
8. The method according to claim 7, characterized in that The feature extraction network includes a first feature extraction subnetwork, a second feature extraction subnetwork, a third feature extraction subnetwork and a fourth feature extraction subnetwork; Inputting the first feature vector into the feature extraction network to obtain a second feature vector includes: Inputting the first feature vector into the first feature extraction sub-network to obtain a first sub-feature vector; Inputting the first sub-feature vector and the first feature vector into the second feature extraction sub-network to obtain a second sub-feature vector; Inputting the second sub-feature vector and the first sub-feature vector into the third feature extraction sub-network to obtain a third sub-feature vector; Inputting the third sub-feature vector and the second sub-feature vector into the fourth feature extraction sub-network to obtain a second feature vector; The processing of the second feature vector based on the output network to obtain predicted attribute information of the current sample image to be processed for each preset attribute includes: The second feature vector and the third sub-feature vector are input into the output network to obtain the predicted attribute information of the current sample image to be processed for each preset attribute.
9. The method according to any one of claims 1 to 8, characterized in that: The sample object is a vehicle, and each preset attribute includes any combination of the following: vehicle color, vehicle type, and license plate color.
10. A multi-task recognition method, characterized in that: The method comprises: Acquire an image to be identified that contains an object to be identified; The image to be identified is input into a pre-trained multi-task recognition model to obtain attribute information of the object to be identified for each preset attribute; wherein the multi-task recognition model is trained based on the method described in any one of claims 1 to 9 above.
11. The method according to claim 10, characterized in that The object to be identified is a vehicle, and each preset attribute includes any combination of the following: vehicle color, vehicle type, and license plate color.
12. A training device for a multi-task recognition model, characterized in that: The device comprises: A first acquisition module, used to acquire a sample image containing a sample object, and to acquire real attribute information of the sample object for each preset attribute; A first processing module, used to process the sample image using the current multi-task recognition model to obtain predicted attribute information of the sample object for each preset attribute; A second acquisition module, configured to obtain a loss value corresponding to each preset attribute based on the difference between the predicted attribute information and the real attribute information of the sample object for each preset attribute; A first calculation module is used to calculate, for each network parameter in the current multi-task recognition model, a change rate of the loss value of the sample object for each preset attribute in the direction of the network parameter, and obtain a first gradient of the network parameter for the preset attribute; A second calculation module, used to calculate a second gradient of the network parameter for the preset attribute; wherein the second gradient of the network parameter for the preset attribute is positively correlated with a loss value of the sample object for the preset attribute; A third calculation module is used to calculate the difference between the first gradient and the second gradient of the network parameter for the preset attribute, and obtain a target gradient of the network parameter for the preset attribute; The adjustment module is used to adjust the network parameters based on the sum of the target gradients of the network parameters for each preset attribute to obtain a trained multi-task recognition model.
13. A multi-task recognition device, characterized in that: The device comprises: A third acquisition module is used to acquire an image to be identified that contains an object to be identified; An input module is used to input the image to be identified into a pre-trained multi-task recognition model to obtain attribute information of the object to be identified for each preset attribute; wherein the multi-task recognition model is trained based on the method described in any one of claims 1 to 9 above.
14. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-11 when executing a program stored in a memory.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
16. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 11 when executed by a processor.