Multi-task model training method and device, electronic equipment and medium

By using the feature extraction network of the single-task model in multi-task model training to calculate the distillation loss and updating the feature extraction network of the multi-task model, the negative migration problem caused by inter-task competition in multi-task joint training is solved, and the performance of the multi-task model is improved.

CN120107753APending Publication Date: 2025-06-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510263033.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

During the multi-task joint training process, the multi-task negative migration problem caused by competition between tasks affects the performance of the multi-task model.

Method used

By inputting perceptual image samples into the feature extraction network in the pre-trained single-task model and the multi-task model to be trained, the distillation loss between the single-task features and the multi-task features is calculated, and the feature extraction network of the multi-task model is updated based on this to improve the performance of the multi-task model.

Benefits of technology

It effectively solves the problem of negative migration of multi-tasks, improves the performance of multi-task models, and enables the feature extraction network of multi-task models to better extract features that meet the needs of sub-tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107753A_ABST
    Figure CN120107753A_ABST
Patent Text Reader

Abstract

The invention provides a multi-task model training method, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes such as automatic driving and the like. The method comprises the following steps: respectively inputting a perception image sample into a feature extraction network in a pre-trained single-task model and a feature extraction network in a to-be-trained multi-task model to obtain a single-task feature and a multi-task feature; one subtask branch of the multi-task model is the same as the task of the single-task model; determining distillation loss between the single-task feature and the multi-task feature based on the task type of the single-task model; based on the distillation loss, performing parameter updating on a feature extraction network of the multi-task model to train the multi-task model; the feature extraction network is shared by each sub-task branch in the multi-task model. According to the invention, the problem of multi-task negative migration is solved, and the performance of the multi-task model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to computer vision, deep learning, large models and other technical fields, and can be applied to scenarios such as autonomous driving. Specifically, it relates to a multi-task model training method. Background Art

[0002] With the development of end-to-end autonomous driving technology, multi-task joint training has become the mainstream of autonomous driving model development. Multi-task joint training refers to training a multi-task model to process multiple subtasks at the same time. In a multi-task model, multiple subtasks share the same backbone network to extract underlying common features for subtask branches to process different tasks. Summary of the invention

[0003] The present invention provides a multi-task model training method, device, electronic device and medium.

[0004] According to one aspect of the present disclosure, a multi-task model training method is provided, the method comprising:

[0005] Inputting the perceived image samples into the feature extraction networks of the pre-trained single-task model and the multi-task model to be trained respectively to obtain single-task features and multi-task features; one of the subtask branches of the multi-task model is the same as the task of the single-task model;

[0006] Determining a distillation loss between the single-task feature and the multi-task feature based on a task type of the single-task model;

[0007] Based on the distillation loss, parameters of a feature extraction network of the multi-task model are updated to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

[0008] According to another aspect of the present disclosure, a multi-task model training device is provided, the device comprising:

[0009] A feature extraction module, used to input the perceived image samples into the feature extraction networks of the pre-trained single-task model and the multi-task model to be trained, respectively, to obtain single-task features and multi-task features; one of the subtask branches of the multi-task model is the same as the task of the single-task model;

[0010] A distillation loss determination module, used to determine the distillation loss between the single-task feature and the multi-task feature based on the task type of the single-task model;

[0011] A first parameter updating module is used to update the parameters of the feature extraction network of the multi-task model based on the distillation loss to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

[0012] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multi-task model training method described in any embodiment of the present disclosure.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the multi-task model training method described in any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the multi-task model training method described in any embodiment of the present disclosure.

[0018] The present invention solves the multi-task negative transfer problem caused by the influence of mutual competition between tasks during multi-task joint training, and improves the performance of the multi-task model.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0021] Figure 1A is a flowchart of a multi-task model training method provided according to an embodiment of the present disclosure;

[0022] Figure 1B is a schematic diagram of a multi-task model training method provided according to an embodiment of the present disclosure;

[0023] Figure 2 is a flowchart of another multi-task model training method provided according to an embodiment of the present disclosure;

[0024] Figure 3A is a flowchart of another multi-task model training method provided according to an embodiment of the present disclosure;

[0025] Figure 3B It is a schematic diagram of a distillation effective feature determination scheme provided for the case where the task type of a single-task model is lane line detection;

[0026] Figure 4 is a structural schematic diagram of a multi-task model training device provided according to an embodiment of the present disclosure;

[0027] Figure 5 A block diagram of an electronic device used to implement a multi-task model training method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0029] Figure 1A 1 is a flowchart of a multi-task model training method provided according to an embodiment of the present disclosure. The embodiment of the present disclosure can be applied to the case of training a multi-task model for autonomous driving. The method can be executed by a multi-task model training device, which can be implemented in software and / or hardware. Figure 1A As shown, the multi-task model training method of this embodiment may include:

[0030] S101, inputting the perceived image samples into the feature extraction networks of the pre-trained single-task model and the multi-task model to be trained respectively to obtain single-task features and multi-task features; one of the sub-task branches of the multi-task model is the same as the task of the single-task model.

[0031] S102, determining a distillation loss between the single-task feature and the multi-task feature based on the task type of the single-task model.

[0032] S103, based on the distillation loss, updating the parameters of the feature extraction network of the multi-task model to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

[0033] Among them, the multi-task model can process deep learning models of multiple related tasks at the same time. The multi-task model includes a feature extraction network and at least two sub-task branches, each sub-task branch includes a sub-task decoder. The tasks of each sub-task branch in the multi-task model are different. The feature extraction network in the multi-task model is shared by each sub-task branch in the multi-task model. Among them, the multi-task feature refers to the underlying common features extracted from the perceived image samples by the feature extraction network in the multi-task model. The underlying common features can be used together by each sub-task branch in the multi-task model. In the autonomous driving scenario, optionally, the single-task feature and the multi-task feature are bird's-eye view features (BEV, Bird Eye View).

[0034] A single-task model refers to an independent deep learning model trained for a specific task. The single-task model includes a feature extraction network and a single-task decoder. The single-task features are extracted from the perceived image samples by the feature extraction network in the single-task model. Single-task models do not share parameters or information. One of the subtask branches of the multi-task model has the same task as the single-task model. Exemplarily, in an autonomous driving scenario, the multi-task model can simultaneously handle tasks such as lane line detection, obstacle detection, and space occupancy detection. The single-task model only focuses on tasks such as lane line detection, obstacle detection, or space occupancy detection.

[0035] The multi-task model to be trained needs to be trained using a pre-trained single-task model and perception image samples. The perception image samples refer to the image data used to train the multi-task model. Optionally, the perception image samples are obtained by capturing images of the vehicle's surroundings from multiple perspectives. The acquisition device and acquisition method of the perception image samples are not limited here. For example, the perception image samples can be acquired by a surround view camera mounted on the vehicle.

[0036] The pre-trained single-task model is the teacher model, and the multi-task model to be trained is the student model. The single-task features of the single-task model are used as soft labels for the multi-task model to guide the feature extraction network of the multi-task model to learn the feature extraction capabilities of the feature extraction network in the single-task model, making the multi-task features more and more similar to the single-task features.

[0037] Among them, distillation loss can quantify the feature difference between single-task features and multi-task features. Distillation loss can evaluate the accuracy of multi-task models in imitating the feature extraction behavior of single-task models. Distillation loss is related to the task type of single-task models.

[0038] It is understandable that for single-task models, the single-task features extracted from different tasks will have obvious task tendencies. For example, the single-task features of the lane detection model may only include information related to the lane, while the single-task features of the obstacle detection model may only include information related to the obstacle.

[0039] The task type of the single-task model is used to determine the distillation loss between the single-task features and the multi-task features. The distillation loss is then used in the back-propagation algorithm to update the parameters of the feature extraction network of the multi-task model. The multi-task features extracted by the feature extraction network of the multi-task model can be closer to the single-task features and more in line with the task tendency.

[0040] The disclosed technical solution uses a pre-trained single-task model to guide the training process of the multi-task model, uses the single-task features extracted by the feature extraction network in the single-task model as soft labels, and performs feature distillation on the multi-task features extracted by the feature extraction network in the multi-task model, thereby avoiding the situation where the feature extraction network of the multi-task model is difficult to optimize the most suitable multi-task features for each sub-task due to competition between tasks, and falls into a suboptimal solution, thereby solving the multi-task negative transfer problem and improving the performance of the multi-task model. In the process of feature distillation, the disclosed technical solution takes into account that the single-task features extracted from different tasks will have obvious task tendencies, and uses the task type of the single-task model to determine the distillation loss between the single-task features and the multi-task features, making full use of the prior knowledge of the single-task model, improving the effect of feature distillation, and making the multi-task features extracted by the feature extraction network of the multi-task model more in line with the sub-task requirements, which is conducive to improving the performance of the multi-task model.

[0041] In an optional embodiment, the method also includes: determining the subtask branch in the multi-task model that is the same as the single-task model task as the subtask branch of the target subtask; wherein the subtask branch includes a task decoder; inputting the multi-task features into the task decoder of the target subtask to obtain a prediction result of the target subtask; determining the true value result of the target subtask based on the sample label of the perceived image sample; and updating the parameters of the subtask branch of the target subtask in the multi-task model based on the prediction result and the true value result.

[0042] The multi-task model includes at least sub-task branches, each of which is used to perform a task. The single-task model has the same task as one of the sub-task branches in the multi-task model. The task of the single-task model is determined as the target sub-task. The sub-task branch in the multi-task model that is the same as the task of the single-task model is determined as the sub-task branch of the target sub-task.

[0043] The task decoder included in the target subtask branch is specially designed for the target subtask and is used to generate a prediction result for the target subtask.

[0044] The multi-task features are input into the task decoder of the target subtask, and the task decoder of the target subtask outputs the prediction result of the target subtask based on the multi-task features. The multi-task features are extracted from the perceptual image samples by the feature extraction network in the multi-task model, and the perceptual image samples are associated with corresponding sample labels. The true value result of the target subtask can be determined based on the sample labels.

[0045] Optionally, based on the prediction results and the true results, the task loss of the target subtask is determined. The task loss is used to quantify the degree of inconsistency between the prediction results and the true results. The task loss can be used to evaluate the effectiveness of multi-task features. The smaller the task loss, the closer the prediction results of the multi-task model for the target subtask are to the true results, and the more effective the multi-task features are for the target subtask. This indirectly shows that the feature extraction network of the multi-task model has extracted the features required for the target subtask.

[0046] When the task loss is determined, the parameters of the subtask branch of the target subtask in the multi-task model are updated based on the task loss.

[0047] The above technical solution verifies the task effectiveness of the multi-task features after performing feature distillation on the multi-task features using the single-task features, closely links feature extraction with task performance, and ensures the accuracy of the multi-task features.

[0048] In an optional embodiment, the method further includes: determining a target perceptual image that matches the task of the single-task model from candidate perceptual images according to the task type of the single-task model; and using the target perceptual image as a perceptual image sample used when training the multi-task model using the single-task model.

[0049] There are significant differences and particularities in the training samples of different tasks. In the case of using a single-task model to guide the training of a multi-task model, the present disclosure supports the use of training samples unique to the single-task model tasks to train the multi-task model, rather than the same source samples that need to meet the requirements of multi-task learning at the same time.

[0050] The candidate perceptual images include perceptual images that are diverse enough to cover the learning requirements of different tasks. The target perceptual images refer to perceptual images that match the single-task model task. The target perceptual images can be training samples unique to the single-model task.

[0051] When the target perception image is determined, the target perception image is used as a perception image sample to train the multi-task model to be trained.

[0052] The above technical solution, in the process of using a pre-trained single-task model to guide the training process of the multi-task model, supports the training of the multi-task model using perceptual images that match the single-task model tasks as perceptual image samples, thereby reducing the requirements for training samples for multi-task model training.

[0053] In an optional embodiment, the method further includes: within a training cycle, using at least two of the single-task models to alternately train the multi-task model.

[0054] Generally speaking, the number of single-task models is the same as the number of subtask branches in the multi-task model. In other words, for each task that can be processed by the multi-task model, the corresponding single-task model can be used to guide the training process of the multi-task model. One training cycle should be able to cover each task that can be processed by the multi-task model. The multi-task model is trained in training cycles until the multi-task model converges.

[0055] Figure 1B is a schematic diagram of a multi-task model training method provided according to an embodiment of the present disclosure. It is worth noting that Figure 1B Only the case where the multi-task model can process two tasks at the same time is shown. Among them, the single-task model 1 has the same task as the sub-task branch 1 in the multi-task model, and the single-task model 2 has the same task as the sub-task branch 2 in the multi-task model.

[0056] See also Figure 1B The process of training the multi-task model using the single-task model 1 is as follows: 1) input the perceptual image samples into the multi-task model to be trained and the pre-trained single-task model 1 respectively, and output the single-task feature 1 through the feature extraction network of the single-task model 1; 2) output the multi-task feature through the feature extraction network of the multi-task model; 3) calculate the distillation loss between the single-task feature 1 and the multi-task feature; use the distillation loss to update the parameters of the feature extraction network of the multi-task model, and input the multi-task feature into the sub-task decoder 1 in the sub-task branch 1 to obtain the sub-task prediction result 1. Optionally, determine the task loss 1 by comparing the single-task prediction result 1 with the true value result 1 corresponding to the perceptual image sample, and use the task loss 1 to update the parameters of the sub-task branch 1 of the multi-task model.

[0057] The process of training the multi-task model using the single-task model 2 is similar to the process of training the multi-task model using the single-task model 1, and will not be repeated here. It is worth noting that the single-task model 1 and the single-task model 2 train the multi-task model alternately. After the single-task model 1 trains the multi-task model, the single-task model 2 trains the multi-task model. The order in which the single-task model trains the multi-task model is not limited here. The multi-task model can be trained using the single-task model 1 first, or the multi-task model can be trained using the single-task model 2 first.

[0058] After the single-task model 1 and the single-task model 2 have trained the multi-task model once, a training cycle ends. The multi-task model is trained in training cycles until the multi-task model converges.

[0059] It is worth noting that Figure 1B The situation shown does not limit the multi-model training method provided by the embodiment of the present disclosure. The multi-model training method provided by the embodiment of the present disclosure is also applicable to a multi-task model that processes more than two tasks at the same time. Exemplarily, the tasks that can be processed by the multi-task model include lane line detection tasks, obstacle detection tasks, and space occupancy detection tasks. Correspondingly, the single-task model includes a lane line detection model, an obstacle detection model, and a space occupancy detection model. It is worth noting that Figure 1B The process of determining the distillation loss is not expanded. The task type of the single-task model needs to be considered in the process of determining the distillation loss.

[0060] The above technical solution uses at least two single-task models to alternately train the multi-task model within one training cycle. This can avoid the situation where the feature extraction network of the multi-task model optimizes the multi-task features to an intermediate state due to the differences in single-task features, making it difficult for the single-task features to be transferred to the multi-task features, which is beneficial to improving the effect of feature distillation.

[0061] Figure 2 It is a flowchart of another multi-task model training method provided according to an embodiment of the present disclosure; this embodiment is an optional scheme proposed on the basis of the above embodiment.

[0062] See also Figure 2 , the multi-task model training method provided in this embodiment includes:

[0063] S201, input the perceived image samples into the feature extraction networks of the pre-trained single-task model and the multi-task model to be trained respectively to obtain single-task features and multi-task features; one of the sub-task branches of the multi-task model is the same as the task of the single-task model.

[0064] S202: Determine distillation effective features from the multi-task features and the single-task features based on the task type of the single-task model.

[0065] Among them, the effective features of distillation are related to the task type of the single-task model. They not only conform to the task tendency of the single-task model tasks, but also take into account the sharing and correlation of multiple tasks. The effective features of distillation can be obtained by extracting single-task features or multi-task features. The feature extraction method corresponding to the effective features of distillation is related to the task type of the single-task model. The effective features of distillation are used to determine the distillation loss.

[0066] In an optional embodiment, the step of determining the distillation effective features from the multi-task features and the single-task features based on the task type of the single-task model also includes: if the task type of the single-task model is space occupancy detection, determining the single-task features and the multi-task features as the distillation effective features.

[0067] Occupancy detection usually divides the 3D space into voxel grids and predicts the probability of each voxel grid being occupied and the target category it may contain. Occupancy detection considers multi-dimensional features such as spatial occupancy information, geometric shape information (such as size and shape outline), temporal dynamic information (such as motion trajectory, speed and acceleration) and attribute information (such as category, size and color), and has high feature richness.

[0068] Therefore, when the task type of the single-task model is space occupancy detection, the single-task features and the multi-task features can be directly determined as distillation effective features.

[0069] The above technical solution provides a practical and feasible solution for determining effective features by distillation when the task type of the single-task model is space occupancy detection. While ensuring the effect of feature distillation, it simplifies the process of determining the distillation loss and enriches the applicable scenarios of the multi-task model training method.

[0070] S203: Determine the distillation loss between the single-task feature and the multi-task feature based on the distillation effective feature.

[0071] The distillation effective features are generated in the single-task features and the multi-task features. The distillation loss between the single-task features and the multi-task features can be determined based on the distillation effective features.

[0072] S204, based on the distillation loss, updating parameters of the feature extraction network of the multi-task model to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

[0073] The distillation loss determined by the effective features of distillation is used to update the parameters of the feature extraction network of the multi-task model, which can ensure that the multi-task model can master the feature extraction and generalization capabilities of the feature extraction network in the single-task model.

[0074] The technical solution disclosed in the present invention determines the distillation loss by utilizing the effective features of distillation that take into account the task type. Compared with determining the distillation loss based on single-task features and multi-task features without considering the task type and without task distinction, the multi-task features extracted by the feature extraction network of the multi-task model can be closer to the single-task features and more in line with the task tendency, thereby effectively improving the feature distillation effect.

[0075] Figure 3A It is a flowchart of another multi-task model training method provided according to an embodiment of the present disclosure; this embodiment is an optional scheme proposed on the basis of the above embodiment.

[0076] See also Figure 3A , the multi-task model training method provided in this embodiment includes:

[0077] S301, input the perceived image samples into the feature extraction networks of the pre-trained single-task model and the multi-task model to be trained respectively to obtain single-task features and multi-task features; one of the sub-task branches of the multi-task model is the same as the task of the single-task model.

[0078] S302: If the task type of the single-task model is obstacle detection or lane detection, then based on the distillation auxiliary information corresponding to the task type, determine the distillation effective feature from the single-task feature and the multi-task feature.

[0079] Compared with space occupancy detection, obstacle detection and lane detection consider fewer feature dimensions, have lower feature richness, and include less comprehensive features. When the task type of a single-task model is obstacle detection or lane detection, directly using multi-task features and single-task features as effective features for distillation to determine distillation loss will result in feature loss.

[0080] Therefore, when the task type of the single-task model is obstacle detection or lane line detection, it is necessary to use the distillation auxiliary information corresponding to the task type to further process the single-task features or multi-task features and determine the distillation effective features in the single-task features and multi-task features.

[0081] Among them, the distillation auxiliary information is related to the task type. Different task types have different distillation auxiliary information used to determine the effective features of distillation.

[0082] S303: Determine the distillation loss between the single-task feature and the multi-task feature based on the distillation effective feature.

[0083] Optionally, the L1 loss between the effective features of the distillation is calculated, and the calculated L1 loss value is determined as the distillation loss between the single-task feature and the multi-task feature.

[0084] S304: Based on the distillation loss, update the parameters of the feature extraction network of the multi-task model to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

[0085] The disclosed technical solution introduces distillation auxiliary information corresponding to the task type when the task type of the single-task model is obstacle detection or lane line detection. The distillation loss feature is used to determine the distillation effective features in the single-task features and the multi-task features, thereby ensuring the accuracy of the distillation effective features. The distillation effective features are used to determine the distillation loss, and the distillation loss is used to update the parameters of the feature extraction network of the multi-task model, which is beneficial to improving the effect of feature distillation, improving the performance of the multi-task model, and enriching the applicable scenarios of the multi-task model training method.

[0086] In an optional embodiment, based on the distillation auxiliary information corresponding to the task type, the distillation effective feature is determined in the single-task feature and the multi-task feature, including: if the task type of the single-task model is obstacle detection, then based on the sample label of the perceived image sample, the position information of the true obstacle in the perceived image sample is determined; the position information of the true obstacle in the perceived image sample is used as the distillation auxiliary information; based on the distillation auxiliary information, a distillation effective area is determined in the single-task feature; and local features of the single-task feature and the multi-task feature in the distillation effective area are determined as the distillation effective features.

[0087] Obstacle detection mainly identifies and locates obstacles in the surrounding environment, and its output is usually a bounding box containing the obstacle. When the task type of the single-task model is obstacle detection, the sample label of the perceived image sample includes a bounding box for showing the location of the true obstacle in the perceived image sample.

[0088] Based on the sample labels of the perceived image samples, the position information of the true value obstacles in the perceived image samples is determined, and the position information of the true value obstacles in the perceived image samples is used as distillation auxiliary information.

[0089] Optionally, the feature area framed by the bounding box in the sample label in the single-task feature is determined as the distillation effective area. Only the local features in the single-task feature that are in the distillation effective area need to be transferred to the multi-task feature.

[0090] The local features of the single-task features and the multi-task features that are in the distillation effective area are determined as distillation effective features. Then, the distillation loss between the single-task features and the multi-task features is determined based on the distillation effective features.

[0091] Optionally, the distillation loss between single-task features and multi-task features is determined using the following formula:

[0092]

[0093] Among them, x represents a position in the single-task feature, loss(bev 1 (x),bev 2 (x),bbox) represents the distillation loss. bbox represents the feature effective area, bev 1 (x) refers to the single-task feature, bev 2 (x) refers to the multi-task feature. The above formula indicates that the L1 loss is calculated only for the local features of the single-task feature and the multi-task feature that are in the effective area of ​​distillation.

[0094] The above technical solution provides a feasible solution for determining effective distillation features when the task type of the single-task model is obstacle detection. It uses the position information of the true obstacle in the perceived image sample as distillation auxiliary information to assist in transferring the features related to obstacle detection in the single-task feature to the multi-task feature. This improves the accuracy of the effective distillation features, helps improve the feature distillation effect, and enriches the applicable scenarios of the multi-task model training method.

[0095] In an optional embodiment, the distillation effective feature is determined in the single-task feature and the multi-task feature based on the distillation auxiliary information corresponding to the task type, including: if the task type of the single-task model is lane detection, the multi-task feature is input into a nonlinear projection layer to obtain a projection layer feature; the projection layer feature is input into a lane detection decoder in the multi-task model to obtain a first detection result; the multi-task feature is input into a lane detection decoder in the multi-task model to obtain a second detection result; the task loss between the first detection result and the second detection result is determined, and the task loss is used as the distillation auxiliary information; based on the distillation auxiliary information, the nonlinear projection layer is parameterized to improve the correlation between the projection layer feature and the lane detection; the projection layer feature and the multi-task feature are determined as the distillation effective features.

[0096] Lane line detection Lane line detection mainly detects and locates lane lines in images or videos. Its output usually includes the location information, type information, and possible instance information of the lane line. When the task type of the single-task model is lane line detection, the sample label of the perceived image sample includes the precise coordinates of the lane line in the image, which is generally represented by line segments rather than regions. When the task type of the single-task model is obstacle detection, the distillation effective feature determination scheme used is no longer applicable.

[0097] Figure 3B This is a schematic diagram of a distillation effective feature determination scheme for the case where the task type of a single-task model is lane line detection. It is worth noting that Figure 3B The other sub-task branches included in the multi-task model are not shown in FIG, only the lane line detection branch is shown. Figure 3B , the multi-task features are input into the nonlinear projection layer to obtain the projection layer features. The projection layer features and the multi-task features have the same dimension, but the amount of information they contain is different. The nonlinear projection layer is used to transform the multi-task features in order to extract the projection layer features including lane line related information from the multi-task features including the full amount of information. Optionally, the nonlinear projection layer is a 1×1 convolutional network.

[0098] After obtaining the projection layer features, the projection layer features are input into the lane line detection decoder in the multi-task model to obtain a first detection result, wherein the accuracy of the first detection result is closely related to the accuracy of the projection layer features.

[0099] The multi-task features are input into the lane detection decoder in the multi-task model to obtain a second detection result, wherein the accuracy of the second detection result is closely related to the accuracy of the multi-task features.

[0100] Determine the task loss between the first detection result and the second detection result, use the task loss as distillation auxiliary information, use the distillation auxiliary information to update the parameters of the nonlinear projection layer, and assist the nonlinear projection layer to extract features related to the lane line, so as to improve the correlation between the projection layer features and the lane line detection.

[0101] After obtaining the projection layer features, the projection layer features and the single-task features are used as distillation effective features. The distillation loss between the single-task features and the multi-task features is calculated using the distillation effective features.

[0102] It is understandable that the projection layer features are learnable features. As the number of multi-model training increases and the parameters of the nonlinear projection layer are continuously updated, the correlation between the projection layer features and lane detection will gradually increase. While the nonlinear projection layer parameters are updated using the task loss, the feature extraction network parameters of the multi-task model are updated using the distillation loss determined based on the projection layer features and the single-task features. The feature extraction network and nonlinear projection layer of the multi-task model converge at the same time.

[0103] The above technical solution provides a feasible solution for determining effective features of distillation for the case where the task type of the single-task model is lane line detection. The task loss is determined by using the first detection result determined by the projection layer features and the second detection result determined by the multi-task features. The task loss is used as auxiliary information for distillation to assist the nonlinear projection layer in extracting features related to lane line information from the multi-task features that include the full amount of information. The features related to lane lines in the multi-task features can be closer to the single-task features extracted by the lane line detection model. The accuracy of the effective features of distillation is improved, which is conducive to improving the effect of feature distillation, and at the same time enriches the applicable scenarios of the multi-task model training method.

[0104] Figure 4 The present invention is a schematic diagram of a multi-task model training device provided according to an embodiment of the present invention. The present invention can be applied to the case of training a multi-task model for autonomous driving. The device can be implemented by software and / or hardware, and the device can implement the multi-task model training method described in any embodiment of the present invention.

[0105] like Figure 4 As shown, the multi-task model training device 400 includes:

[0106] A feature extraction module 401 is used to input the perceived image samples into the feature extraction networks of the pre-trained single-task model and the multi-task model to be trained, respectively, to obtain single-task features and multi-task features; one of the subtask branches of the multi-task model is the same as the task of the single-task model;

[0107] A distillation loss determination module 402, configured to determine a distillation loss between the single-task feature and the multi-task feature based on a task type of the single-task model;

[0108] The first parameter updating module 403 is used to update the parameters of the feature extraction network of the multi-task model based on the distillation loss to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

[0109] The disclosed technical solution uses a pre-trained single-task model to guide the training process of the multi-task model, uses the single-task features extracted by the feature extraction network in the single-task model as soft labels, and performs feature distillation on the multi-task features extracted by the feature extraction network in the multi-task model, thereby avoiding the situation where the feature extraction network of the multi-task model is difficult to optimize the most suitable multi-task features for each sub-task due to competition between tasks, and falls into a suboptimal solution, thereby solving the multi-task negative transfer problem and improving the performance of the multi-task model. In the process of feature distillation, the disclosed technical solution takes into account that the single-task features extracted from different tasks will have obvious task tendencies, and uses the task type of the single-task model to determine the distillation loss between the single-task features and the multi-task features, making full use of the prior knowledge of the single-task model, improving the effect of feature distillation, and making the multi-task features extracted by the feature extraction network of the multi-task model more in line with the sub-task requirements, which is conducive to improving the performance of the multi-task model.

[0110] Optionally, the distillation loss determination module 402 includes: an effective feature determination submodule, used to determine the distillation effective features among the multi-task features and the single-task features based on the task type of the single-task model; and a distillation loss determination submodule, used to determine the distillation loss between the single-task features and the multi-task features based on the distillation effective features.

[0111] Optionally, the effective feature determination submodule is specifically used to: if the task type of the single-task model is obstacle detection or lane line detection, then based on the distillation auxiliary information corresponding to the task type, determine the distillation effective feature from the single-task feature and the multi-task feature.

[0112] Optionally, the effective feature determination submodule includes: a position information determination unit, which is used to determine the position information of the true obstacle in the perceived image sample based on the sample label of the perceived image sample if the task type of the single-task model is obstacle detection; an auxiliary information determination unit, which is used to use the position information of the true obstacle in the perceived image sample as the distillation auxiliary information; an effective area determination unit, which is used to determine the distillation effective area in the single-task feature based on the distillation auxiliary information; and a first effective feature determination unit, which is used to determine the local features of the single-task feature and the multi-task feature in the distillation effective area as the distillation effective features.

[0113] Optionally, the effective feature determination submodule includes: a feature transformation unit, which is used to input the multi-task feature into a nonlinear projection layer to obtain a projection layer feature if the task type of the single-task model is lane line detection; a first detection result determination unit, which is used to input the projection layer feature into a lane line detection decoder in the multi-task model to obtain a first detection result; a second detection result determination unit, which is used to input the multi-task feature into a lane line detection decoder in the multi-task model to obtain a second detection result; an auxiliary information determination unit, which is used to determine the task loss between the first detection result and the second detection result, and use the task loss as the distillation auxiliary information; a parameter updating unit, which is used to update the parameters of the nonlinear projection layer based on the distillation auxiliary information to improve the correlation between the projection layer feature and the lane line detection; a second effective feature determination unit, which is used to determine the projection layer feature and the multi-task feature as the distillation effective feature.

[0114] Optionally, the effective feature determination submodule is further used to: if the task type of the single-task model is space occupancy detection, determine the single-task feature and the multi-task feature as the distillation effective feature.

[0115] Optionally, the device also includes: a task branch determination module, used to determine the subtask branch in the multi-task model that is the same as the single-task model task as the subtask branch of the target subtask; wherein the subtask branch includes a task decoder; a prediction result determination module, used to input the multi-task feature into the task decoder of the target subtask to obtain the prediction result of the target subtask; a true value result determination module, used to determine the true value result of the target subtask based on the sample label of the perceived image sample; a second parameter updating module, used to update the parameters of the subtask branch of the target subtask in the multi-task model based on the prediction result and the true value result.

[0116] Optionally, the device also includes: a perceptual image selection module, used to determine a target perceptual image that matches the task of the single-task model from candidate perceptual images according to the task type of the single-task model; and an image sample determination module, used to use the target perceptual image as a perceptual image sample used when training the multi-task model using the single-task model.

[0117] Optionally, the device is further used to: within a training cycle, use at least two of the single-task models to alternately train the multi-task model.

[0118] The multi-task model training device provided in the embodiments of the present disclosure can execute the multi-task model training method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the multi-task model training method.

[0119] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user data involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0120] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0121] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0122] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0123] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0124] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as a multi-task model training method. For example, in some embodiments, the multi-task model training method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the multi-task model training method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the multi-task model training method in any other appropriate manner (e.g., by means of firmware).

[0125] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0126] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable multi-task model training device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or completely on a remote machine or server.

[0127] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0129] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0130] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0131] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, knowledge graph technology, and other major directions.

[0132] Cloud computing refers to a technology system that uses network access to elastically scalable shared physical or virtual resource pools. Resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for technical applications such as artificial intelligence and blockchain, as well as model training.

[0133] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0134] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A multi-task model training method, the method comprising: Input the perceived image samples into the feature extraction networks of the pre-trained single-task model and the multi-task model to be trained respectively to obtain single-task features and multi-task features; One of the subtask branches of the multi-task model is the same as the task of the single-task model; Determining a distillation loss between the single-task feature and the multi-task feature based on a task type of the single-task model; Based on the distillation loss, parameters of a feature extraction network of the multi-task model are updated to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

2. The method according to claim 1, wherein: The determining the distillation loss between the single-task feature and the multi-task feature based on the task type of the single-task model includes: Based on the task type of the single-task model, determining distillation effective features from the multi-task features and the single-task features; Based on the distillation effective features, the distillation loss between the single-task features and the multi-task features is determined.

3. The method according to claim 2, wherein: Based on the task type of the single-task model, determining distillation effective features from the multi-task features and the single-task features includes: If the task type of the single-task model is obstacle detection or lane detection, the distillation effective feature is determined from the single-task feature and the multi-task feature based on the distillation auxiliary information corresponding to the task type.

4. The method according to claim 3, wherein: The determining the distillation effective feature from the single-task feature and the multi-task feature based on the distillation auxiliary information corresponding to the task type includes: If the task type of the single-task model is obstacle detection, determining position information of a true obstacle in the perceived image sample based on the sample label of the perceived image sample; Using the position information of the true obstacle in the perceived image sample as the distillation auxiliary information; Based on the distillation auxiliary information, determining a distillation effective area in the single-task feature; The local features of the single-task feature and the multi-task feature in the distillation effective area are determined as the distillation effective features.

5. The method according to claim 3, wherein: The determining the distillation effective feature from the single-task feature and the multi-task feature based on the distillation auxiliary information corresponding to the task type includes: If the task type of the single-task model is lane detection, inputting the multi-task features into a nonlinear projection layer to obtain projection layer features; Inputting the projection layer features into a lane detection decoder in the multi-task model to obtain a first detection result; Inputting the multi-task feature into a lane detection decoder in the multi-task model to obtain a second detection result; determining a task loss between the first detection result and the second detection result, and using the task loss as the distillation auxiliary information; Based on the distilled auxiliary information, updating parameters of the nonlinear projection layer to improve the correlation between the projection layer features and lane line detection; The projection layer features and the multi-task features are determined as the distillation effective features.

6. The method according to claim 2, wherein: The determining of distillation effective features from the multi-task features and the single-task features based on the task type of the single-task model further includes: If the task type of the single-task model is space occupancy detection, the single-task feature and the multi-task feature are determined as the distillation effective features.

7. The method according to claim 1, further comprising: Determine the subtask branch in the multi-task model that is the same as the task in the single-task model as the subtask branch of the target subtask; wherein the subtask branch includes a task decoder; Inputting the multi-task features into a task decoder of the target subtask to obtain a prediction result of the target subtask; Determining a true value result of the target subtask based on the sample labels of the perceived image samples; Based on the prediction result and the true value result, parameters of the subtask branch of the target subtask in the multi-task model are updated.

8. The method according to claim 1, further comprising: Determining, according to the task type of the single-task model, a target perception image matching the task of the single-task model from candidate perception images; The target perception image is used as a perception image sample used when the single-task model is used to train the multi-task model.

9. The method according to claim 1, further comprising: In one training cycle, at least two of the single-task models are used to alternately train the multi-task model.

10. A multi-task model training device, comprising: A feature extraction module is used to input the perceived image samples into the feature extraction networks in the pre-trained single-task model and the multi-task model to be trained, respectively, to obtain single-task features and multi-task features; One of the subtask branches of the multi-task model is the same as the task of the single-task model; A distillation loss determination module, used to determine the distillation loss between the single-task feature and the multi-task feature based on the task type of the single-task model; A first parameter updating module is used to update the parameters of the feature extraction network of the multi-task model based on the distillation loss to train the multi-task model; wherein the feature extraction network is shared by each sub-task branch in the multi-task model.

11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multi-task model training method described in any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the multi-task model training method according to any one of claims 1-9.

13. A computer program product, comprising a computer program, which, when executed by a processor, implements the multi-task model training method according to any one of claims 1-9.

Citation Information

Cited By

  • Knowledge distillation method and equipment for nuclear accident deduction and fault diagnosis of reactor

    CN120745746A