Multi-task model training method and device, electronic equipment and readable storage medium

By building a multi-task model, the problem of excessive computing resources in computer vision products is solved, and more efficient visual processing task execution is achieved.

CN120375128APending Publication Date: 2025-07-25北京数原数字化城市研究中心
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410030104.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Performing visual processing tasks in computer vision products will take up more computing resources.

Method used

Build a multi-task model, including the basic feature extraction module, feature fusion module and multiple task prediction modules, train these modules to perform different visual processing tasks, and calculate the target loss function value to determine the optimal model.

Benefits of technology

It reduces the use of computing resources, reduces the cost of landing the model, and improves the performance and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375128A_ABST
    Figure CN120375128A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task model training method and device, electronic equipment and a readable storage medium, and relates to the field of artificial intelligence. The method comprises the steps that a target image set is input into a multi-task model for training, and training comprises the step of executing at least two different visual processing tasks through a basic feature extraction module, a feature fusion module and at least two task prediction modules in n task prediction modules, calculating a target loss function value based on the loss function value corresponding to each task prediction module in the at least two task prediction modules; the target multi-task model is a multi-task model used for executing a plurality of different visual processing tasks, and the target loss function value is minimum. In this way, when the obtained target multi-task model processes different visual processing tasks, only one-time basic feature extraction and feature fusion need to be carried out, and compared with the prior art that basic feature extraction and feature fusion need to be carried out on each visual processing task, occupation of computing power resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence, and specifically relates to a training method, device, electronic device, and readable storage medium for a multi-task model. Background Art

[0002] With the continuous development and improvement of artificial intelligence technology, various products centered on artificial intelligence technology have gradually entered the stage of implementation.

[0003] When implementing a computer vision product, since a computer vision product usually involves multiple models, and each model in the computer vision product is responsible for completing a visual processing task. Thus, when a computer vision product involves a large number of visual processing tasks, the computer vision product will also contain a large number of models, and each model requires a certain amount of computing power resources. The more models there are in the computer vision product, the more computing power resources it will occupy.

[0004] In the process of implementing this application, the inventor found that there are at least the following problems in the prior art: Executing visual processing tasks occupies a large amount of computing power resources. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a training method for a multi-task model, which can solve the problem that executing visual processing tasks occupies a large amount of computing power resources.

[0006] To solve the above technical problems, this application is implemented as follows:

[0007] In a first aspect, the embodiments of this application provide a training method for a multi-task model, and the method includes:

[0008] Construct a multi-task model, where the multi-task model includes a basic feature extraction module, a feature fusion module, and n task prediction modules, n is a positive integer greater than 1, and the n task prediction modules are respectively used to execute n different visual processing tasks;

[0009] Input a target image set into the multi-task model for training. The training includes executing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating the target loss function value of the multi-task model. The target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules;

[0010] When the target loss function value of the multi-task model is the smallest, determine the multi-task model as the target multi-task model, and the target multi-task model is a model for executing multiple different visual processing tasks.

[0011] In a second aspect, an embodiment of the present application provides a training device for a multi-task model, including:

[0012] A construction module, configured to construct a multi-task model, where the multi-task model includes a basic feature extraction module, a feature fusion module, and n task prediction modules, n is a positive integer greater than 1, and the n task prediction models are respectively used to execute n different visual processing tasks;

[0013] A training module, configured to input a target image set into the multi-task model for training, where the training includes executing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating a target loss function value of the multi-task model, and the target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules;

[0014] A determination module, configured to determine the multi-task model as a target multi-task model when the target loss function value of the multi-task model is the smallest, and the target multi-task model is a model used to execute multiple different visual processing tasks.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0017] In the embodiment of the present application, by inputting a target image set into a multi-task model for training, the training includes executing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating a target loss function value based on the loss function value corresponding to each of the at least two task prediction modules; the target multi-task model is a multi-task model used to execute multiple different visual processing tasks and having the smallest target loss function value. The obtained target multi-task model only needs to perform basic feature extraction and feature fusion once when processing different visual processing tasks, compared with the prior art that needs to perform basic feature extraction and feature fusion for each visual processing task respectively, reducing the occupation of computing power resources. Description of the Drawings

[0018] Figure 1 It is a flowchart of the training method of the multi-task model in the embodiment of the present application;

[0019] Figure 2 It is a schematic diagram of the training process of a multi-task model in the embodiment of the present application;

[0020] Figure 3 It is a flowchart of image acquisition, image cleaning, and image annotation in the embodiment of the present application;

[0021] Figure 4 It is a schematic structural diagram of the training device of the multi-task model in the embodiment of the present application;

[0022] Figure 5 It is a schematic structural diagram of an electronic device further provided by the embodiment of the present application. Specific Embodiments

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0024] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such use can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0025] Next, in conjunction with the accompanying drawings, the training method of the multi-task model provided by the embodiment of the present application will be described in detail through specific embodiments and their application scenarios.

[0026] The training method of the multi-task model provided by the present application can be applied to the training device of the multi-task model, and the training device of the multi-task model is used to train the multi-task model. Among them, the multi-task model includes a basic feature extraction module, a feature fusion module, and n task prediction modules, where n is a positive integer greater than 1.

[0027] For the purpose of clearly explaining the technical solution, hereinafter, taking the training of 3 task prediction modules by the above-mentioned training device of the multi-task model as an example, the technical solution will be described.

[0028] When the training device of the multi-task model trains three task prediction modules, exemplarily, the training device of the multi-task model can train object detection, semantic segmentation, and depth estimation.

[0029] When the training device of the multi-task model trains four task prediction modules, exemplarily, the training device of the multi-task model can train object detection, semantic segmentation, depth estimation, and object recognition.

[0030] It should be understood that object detection refers to detecting specific objects in an image and obtaining the position and size information of the objects; semantic segmentation refers to classifying different parts of an image and can effectively distinguish different objects in the image; depth estimation refers to performing 3D reconstruction on the objects in the image to obtain their accurate three-dimensional structure in the real world; object recognition refers to accurately classifying the objects in the image, such as dogs, cats, etc.

[0031] Specifically, for a training method of a multi-task model provided in an embodiment of the present application, please refer to Figure 1 , Figure 1 which is a flowchart of the training method of the multi-task model in the embodiment of the present application, and includes the following steps:

[0032] S101, construct a multi-task model, the multi-task model includes a basic feature extraction module, a feature fusion module, and n task prediction modules, where n is a positive integer greater than 1, and the n task prediction modules are respectively used to execute n different visual processing tasks;

[0033] In S101, the above multi-task model is used to execute multiple visual processing tasks.

[0034] It should be understood that the above basic feature extraction module is used to extract the feature vector of the image; the above feature fusion vector is used to obtain the feature fusion vector of the above feature vector; the task prediction module executes the visual processing task according to the above feature fusion vector.

[0035] Further, the output of the above basic feature extraction module corresponds to the input of the above feature fusion module, and the output of the above feature fusion module corresponds to the input of the above n task prediction modules.

[0036] It should be understood that the n visual processing tasks may include but are not limited to at least one of the following: object detection, semantic segmentation, depth estimation, and object recognition; no further limitation is made here.

[0037] Exemplarily, when n is equal to 3, the n visual processing tasks may include: object detection, semantic segmentation, and depth estimation; further, the n task prediction modules are 3 task prediction modules, where the first task prediction module is used to perform the above-mentioned object detection visual processing task, the second task prediction module is used to perform the above-mentioned semantic segmentation visual processing task, and the third task prediction module is used to perform the above-mentioned depth estimation visual processing task.

[0038] When n is equal to 4, the n visual processing tasks may include: object detection, semantic segmentation, depth estimation, and object recognition; further, the n task prediction modules are 4 task prediction modules, where the first task prediction module is used to perform the above-mentioned object detection visual processing task, the second task prediction module is used to perform the above-mentioned semantic segmentation visual processing task, the third task prediction module is used to perform the above-mentioned depth estimation visual processing task, and the fourth task prediction module is used to perform the above-mentioned object recognition visual processing task.

[0039] S102. Input the target image set into the multi-task model for training. The training includes performing at least two different visual processing tasks through the basic feature extraction module, the feature fusion module, and at least two task prediction modules among the n task prediction modules, and calculating the target loss function value of the multi-task model. The target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules.

[0040] In S102, a specific implementation manner of inputting the target image set into the multi-task model for training may be to configure the above target image set into the configuration file of the above multi-task model.

[0041] Optionally, the above configuration file can also be used to adjust and configure parameters such as the parameter (learning rate) for controlling the update speed of the control parameter, the parameter (max_epoch) for controlling the total number of times of model training, the parameter (weight_decay) for controlling the fitting degree of the model, the parameter (momentu) for controlling the acceleration during the parameter update process, and the parameter (dropout) for controlling the complexity of the neural network.

[0042] It should be understood that the implementation manner in which the above target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules includes:

[0043] A specific implementation manner is: adding the loss function values corresponding to each of the at least two task prediction modules to obtain the above target loss function value.

[0044] Another specific implementation manner is as follows: First, according to the importance of the visual processing task corresponding to each of the at least two task prediction modules, determine the weight value of each of the at least two task prediction modules, and then add the products of the weight value of each task prediction module and the loss function value corresponding to each task prediction module to obtain the above-mentioned target loss function value.

[0045] S103. When the target loss function value of the multi-task model is the smallest, determine the multi-task model as the target multi-task model, where the target multi-task model is a model for performing multiple different visual processing tasks.

[0046] It should be understood that the target loss function value of the multi-task model is the smallest, that is, when verifying the multi-task model twice in succession, the target loss function value of the multi-task model during the second verification is greater than or equal to the target loss function value of the multi-task model during the first verification. At this time, the target loss function value of the multi-task model is the smallest, and the multi-task model at this time is the above-mentioned target multi-task model.

[0047] In S103, when the target loss function value of the multi-task model is the smallest, determine the multi-task model as the target multi-task model, where the target multi-task model is a model for performing multiple different visual processing tasks.

[0048] When n is equal to 3, please refer to Figure 2 , Figure 2 is a schematic diagram of the training process of a multi-task model provided by an embodiment of the present application. Figure 2 In, the multi-task model includes 3 task prediction modules respectively used for object detection, semantic segmentation, and depth estimation.

[0049] Refer to Figure 2 for description. The images in the target image set are subjected to feature extraction by the basic feature extraction module 201 to obtain feature vectors; the feature vectors are input into the feature fusion module 202 for feature fusion to obtain feature fusion vectors; and then the feature fusion vectors are respectively input into the 3 task prediction modules, namely the first task prediction module 203, the second task prediction module 204, and the third task prediction module 205. The first task prediction module 203 is used for object detection, the second task prediction module 204 is used for semantic segmentation, and the third task prediction module 205 is used for depth estimation.

[0050] Among them, the feature fusion module 202 includes a Spatial Pyramid Pooling–Fast (SPPF) sub-module 2021.

[0051] In an embodiment of the present application, by inputting a target image set into a multi-task model for training, the training includes performing at least two different visual processing tasks through at least two of a basic feature extraction module, a feature fusion module, and n task prediction modules, and calculating a target loss function value based on the loss function value corresponding to each of the at least two task prediction modules; the target multi-task model is a multi-task model for performing multiple different visual processing tasks and having the minimum target loss function value. The obtained target multi-task model only needs to perform basic feature extraction and feature fusion once when processing different visual processing tasks. Compared with the prior art, which needs to perform basic feature extraction and feature fusion for each visual processing task separately, the occupation of computing power resources is reduced. At the same time, the implementation cost of the model is also reduced, and the performance and efficiency of the model are improved.

[0052] Optionally, in some embodiments, the inputting the target image set into the multi-task model for training includes:

[0053] Determining at least two activated task prediction modules in the multi-task model;

[0054] Inputting the target image set into the multi-task model, extracting feature vectors from the images in the target image set through the basic feature extraction module, and performing feature fusion on the feature vectors extracted from the images through the feature fusion module to obtain a feature fusion vector, performing at least two different visual processing tasks on the feature fusion vector based on the at least two activated task prediction modules, and calculating the target loss function value of the multi-task model.

[0055] It should be understood that the above-mentioned activated task prediction modules refer to the task prediction modules in the multi-task model activated by the user according to the visual processing tasks to be performed through a configuration file.

[0056] Exemplarily, the visual processing tasks to be performed include object detection and semantic segmentation. Therefore, the user needs to set the two visual processing tasks of object detection and semantic segmentation to true through the configuration file, so as to obtain two activated task prediction modules corresponding to the two visual processing tasks of object detection and semantic segmentation.

[0057] Furthermore, extracting feature vectors from the images in the target image set through the basic feature extraction module, and performing feature fusion on the feature vectors extracted from the images through the feature fusion module to obtain a feature fusion vector, performing object detection and semantic segmentation on the feature fusion vector based on the two activated task prediction modules corresponding to object detection and semantic segmentation, and calculating the target loss function value of the multi-task model.

[0058] In the embodiments of the present application, through the above manner, according to the visual processing tasks to be executed, the task prediction modules corresponding to the visual processing tasks to be executed in the multitask model can be correspondingly activated. Thereby, the execution of the visual processing tasks corresponding to the unactivated task prediction modules in the multitask model is avoided, and further, the occupation amount of computing power resources by the multitask model is reduced.

[0059] Optionally, in some embodiments, the performing at least two different visual processing tasks on the feature fusion vector based on the at least two activated task prediction modules and calculating the target loss function value of the multitask model includes:

[0060] Performing the i-th visual processing task on the feature fusion vector based on the i-th activated task prediction module among the at least two activated task prediction modules, and calculating the loss function value corresponding to the i-th activated task prediction module, where i is a positive integer;

[0061] When the loss function value corresponding to the i-th activated task prediction module is the smallest, freezing the i-th activated task prediction module;

[0062] In the case where the i-th activated task prediction module is frozen, incrementing i by 1, and looping to execute the step of performing the i-th visual processing task on the feature fusion vector based on the i-th activated task prediction module among the at least two activated task prediction modules and calculating the loss function value corresponding to the i-th activated task prediction module until the at least two activated task prediction modules are trained.

[0063] It should be understood that freezing the i-th activated task prediction module means that this module no longer undergoes training at this time.

[0064] It should be noted that the performing at least two different visual processing tasks on the feature fusion vector based on the at least two activated task prediction modules and calculating the target loss function value of the multitask model includes:

[0065] First, performing the visual processing task corresponding to the first activated task prediction module among the at least two activated task prediction modules on the feature fusion vector, and calculating the loss function value of the first activated task prediction module after the visual processing task corresponding to the first activated task prediction module is completed;

[0066] Secondly, with the first activated task prediction module frozen, the second activated task prediction module among the at least two activated task prediction modules performs the corresponding visual processing task of the second activated task prediction module on the feature fusion vector, and calculates the loss function value of the second activated task prediction module after the corresponding visual processing task of the second activated task prediction module is completed;

[0067] And so on, until all of the at least two activated task prediction modules are trained.

[0068] Finally, according to the loss function values of each activated task prediction module among the at least two activated task prediction modules, the target loss function value of the multi-task model is calculated.

[0069] Exemplarily, the at least two activated task prediction modules include: the activated task prediction modules corresponding to the three visual processing tasks of object detection, semantic segmentation, and depth estimation.

[0070] First, the object detection corresponding to the activated task prediction module of object detection can be performed on the feature fusion vector, and the loss function value of the activated task prediction module corresponding to object detection can be calculated; then, with the activated task prediction module corresponding to object detection frozen, the semantic segmentation corresponding to the activated task prediction module of semantic segmentation can be performed on the feature fusion vector, and the loss function value of the activated task prediction module corresponding to semantic segmentation can be calculated; then, with the activated task prediction modules corresponding to object detection and semantic segmentation frozen, the depth estimation corresponding to the activated task prediction module of depth estimation can be performed on the feature fusion vector, and the loss function value of the activated task prediction module corresponding to depth estimation can be calculated; finally, based on the loss function values of the activated task prediction modules corresponding to object detection, semantic segmentation, and depth estimation respectively, the target loss function value of the multi-task model is calculated.

[0071] In the embodiments of the present application, through the above method, the activated task prediction modules in the multi-task model can be trained one by one, improving the training efficiency of the multi-task model.

[0072] Optionally, in some embodiments, before inputting the target image set into the multi-task model for training, it further includes:

[0073] Collecting a first image set, the first image set including at least one first image;

[0074] When the first image set includes at least two first images, image cleaning is performed on the first image set to obtain a second image set, and the second image set includes at least one second image;

[0075] When the second image set includes at least two second images, according to the activated task prediction module, each second image in the second image set is labeled to obtain a target image set.

[0076] In an embodiment of the present application, a specific implementation manner of image acquisition, image cleaning, and image annotation is shown in Figure 3 , Figure 3 which is a flowchart of image acquisition, image cleaning, and image annotation provided by an embodiment of the present application, including the following steps:

[0077] First, image acquisition is performed. The acquired images are subjected to image cleaning, and then it is determined whether the image cleaning is qualified. If the image cleaning is not qualified, return to continue image cleaning. If the image cleaning is qualified, enter the next step of image annotation, and then determine whether the image annotation is qualified. If the image annotation is not qualified, return to continue image annotation. If the image annotation is qualified, enter the end.

[0078] It should be understood that the standard for determining whether the image cleaning is qualified may be that the proportion of invalid images such as redundant images, blurred images, and duplicate images in the total number of images is less than a preset value; the standard for determining whether the image annotation is qualified may be: whether it conforms to the target category and annotation criteria of the visual processing task to be executed.

[0079] In an embodiment of the present application, the first image set can be obtained by taking pictures at preset intervals, or by extracting video frames from a recorded video at preset intervals.

[0080] Optionally, in some embodiments, the first image set can also be obtained according to the visual processing task to be executed.

[0081] For example, in one embodiment, when only the visual processing task of face target detection needs to be executed, images containing more faces can be collected as much as possible.

[0082] For example, in another embodiment, when two visual processing tasks of face target detection and human semantic segmentation need to be executed, images containing more faces and more humans can be collected as much as possible.

[0083] Optionally, in some embodiments, the first images included in the first image set should cover different scenarios, different time periods, different climates, and different seasons as much as possible.

[0084] It should be understood that the above image cleaning refers to removing invalid data such as redundant images, blurred images, and duplicate images, and at the same time, images containing more visual processing tasks and more targets should be selected as much as possible.

[0085] It should be noted that according to the activated task prediction module, each of the second images in the second image set is labeled to obtain a target image set; according to the activated task prediction module, the corresponding visual processing task is determined, and the target category and labeling criterion can be known according to the corresponding visual processing task, and labeling is performed according to the target category and labeling criterion.

[0086] In the embodiment of the present application, the target image set obtained by the above method has the following advantages: high diversity, conciseness and high quality of image data; and by training the multi-task model with the target image set, the generalization performance and accuracy of the multi-task model can be improved.

[0087] Optionally, in some embodiments, the inputting the target image set into the multi-task model for training includes:

[0088] Performing a Shuffle process on the target image set to obtain a third image set;

[0089] Inputting the third image set into the multi-task model for training.

[0090] It should be understood that the main function of the Shuffle process is to randomly disrupt the arrangement order of the images in the target image set. The process of the Shuffle process includes reordering, rearranging, and reclassifying the target image set, etc., to ensure that the effect of each run is completely random.

[0091] It should be noted that the training process of inputting the third image set into the multi-task model for training includes: performing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating the target loss function value of the multi-task model, and the target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules.

[0092] In the embodiment of the present application, by the above method, since the third image set is obtained after randomly disrupting the target image set, the generalization performance and robustness of the multi-task model trained by the third image set are improved.

[0093] Optionally, the inputting the target image set into the multi-task model for training includes:

[0094] Perform data augmentation on the target image set to obtain a fourth image set. The data augmentation includes common task augmentation and independent task augmentation. The common task augmentation is to perform data augmentation on all visual processing tasks, and the independent task augmentation is to perform data augmentation on a single visual processing task;

[0095] Input the fourth image set into the multi-task model for training.

[0096] It should be understood that the common task augmentation includes but is not limited to at least one of the following: color model augmentation (HSV Augment) and flip augmentation (Flip Augment). HSV Augment refers to making the images in the target image set look more vivid by changing the color model in the target image set, that is, hue, saturation, and brightness; while Flip Augment refers to horizontally flipping, vertically flipping, or rotating the images in the target image set by a certain angle to enhance the images, so as to obtain a fourth image set with richer image data.

[0097] It should be understood that the independent task augmentation includes but is not limited to at least one of the following: Mixup and Mosaic. The above Mixup is a data augmentation method that mixes images with different annotations in a random mixing form to obtain a more diverse and rich fourth image set; the above Mosaic refers to splitting the images in the target image set into different small pieces and then reassembling them together, so as to transform the images in the target image set into a fourth image set with more rich details.

[0098] It should be noted that the training process of inputting the fourth image set into the multi-task model for training includes: performing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating the target loss function value of the multi-task model. The target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules.

[0099] In the embodiments of the present application, through the above method, since the fourth image set has more image data than the target image set, the multi-task model trained by the fourth image set can improve the detection accuracy of the above multi-task model.

[0100] Optionally, in some embodiments, the target image set is divided into a test set and a validation set. The feature extraction of the images in the target image set by the basic feature extraction module includes:

[0101] Feature extraction is performed on the images in the test set of the target image set by the basic feature extraction module;

[0102] Calculating the target loss function value of the multi-task model includes:

[0103] Calculating the loss function value between the processing results of the at least two activated task prediction modules performing visual processing tasks and the images corresponding to the validation set.

[0104] In the embodiments of the present application, by dividing the target images into a test set and a validation set respectively, feature extraction is performed on the images in the test set of the target image set by the basic feature extraction module; and calculating the loss function value between the processing results of the at least two activated task prediction modules performing visual processing tasks and the images corresponding to the validation set. The detection accuracy of the above multi-task model can be improved.

[0105] Optionally, in some embodiments, calculating the loss function value between the processing results of the at least two activated task prediction modules performing visual processing tasks and the images corresponding to the validation set includes:

[0106] Calculating the loss function value between the processing results of the at least two activated task prediction modules performing visual processing tasks and the images corresponding to the validation set according to the activated task prediction modules and a preset validation period.

[0107] It should be noted that the loss function value between the processing results of the at least two activated task prediction modules performing visual processing tasks and the images corresponding to the validation set will gradually decrease as the training progresses, and the decrease rates of the corresponding loss function values of different activated task prediction modules are also different.

[0108] It should be noted that the decrease rate of the loss function value of the activated task prediction module corresponding to the visual processing task of object detection during training is faster than the decrease rate of the loss function value of the activated task prediction module corresponding to the visual processing task of semantic segmentation during training.

[0109] Exemplarily, when the preset validation period is 1 minute; calculate the loss function value between the visual processing task of object detection and the images corresponding to the validation set every minute; calculate the loss function value between the visual processing task of semantic segmentation and the images corresponding to the validation set every two minutes.

[0110] In another embodiment, since the change process of the decline rate of the loss function value of the activated task prediction module corresponding to the visual processing task of object detection is fast first and then slow during the training process; therefore, the loss function value between the visual processing task of object detection and the image corresponding to the validation set can be calculated every one minute in the early stage of training, and in the later stage of training, the loss function value between the visual processing task of object detection and the image corresponding to the validation set can be calculated every two minutes.

[0111] In the embodiments of the present application, through the above method, the verification steps of the multi-task model can be simplified, and the training speed of the multi-task module can be accelerated.

[0112] It should be noted that for the training method of the multi-task model provided in the embodiments of the present application, the execution subject can be the training device 300 of the multi-task model, or the control module in the training device 300 of the multi-task model for executing the training method of loading the multi-task model. In the embodiments of the present application, the training method of loading the multi-task model is executed by the training device 300 of the multi-task model as an example to illustrate the training method of the multi-task model provided in the embodiments of the present application.

[0113] The embodiments of the present application provide a training device 300 for a multi-task model. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the training device 300 for the multi-task model in the embodiments of the present application, including:

[0114] A construction module 301, configured to construct a multi-task model, where the multi-task model includes a basic feature extraction module, a feature fusion module, and n task prediction modules, n is a positive integer greater than 1, and the n task prediction models are respectively used to execute n different visual processing tasks;

[0115] A training module 302, configured to input a target image set into the multi-task model for training, where the training includes executing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating the target loss function value of the multi-task model, where the target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules;

[0116] A determination module 303, configured to determine the multi-task model as a target multi-task model when the target loss function value of the multi-task model is the smallest, where the target multi-task model is a model for executing multiple different visual processing tasks.

[0117] Optionally, in some embodiments, the training module 302 includes:

[0118] A determination sub-module, configured to determine at least two activated task prediction modules in the multi-task model;

[0119] A first input sub-module, configured to input the target image set into the multi-task model, extract features of the images in the target image set through the basic feature extraction module to obtain feature vectors, and perform feature fusion on the feature vectors obtained by extracting features of the images through the feature fusion module to obtain feature fusion vectors, perform at least two different visual processing tasks on the feature fusion vectors based on the at least two activated task prediction modules, and calculate the target loss function value of the multi-task model.

[0120] Optionally, in some embodiments, the first input sub-module is further configured to:

[0121] Perform the i-th visual processing task on the feature fusion vector based on the i-th activated task prediction module among the at least two activated task prediction modules, and calculate the loss function value corresponding to the i-th activated task prediction module, where i is a positive integer;

[0122] When the loss function value corresponding to the i-th activated task prediction module is the smallest, freeze the i-th activated task prediction module;

[0123] When the i-th activated task prediction module is frozen, increment i by 1, and loop to perform the step of performing the i-th visual processing task on the feature fusion vector based on the i-th activated task prediction module among the at least two activated task prediction modules and calculating the loss function value corresponding to the i-th activated task prediction module until the at least two activated task prediction modules are trained.

[0124] Optionally, in some embodiments, the training device 300 of the multi-task model further includes:

[0125] An acquisition module, configured to acquire a first image set, where the first image set includes at least one first image;

[0126] An image cleaning module, configured to perform image cleaning on the first image set to obtain a second image set including at least one second image when the first image set includes at least two first images;

[0127] A labeling module, configured to label each second image in the second image set according to the activated task prediction module to obtain the target image set when the second image set includes at least two second images.

[0128] Optionally, in some embodiments, the input module includes:

[0129] A shuffling sub-module for performing a Shuffle process on the target image set to obtain a third image set;

[0130] A second input sub-module for inputting the third image set into the multi-task model for training.

[0131] Optionally, in some embodiments, the target image set is divided into a test set and a validation set, and the first input sub-module is further configured to:

[0132] Extract features from the images in the test set of the target image set through the basic feature extraction module;

[0133] Calculating the target loss function value of the multi-task model includes:

[0134] Calculating the loss function value between the processing results of the at least two activated task prediction modules performing visual processing tasks and the images corresponding to the validation set.

[0135] The training device 300 of the multi-task model in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. This device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0136] The training device 300 of the multi-task model in the embodiments of the present application can be a device with an operating system. This operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0137] The training device 300 of the multi-task model provided by the embodiments of the present application can implement Figure 1 Each process implemented by the training device 300 of the multi-task model in the method embodiments. To avoid repetition, it will not be elaborated here.

[0138] Figure 5 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0139] The electronic device 1000 includes, but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010 and other components.

[0140] Those skilled in the art can understand that the electronic device 1000 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 1010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 5 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0141] Among them, the processor 1010 is used to construct a multi-task model, and the multi-task model includes a basic feature extraction module, a feature fusion module, and n task prediction modules, where n is a positive integer greater than 1, and the n task prediction modules are respectively used to execute n different visual processing tasks;

[0142] The input unit 1004 is used to input a target image set into the multi-task model for training. The training includes performing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating the target loss function value of the multi-task model. The target loss function value is determined based on the loss function value corresponding to each of the at least two task prediction modules;

[0143] The processor 1010 is used to determine the multi-task model as a target multi-task model when the target loss function value of the multi-task model is the smallest. The target multi-task model is a model for performing multiple different visual processing tasks.

[0144] In the embodiments of the present application, by inputting a target image set into a multi-task model for training, the training includes performing at least two different visual processing tasks through at least two of a basic feature extraction module, a feature fusion module, and n task prediction modules, and calculating a target loss function value based on the loss function value corresponding to each task prediction module in the at least two task prediction modules; the target multi-task model is a multi-task model for performing multiple different visual processing tasks and having the minimum target loss function value. The obtained target multi-task model only needs to perform basic feature extraction and feature fusion once when processing different visual processing tasks. Compared with the prior art that needs to perform basic feature extraction and feature fusion for each visual processing task respectively, it reduces the occupation of computing power resources. At the same time, it also reduces the implementation of the model.

[0145] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned training method embodiment of the multi-task model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0146] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc.

[0147] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0149] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A training method for a multi-task model, characterized in that, The method includes: Constructing a multi-task model, the multi-task model including a basic feature extraction module, a feature fusion module, and n task prediction modules, where n is a positive integer greater than 1, and the n task prediction modules are respectively used to perform n different visual processing tasks; Inputting a target image set into the multi-task model for training, the training including performing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating a target loss function value of the multi-task model, the target loss function value being determined based on the loss function values corresponding to each of the at least two task prediction modules; When the target loss function value of the multi-task model is minimized, determining the multi-task model as a target multi-task model, the target multi-task model being a model for performing multiple different visual processing tasks.

2. The training method of the multi-task model according to claim 1, wherein The inputting the target image set into the multi-task model for training includes: Determining at least two activated task prediction modules in the multi-task model; Inputting the target image set into the multi-task model, extracting feature vectors from the images in the target image set through the basic feature extraction module, and performing feature fusion on the feature vectors extracted from the images through the feature fusion module to obtain a feature fusion vector, performing at least two different visual processing tasks on the feature fusion vector based on the at least two activated task prediction modules, and calculating the target loss function value of the multi-task model.

3. The training method of the multi-task model according to claim 2, characterized in that The performing at least two different visual processing tasks on the feature fusion vector based on the at least two activated task prediction modules and calculating the target loss function value of the multi-task model includes: Performing a first visual processing task on the feature fusion vector based on the i-th activated task prediction module among the at least two activated task prediction modules, and calculating the loss function value corresponding to the i-th activated task prediction module, where i is a positive integer; When the loss function value corresponding to the i-th activated task prediction module is minimized, freezing the i-th activated task prediction module; In the case where the i-th activated task prediction module is frozen, incrementing i by 1, and repeatedly performing the step of performing a first visual processing task on the feature fusion vector based on the i-th activated task prediction module among the at least two activated task prediction modules and calculating the loss function value corresponding to the i-th activated task prediction module until the at least two activated task prediction modules are trained.

4. The training method of the multi-task model according to claim 2, wherein Before inputting the target image set into the multi-task model for training, it further includes: Collecting a first image set, the first image set including at least one first image; In the case where the first image set includes at least two first images, performing image cleaning on the first image set to obtain a second image set, the second image set including at least one second image; In the case where the second image set includes at least two second images, according to the activated task prediction module, each of the second images in the second image set is labeled to obtain the target image set.

5. The training method of the multi-task model according to claim 1, characterized in that The inputting the target image set into the multi-task model for training includes: Performing a Shuffle process on the target image set to obtain a third image set; Inputting the third image set into the multi-task model for training.

6. The training method of the multi-task model according to claim 2, characterized in that The target image set is divided into a test set and a validation set. The feature extraction of the images in the target image set by the basic feature extraction module includes: Performing feature extraction on the images in the test set of the target image set by the basic feature extraction module; The calculation of the target loss function value of the multi-task model includes: Calculating the loss function value between the processing results of the at least two activated task prediction modules performing visual processing tasks and the images corresponding to the validation set.

7. A training device for a multi-task model, characterized in that, The device includes: A construction module for constructing a multi-task model, the multi-task model including a basic feature extraction module, a feature fusion module, and n task prediction modules, where n is a positive integer greater than 1, and the n task prediction models are respectively used to perform n different visual processing tasks; A training module for inputting a target image set into the multi-task model for training, the training including performing at least two different visual processing tasks through at least two of the basic feature extraction module, the feature fusion module, and the n task prediction modules, and calculating the target loss function value of the multi-task model, the target loss function value being determined based on the loss function value corresponding to each of the at least two task prediction modules; A determination module for determining the multi-task model as the target multi-task model when the target loss function value of the multi-task model is the smallest, the target multi-task model being a model for performing multiple different visual processing tasks.

8. The training device for the multi-task model according to claim 7, wherein The training module includes: A determination sub-module for determining at least two activated task prediction modules in the multi-task model; An input sub-module for inputting the target image set into the multi-task model, performing feature extraction on the images in the target image set through the basic feature extraction module to obtain feature vectors, and performing feature fusion on the feature vectors obtained by the feature extraction of the images through the feature fusion module to obtain feature fusion vectors, performing at least two different visual processing tasks on the feature fusion vectors based on the at least two activated task prediction modules, and calculating the target loss function value of the multi-task model.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the training method of the multi-task model according to any one of claims 1-6 are implemented.

10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the training method of the multi-task model according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Multi-task target detection method and device, automatic driving system and storage medium

    CN114821269A

  • Target detection method and device, equipment and storage medium

    CN115761698A

  • Model training method and system and storage medium

    CN115860087A

  • Training method of multi-task combination model based on small samples

    CN116070119A

  • Multi-task model training method and device, task prediction method and device, equipment and medium

    CN116758511A