Multi-task model training method, mobile inspection method, equipment and medium

Through the multi-task model training method, the main and side tasks are merged into a unified model, which solves the problems of diverse algorithm scenarios and scarce data in mobile inspection, and improves the accuracy of side tasks and the computing efficiency of edge devices.

CN120236185APending Publication Date: 2025-07-01HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311871492.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

There are problems in mobile inspections with diverse algorithm scenarios and scarce data, which makes algorithm tasks difficult to train and low computing efficiency when deployed on edge devices.

Method used

The multi-task model training method is adopted, tasks with more training data are used as the main task and tasks with less data are used as the side tasks. By learning the main and side tasks in stages, it is merged into a unified model for prediction, and the main task ability is used to improve the accuracy of the side tasks.

Benefits of technology

Improves the accuracy of the few-sample side tasks, reduces training difficulty, and improves computing efficiency on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236185A_ABST
    Figure CN120236185A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task model training method, a mobile inspection method, equipment and a medium, and the training method comprises the steps: obtaining a first data set containing a main line task label, and carrying out the training of a pre-training model through the first data set, and obtaining a main line task model; obtaining a second data set containing a branch task label, and adding the branch task label for the first data set according to the second data set and a pre-training model to obtain an updated first data set; the training data of the first data set is more than the training data of the second data set; adding a main line task label for the second data set according to the main line task model to obtain an updated second data set; and training the main line task model by using the updated first data set and the updated second data set to obtain a multi-task model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly relates to a multi-task model training method, a mobile inspection method, a device, a device and a medium. Background Art

[0002] Mobile inspection mainly faces the urban governance scenario, and realizes the automatic discovery of illegal and irregular events through mobile vehicles, providing basic perception capabilities for the intelligent upgrade of urban governance. The algorithm scenarios faced by mobile inspection are diverse and complex. Compared with consumer scenarios or visual detection tasks in specific fields, the scenario customization degree of mobile inspection is higher. In addition to common algorithm scenarios, such as personnel and vehicle detection, it also involves algorithm scenarios related to government affairs, with high data sensitivity and scarce quantity, resulting in great difficulty in algorithm task training.

[0003] Currently, in the face of complex and diverse algorithm scenarios, the conventional solution is to rely on supervised learning to independently form multiple algorithm task models to complete detections respectively. However, due to the problem of scarce data in some special algorithm scenarios, its accuracy is low, and mobile inspection generally needs to deploy algorithm task models on edge devices, while the computing power of edge devices is relatively lower than that of common server hosts. Running multiple algorithm tasks at the same time will be limited by computing power, seriously affecting the computing efficiency of complex and diverse algorithm scenarios. Summary of the Invention

[0004] The purpose of this application is to propose a multi-task model training method, a mobile inspection method, a device and a medium for the above-mentioned deficiencies of the prior art, and this purpose is achieved through the following technical solutions.

[0005] The first aspect of this application proposes a multi-task model training method, and the method includes:

[0006] Obtain a first data set containing main task annotations, and use the first data set to train a pre-trained model to obtain a main task model;

[0007] Obtain a second data set containing branch task annotations, and add branch task annotations to the first data set according to the second data set and the pre-trained model to obtain an updated first data set; the training data of the first data set is more than the training data of the second data set;

[0008] Add main task annotations to the second data set according to the main task model to obtain an updated second data set;

[0009] Use the updated first data set and the updated second data set to train the main task model to obtain a multi-task model.

[0010] Based on the multi-task model training method described in the above first aspect, it has at least the following beneficial effects or advantages:

[0011] By using a combined model for tasks with more training data and tasks with less training data in the actual application scenario instead of training multiple models separately, the number of models is reduced. That is, the task with more training data is used as the main task, and the task with less training data is used as the secondary task. The model is made to learn the main task and the secondary task in stages. That is, first, the pre-trained model is made to learn the main task using the first dataset with main task annotations to obtain the main task model. Then, the second dataset with secondary task annotations and the pre-trained model are used to add secondary task annotations to the first dataset, and the main task model is used to add main task annotations to the second dataset. Furthermore, the main task model is trained using the first dataset and the second dataset that have both secondary task annotations and main task annotations, so that the model can utilize the main task ability when learning the secondary task and keep the main task ability from degrading. As a result, the accuracy of the few-shot secondary tasks of the finally trained multi-task model is effectively improved, and the training difficulty of the few-shot tasks is reduced.

[0012] The second aspect of this application proposes a mobile inspection method, and the method includes:

[0013] Obtain inspection images collected by a mobile vehicle;

[0014] Input the inspection images into the multi-task model obtained from the above first aspect;

[0015] Obtain the main task prediction results and secondary task prediction results output by the multi-task model.

[0016] Based on the mobile inspection method described in the above second aspect, it has at least the following beneficial effects or advantages:

[0017] Since the multi-task model combines the algorithm tasks required for mobile inspection into a unified model for prediction and can also ensure the accuracy of data-scarce tasks, by inputting the inspection images into the unified multi-task model, the prediction results of all algorithm tasks, that is, the main task prediction results and secondary task prediction results, can be obtained. In this way, when the multi-task model is deployed to the edge device for mobile inspection, since the single model consumes less computing power, the computing efficiency of mobile inspection under the edge device can be effectively improved.

[0018] The third aspect of this application proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the method described in the above first aspect or second aspect.

[0019] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method described in the first aspect or the second aspect above.

[0020] The above description is only an overview of the technical solution of the present application. In order to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specific embodiments of the present application are given. Description of the Drawings

[0021] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0022] Figure 1 It is a flowchart of an embodiment of a multi-task model training method shown according to an exemplary embodiment;

[0023] Figure 2 It is a classification schematic diagram of an algorithm task shown according to an exemplary embodiment;

[0024] Figure 3 It is a flowchart of the implementation of a main task model shown according to an exemplary embodiment;

[0025] Figure 4 It is a flowchart of implementing a multi-task model based on a main task model shown according to an exemplary embodiment;

[0026] Figure 5 It is a flowchart of an embodiment of a mobile inspection method shown according to an exemplary embodiment;

[0027] Figure 6 It is a schematic diagram of the hardware structure of an electronic device shown according to an exemplary embodiment;

[0028] Figure 7 It is a schematic diagram of the structure of a storage medium shown according to an exemplary embodiment. Detailed Description of the Embodiments

[0029] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0030] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0031] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0032] As mentioned above, the mobile patrol inspection faces scenarios that require a large number of algorithm tasks, and the training data resources corresponding to different algorithm tasks are not balanced. For example, for common urban road surface segmentation, vehicle or personnel detection, the data volume is large and the model fitting effect is good. However, for the specific detection tasks of mobile patrol inspection, such as detecting randomly piled materials and drying along the street, the data is relatively small. Full-parameter training is prone to overfitting, resulting in a poor model fitting effect. Moreover, mobile patrol inspection is generally carried on a vehicle-mounted computing box (i.e., an edge device), and the computing power of the vehicle-mounted computing box is relatively low compared to common server hosts. At the same time, running the models of each algorithm task will be affected by the computing power limit, resulting in relatively low computing efficiency.

[0033] To solve the above technical problems, the task model training method proposed in this application combines tasks with more training data and tasks with less training data in the actual application scenario using a combined model instead of training multiple models separately, so as to reduce the number of models. That is, the task with more training data is used as the main task, and the task with less training data is used as the secondary task. The model is made to learn the main task and the secondary task in stages. That is, first, the pre-trained model is made to learn the main task using the first data set with main task annotations to obtain the main task model. Then, the second data set with secondary task annotations and the pre-trained model are used to add secondary task annotations to the first data set, and the main task model is used to add main task annotations to the second data set. Furthermore, the main task model is trained using the first data set and the second data set that have both secondary task annotations and main task annotations, so that the model can utilize the main task ability when learning the secondary task and keep the main task ability from degrading, enabling the finally trained multi-task model to effectively improve the accuracy of the few-shot secondary task and reduce the training difficulty of the few-shot task.

[0034] Therefore, by inputting the inspection images of mobile inspection into the multi-task model obtained through the above training, the prediction results of all algorithm tasks required for mobile inspection output by the multi-task model can be obtained. Moreover, even if the multi-task model is deployed to the edge devices of mobile inspection, since the computing power consumed by a single model is small, the computing efficiency of mobile inspection under edge devices can be effectively improved.

[0035] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the aforementioned technical problems. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of this application in detail with reference to the accompanying drawings.

[0036] Embodiment 1:

[0037] Figure 1 FIG. is a flowchart of an embodiment of a multi-task model training method shown according to an exemplary embodiment. The embodiments of this application involve two types of tasks, namely the main task and the branch task, as Figure 2 shown. The main task is a common task with relatively more training data, such as vehicle detection, person detection, road surface segmentation, etc. It is relatively easy to collect a large amount of training data, and the model training effect is good. The branch task is a special task with relatively less training data, such as road garbage detection, random material stacking detection, road crack segmentation, etc. It is not easy to collect training data, the amount of training data is relatively small, and the model training is prone to overfitting.

[0038] As Figure 1 shown, the training method includes the following steps:

[0039] Step 101: Obtain a first data set containing main task annotations, and use the first data set to train a pre-trained model to obtain a main task model.

[0040] In this step, the first data set includes training samples and main task annotations of the training samples. The training samples are images, on which target boxes and the categories of the target boxes, such as people, vehicles, etc., are marked, and / or semantic information is marked on each pixel of the image, such as -0 representing the road surface and -1 representing non-road surface.

[0041] In this embodiment, the pre-trained model adopts a lightweight multi-task algorithm framework, including feature extraction and task prediction heads. However, the present application does not make specific limitations on the network model structure. Taking the Anchor-based detection model as an example, it includes a backbone network, a feature pyramid, and a detection head. The backbone network performs initial feature extraction on the input image to obtain a multi-dimensional feature map. The feature pyramid further performs feature extraction at different scales on the multi-dimensional feature map to obtain a multi-scale feature map. Then, the detection head uses the multi-scale feature map to perform corresponding detection tasks and outputs detection results. Of course, if a segmentation task is involved, a segmentation head can be further added, and the segmentation head uses the multi-scale feature map to perform corresponding segmentation tasks and outputs segmentation results.

[0042] Furthermore, since the first data set only has main task annotations, the pre-trained model is trained with the first data set, and the obtained main task model is only applicable to predicting the main task. And since the main task is a common task with relatively more training data, the fitting effect of the main task model trained with the first data set is relatively good, and the prediction accuracy is high.

[0043] In an alternative embodiment, the main tasks may include two types: detection tasks and segmentation tasks. Accordingly, the main task annotations correspondingly include main detection task annotations and main segmentation task annotations.

[0044] Based on this, the process of obtaining the first data set containing main task annotations includes:

[0045] First, obtain a first sub-data set containing main detection task annotations and a second sub-data set containing main segmentation task annotations; wherein, both the first sub-data set and the second sub-data set are the original training data for the corresponding tasks.

[0046] Then, according to the pre-trained model, add main segmentation task annotations to the first sub-data set to obtain an updated first sub-data set, and according to the pre-trained model, add main detection task annotations to the second sub-data set to obtain an updated second sub-data set, so that the original training data can be used for both main detection task training and main segmentation task training, and the manual annotation cost is reduced.

[0047] Finally, merge the updated first sub-data set and the updated second sub-data set into the first data set.

[0048] In a specific embodiment, as described above, the main tasks are of two types: detection tasks and segmentation tasks. Correspondingly, the pre-trained model includes a detection head and a segmentation head. Among them, the detection head is used to implement the prediction of the main detection task, and the segmentation head is used to implement the prediction of the main segmentation task.

[0049] When adding main line segmentation task annotations to the first sub-dataset according to the prediction model, the part of the pre-trained model except the segmentation head can be trained using the first sub-dataset to obtain a main line detection task model. Then, the second sub-dataset is predicted through this main line detection task model, and the obtained prediction results are used to add main line detection task annotations to the second sub-dataset, thereby realizing the automatic addition of main line detection task annotations for the second sub-dataset. Among them, the part of the pre-trained model except the segmentation head includes a feature extraction part and a detection head. During the training process, mainly the parameters of the detection head are optimized, and the parameters of the feature extraction part can be fine-tuned or not adjusted.

[0050] When adding main line segmentation task annotations to the first sub-dataset according to the pre-trained model, the part of the pre-trained model except the detection head can be trained using the second sub-dataset to obtain a main line segmentation task model. Then, the first sub-dataset is predicted through this main line segmentation task model, and the obtained prediction results are used to add main line segmentation task annotations to the first sub-dataset, thereby realizing the automatic addition of segmentation task annotations for the first sub-dataset. Among them, the part of the pre-trained model except the detection head includes a feature extraction part and a segmentation head. During the training process, mainly the parameters of the segmentation head are optimized, and the parameters of the feature extraction part can be fine-tuned or not adjusted.

[0051] It should be noted here that since the main line detection task model is obtained by training with the first sub-dataset with more training data, the prediction accuracy of the main line detection task model is relatively high. Therefore, the quality of the prediction results of the second sub-dataset in the main line detection task model is relatively good, and thus the main line detection task annotations added to the second sub-dataset belong to high-quality annotations. Similarly, the main line segmentation task annotations added to the first sub-dataset also belong to high-quality annotations.

[0052] Based on the description of the above embodiments, see Figure 3 As shown, the specific implementation process of the main line task model includes:

[0053] First, prepare the training data required for the main line task. Assume that there are two types of tasks in the main line task: detection task and segmentation task. Usually, the training data volume of a single task is greater than 5000 cases. The training data corresponding to the detection task is the first sub-dataset data_det_1, and the training data corresponding to the segmentation task is the second sub-dataset data_seg_1, and the data volume of each task is 5000 cases.

[0054] Then, use data_det_1 to train the part of the pre-trained model model_init except for the segmentation head to obtain the main-line detection task model model_det_1. Use model_det_1 to predict data_seg_1, and use the prediction results to add high-quality main-line detection task annotations to data_seg_1.

[0055] Similarly, use data_seg_1 to train the part of the pre-trained model model_init except for the detection head to obtain the main-line segmentation task model model_seg_1. Use model_seg_1 to predict data_det_1, and use the prediction results to add high-quality main-line segmentation task annotations to data_det_1.

[0056] Finally, merge the training data generated in the above two steps into the first dataset data_master, with a total data volume of 10,000 cases. Each training data in it contains high-quality main-line detection task annotations and main-line segmentation task annotations, so as to use data_master to train model_init to obtain the main-line task model model_master.

[0057] Step 102: Obtain a second dataset containing branch task annotations, and add branch task annotations to the first dataset according to the second dataset and the pre-trained model to obtain the updated first dataset.

[0058] In this step, the second dataset includes training samples and branch task annotations of the training samples. The training samples are images, on which target boxes and the categories of the target boxes are marked, such as road garbage, randomly piled materials, etc., and / or semantic information is marked for each pixel on the image, such as -0 indicating a road crack and -1 indicating a non-road crack.

[0059] It should be noted that, as mentioned above, the branch task is a special task with relatively few training data, so the training data of the second dataset is less than that of the first dataset mentioned above.

[0060] In an optional implementation manner, the branch task can also include two types: detection task and segmentation task. Accordingly, the branch task annotation can correspondingly include branch detection task annotation and branch segmentation task annotation.

[0061] Based on this, the process of obtaining the second dataset containing branch task annotations includes:

[0062] First, obtain a third sub-dataset containing branch detection task annotations and a fourth sub-dataset containing branch segmentation task annotations; among them, both the third sub-dataset and the fourth sub-dataset are the original training data for the corresponding tasks.

[0063] Then, according to the pre-trained model, add branch line segmentation task annotations to the third sub-dataset to obtain the updated third sub-dataset, and add branch line detection task annotations to the fourth sub-dataset according to the pre-trained model to obtain the updated fourth sub-dataset, so that the original training data can be used for both branch line detection task training and branch line segmentation task training, and reduce the manual annotation cost;

[0064] Finally, merge the updated third sub-dataset and the updated fourth sub-dataset into the second dataset.

[0065] In a specific embodiment, as described above, there are two types of branch line tasks, namely detection task and segmentation task. Correspondingly, the pre-trained model includes a detection head and a segmentation head. Among them, the detection head is used to implement the prediction of the branch line detection task, and the segmentation head is used to implement the prediction of the branch line segmentation task. It should be noted here that the network structures of the detection head and the segmentation head for the main line task can be the same as those of the detection head and the segmentation head for the branch line task, but the last output layers of the detection head and the segmentation head included in the pre-trained model for branch line task learning are set according to the output quantity required by the branch line task, and the last output layers of the detection head and the segmentation head included in the pre-trained model for main line task learning are set according to the output quantity required by the main line task.

[0066] When adding branch line detection task annotations to the fourth sub-dataset according to the pre-trained model, the part of the pre-trained model except the segmentation head can be trained using the third sub-dataset to obtain a branch line detection task model, and then the fourth sub-dataset can be predicted by the branch line detection task model, and the obtained prediction results can be used to add branch line detection task annotations to the fourth sub-dataset, so as to realize the automatic addition of branch line detection task annotations for the fourth sub-dataset.

[0067] Among them, the part of the pre-trained model except the segmentation head includes a feature extraction part and a detection head. During the training process, mainly the parameters of the detection head are optimized, and the parameters of the feature extraction part can be slightly adjusted or not adjusted. When using the prediction results to add branch line detection task annotations to the fourth sub-dataset, the branch line detection results with a confidence level exceeding the preset value in the prediction results can be used as branch line detection task annotations.

[0068] When adding branch line segmentation task annotations to the third sub-dataset according to the pre-trained model, the part of the pre-trained model except the detection head can be trained using the fourth sub-dataset to obtain a branch line segmentation task model. Then, the third sub-dataset is predicted using the branch line segmentation task model, and the obtained prediction results are used to add branch line segmentation task annotations to the third sub-dataset, thereby realizing the automatic addition of branch line segmentation task annotations to the third sub-dataset. Among them, the part of the pre-trained model except the detection head includes the feature extraction part and the segmentation head. During the training process, mainly the parameters of the segmentation head are optimized, and the parameters of the feature extraction part can be fine-tuned or not adjusted. When using the prediction results to add branch line segmentation task annotations to the third sub-dataset, the branch line segmentation results with a confidence level exceeding the preset value in the prediction results can be used as branch line segmentation task annotations.

[0069] It should be noted here that since the branch line detection task model is obtained by training with the third sub-dataset with relatively few training data, the prediction accuracy of the branch line detection task model will be relatively low. As a result, the quality of the prediction results of the fourth sub-dataset in the branch line detection task model is relatively low, and thus the branch line detection task annotations added to the fourth sub-dataset belong to low-quality annotations. Similarly, the branch line segmentation task annotations added to the third sub-dataset also belong to low-quality annotations.

[0070] Furthermore, after obtaining the second dataset, it is also necessary to use the above-mentioned branch line detection task model and branch line segmentation task model to add branch line task annotations to the first dataset.

[0071] Specifically, the first dataset is predicted using the branch line detection task model, and the obtained prediction results are used to add branch line detection task annotations to the first dataset; and, the first dataset is predicted using the branch line segmentation task model, and the obtained prediction results are used to add branch line segmentation task annotations to the first dataset.

[0072] Among them, as mentioned above, the prediction accuracies of both the branch line detection task model and the branch line segmentation task model are relatively low. Therefore, both the branch line detection task annotations and the branch line segmentation task annotations added to the first dataset belong to low-quality annotations, that is, the branch line task annotations added to the first dataset also belong to low-quality annotations.

[0073] Step 103: Add main line task annotations to the second dataset according to the main line task model to obtain the updated second dataset.

[0074] Specifically, the second dataset is predicted using the main line task model, and the obtained prediction results are used to add main line task annotations to the second dataset to realize the automatic addition of main line task annotations to the second dataset.

[0075] Among them, as mentioned above, since the fitting effect of the main task model is relatively good and the prediction accuracy is high, the main task annotations added to the second dataset belong to high-quality annotations.

[0076] Step 104: Use the updated first dataset and the updated second dataset to train the main task model to obtain a multi-task model.

[0077] In this step, the training data in the first dataset and the second dataset are both data with both branch task annotations and main task annotations. After modifying the last output layer in the main task model according to the outputs required by the branch tasks and the main task, the main task model is trained by combining the first dataset and the second dataset together, so that the model can utilize the main task ability when learning new branch tasks and keep the main task ability from degrading. Thus, the easily collected conventional data can be effectively used as supplementary training to obtain a good improvement in effect.

[0078] It should be noted that during the training process of the main task model using the first dataset and the second dataset with high-quality annotations and low-quality annotations, in order to reduce the impact of low-quality annotations on model learning, when calculating the loss, the loss weight of the branch task annotations added to the first dataset is less than the loss weight of the main task annotations in the first dataset, and the loss weight of the branch task annotations added to the second dataset is less than the loss weights of the existing branch task annotations and the added main task annotations in the second dataset. This can reduce the impact of these low-quality annotations.

[0079] Based on the descriptions of the above Step 101 to Step 104, see Figure 4 As shown, the process of implementing a multi-task model based on the main task model specifically includes:

[0080] First, prepare the training data required for the branch tasks. Assume that there are also two types of branch tasks, namely detection tasks and segmentation tasks. Generally, the training data required for a single task is more than 1000 cases. The training data corresponding to the detection task is the third sub-dataset data_det_2, and the training data corresponding to the segmentation task is the fourth sub-dataset data_seg_2. The data volume of each task is 1000 cases.

[0081] Then, use data_det_2 to train the part of the pre-trained model except the segmentation head to obtain a branch detection task model model_det_2. Use model_det_2 to predict the above first dataset data_master and data_seg_2, and use the prediction results to add low-quality branch detection task annotations to data_master and data_seg_2.

[0082] Similarly, use data_seg_2 to train the part of the pre-trained model model_init except for the detection head to obtain the branch segmentation task model model_seg_2. Use model_seg_2 to predict the above first dataset data_master and data_det_2, and use the prediction results to add low-quality branch segmentation task annotations to data_master and data_det_2.

[0083] Then, use the above main task model model_master to predict data_det_2 and data_seg_2 respectively, and use the prediction results to add high-quality main task annotations to data_det_2 and data_seg_2 respectively.

[0084] Furthermore, merge the training data generated in the above steps into the second dataset data_branch, and merge data_master and data_branch with the added branch task annotations. The total data volume is 12,000 cases.

[0085] Finally, use the merged data_master and data_branch to train model_master. Among them, the loss weight of the branch task annotation added to data_master is 1 / 2 of the original weight, and the loss weights of the branch segmentation task annotation added to data_det_2 and the branch detection task annotation added to data_seg_2 are both 1 / 2 of the original weight. After training, the final multi-task model model_branch is obtained.

[0086] So far, the above Figure 1 shown training process is completed. By using a merged model instead of training multiple models separately for tasks with more training data and tasks with less training data in the actual application scenario, the number of models is reduced. That is, the task with more training data is used as the main task, and the task with less training data is used as the branch task. And let the model learn the main task and the branch task in stages. That is, first use the first dataset containing the main task annotations to let the pre-trained model learn the main task to obtain the main task model. Then, use the second dataset containing the branch task annotations and the pre-trained model to add branch task annotations to the first dataset, and use the main task model to add main task annotations to the second dataset. Furthermore, use the first dataset and the second dataset with both branch task annotations and main task annotations to train the main task model, so that the model can utilize the main task ability when learning the branch task and keep the main task ability from degrading. This makes the final trained multi-task model effectively improve the accuracy of the few-shot branch task and reduce the training difficulty of the few-shot task.

[0087] Embodiment 2:

[0088] Based on the above Figure 1 shown embodiment, Figure 5 FIG. is a flowchart of an embodiment of mobile inspection shown according to an exemplary embodiment, including the following steps:

[0089] Step 501: Obtain inspection images collected by the mobile vehicle.

[0090] Step 502: Input the inspection image into the multi-task model to obtain the main task prediction result and the branch task prediction result output by the multi-task model.

[0091] In this step, the multi-task model is the model obtained in the above Figure 1 shown embodiment. The line task prediction result is the prediction result of the regular task in the tasks required for mobile inspection, such as the prediction results of the person detection task, the vehicle detection task, and the road segmentation task. The branch task prediction result is the prediction result of the special task in mobile inspection, such as the prediction results of the road garbage detection task, the disorderly stacking of materials detection task, and the road crack segmentation task.

[0092] Based on Embodiment 2 above, since the multi-task model realizes the combination of the algorithm tasks required for mobile inspection into a unified model for prediction, and can also ensure the accuracy of the data-scarce tasks, by inputting the inspection image into the unified multi-task model, the prediction results of all algorithm tasks can be obtained, that is, the main task prediction result and the branch task prediction result. In this way, when the multi-task model is deployed to the edge device of mobile inspection, since the single model consumes less computing power, the computing efficiency of mobile inspection under the edge device can be effectively improved.

[0093] The execution entity of the embodiments of this application can be an application program, service, instance, functional module in software form, virtual machine (VM), container, cloud server, etc., or a hardware device with data processing functions (such as a server or terminal device) or a hardware chip (such as a CPU, GPU, FPGA, NPU, AI acceleration card, or DPU). The device for implementing multi-task model training or mobile patrol inspection can be deployed on the computing device of the application party providing the corresponding service or on a cloud computing platform providing computing power, storage, and network resources. The service mode provided by the cloud computing platform to the outside world can be IaaS (Infrastructure as a Service), PaaS (Platform as a Service), SaaS (Software as a Service), or DaaS (Data as a Service). Taking the platform providing SaaS (Software as a Service) as an example, the cloud computing platform can use its own computing resources to provide multi-task model training or the training of mobile patrol inspection models or the function execution of multi-task model training or mobile patrol inspection modules. The specific application architecture can be built according to service requirements. For example, the platform can provide construction services based on the above models to application parties or individuals using the platform resources, and further call the above models and implement the functions of online or offline multi-task model training or mobile patrol inspection based on multi-task model training or mobile patrol inspection requests submitted by relevant client or server devices.

[0094] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0095] The embodiments of this application also provide an electronic device corresponding to the multi-task model training method or mobile patrol inspection method provided in the foregoing embodiments to execute the above multi-task model training method or mobile patrol inspection method.

[0096] Figure 6The following is a hardware structure diagram of an electronic device shown according to an exemplary embodiment. The electronic device includes: a communication interface 601, a processor 602, a memory 603, and a bus 604; among them, the communication interface 601, the processor 602, and the memory 603 complete their mutual communication through the bus 604. By reading and executing the machine-executable instructions corresponding to the control logic of the multi-task model training method or the mobile inspection method in the memory 603, the processor 602 can execute the multi-task model training method or the mobile inspection method described above. For the specific content of this method, refer to the above embodiments and will not be repeated here.

[0097] The memory 603 mentioned in this application can be any electronic, magnetic, optical, or other physical storage device, and can contain stored information such as executable instructions, data, etc. Specifically, the memory 603 can be RAM (Random Access Memory), flash memory, a storage drive (such as a hard disk drive), any type of storage disk (such as an optical disk, DVD, etc.), or a similar storage medium, or a combination thereof. Through at least one communication interface 601 (which can be wired or wireless), a communication connection is achieved between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0098] The bus 604 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 603 is used to store programs, and the processor 602 executes the programs after receiving execution instructions.

[0099] The processor 602 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 602 or by instructions in software form. The above-mentioned processor 602 can be a general-purpose processor, including a network processor (abbreviated as NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor.

[0100] The electronic device provided by the embodiment of the present application and the multi-task model training method or mobile inspection method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by them.

[0101] The embodiment of the present application also provides a computer-readable storage medium corresponding to the multi-task model training method or mobile inspection method provided by the foregoing embodiment. Please refer to Figure 7 As shown, the computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the multi-task model training method or mobile inspection method provided by any of the foregoing embodiments.

[0102] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0103] The computer-readable storage medium provided by the above embodiment of the present application and the multi-task model training method or mobile inspection method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application program stored therein.

[0104] Those skilled in the art will readily think of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0105] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.

[0106] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for training a multi-task model, characterized in that, The method includes: Obtaining a first dataset containing main task annotations, and training a pre-trained model using the first dataset to obtain a main task model; Obtaining a second dataset containing sub-task annotations, and adding sub-task annotations to the first dataset according to the second dataset and the pre-trained model to obtain an updated first dataset; the training data of the first dataset is more than the training data of the second dataset; Adding main task annotations to the second dataset according to the main task model to obtain an updated second dataset; Training the main task model using the updated first dataset and the updated second dataset to obtain a multi-task model.

2. The method according to claim 1, characterized in that The main task annotations include main detection task annotations and main segmentation task annotations; the obtaining of the first dataset containing main task annotations includes: Obtaining a first sub-dataset containing main detection task annotations and a second sub-dataset containing main segmentation task annotations; Adding main detection task annotations to the second sub-dataset according to the pre-trained model to obtain an updated second sub-dataset; Adding main segmentation task annotations to the first sub-dataset according to the pre-trained model to obtain an updated first sub-dataset; Merging the updated first sub-dataset and the updated second sub-dataset into the first dataset.

3. The method according to claim 2, wherein The pre-trained model includes a detection head and a segmentation head; The adding of main detection task annotations to the first sub-dataset according to the pre-trained model includes: training the part of the pre-trained model except the segmentation head using the first sub-dataset to obtain a main detection task model; predicting the second sub-dataset through the main detection task model, and adding main detection task annotations to the second sub-dataset using the obtained prediction results; The adding of main segmentation task annotations to the first sub-dataset according to the pre-trained model includes: training the part of the pre-trained model except the detection head using the second sub-dataset to obtain a main segmentation task model; predicting the first sub-dataset through the main segmentation task model, and adding main segmentation task annotations to the first sub-dataset using the obtained prediction results.

4. The method according to claim 1, wherein The sub-task annotations include sub-detection task annotations and sub-segmentation task annotations; the obtaining of the second dataset containing sub-task annotations includes: Obtaining a third sub-dataset containing sub-detection task annotations and a fourth sub-dataset containing sub-segmentation task annotations; Adding sub-segmentation task annotations to the third sub-dataset according to the pre-trained model to obtain an updated third sub-dataset; Adding sub-detection task annotations to the fourth sub-dataset according to the pre-trained model to obtain an updated fourth sub-dataset; Merging the updated third sub-dataset and the updated fourth sub-dataset into the second dataset.

5. The method according to claim 4, characterized in that, The pre-trained model includes a detection head and a segmentation head; Adding branch line detection task annotations to the fourth sub-dataset according to the pre-trained model includes: training the part of the pre-trained model except the segmentation head using the third sub-dataset to obtain a branch line detection task model; predicting the fourth sub-dataset through the branch line detection task model, and adding branch line detection task annotations to the fourth sub-dataset using the obtained prediction results; Adding branch line segmentation task annotations to the third sub-dataset according to the pre-trained model includes: training the part of the pre-trained model except the detection head using the fourth sub-dataset to obtain a branch line segmentation task model; predicting the third sub-dataset through the branch line segmentation task model, and adding branch line segmentation task annotations to the third sub-dataset using the obtained prediction results.

6. The method according to claim 5, wherein Adding branch line task annotations to the first dataset according to the second dataset and the pre-trained model includes: Predicting the first dataset through the branch line detection task model, and adding branch line detection task annotations to the first dataset using the obtained prediction results; Predicting the first dataset through the branch line segmentation task model, and adding branch line segmentation task annotations to the first dataset using the obtained prediction results.

7. The method according to any one of claims 4 to 6, characterized in that, During the training process of the main line task model, the loss weight of the branch line task annotations added to the first dataset is less than the loss weight of the main line task annotations in the first dataset, and the loss weight of the branch line task annotations added to the second dataset is less than the loss weight of the existing branch line task annotations and the added main line task annotations in the second dataset.

8. The method according to claim 1, wherein Adding main line task annotations to the second dataset according to the main line task model includes: Predicting the second dataset through the main line task model, and adding main line task annotations to the second dataset using the obtained prediction results.

9. A mobile inspection method, characterized in that, The method includes: Obtaining inspection images collected by a mobile vehicle; Inputting the inspection images into the multi-task model obtained from any one of claims 1 to 8; Obtaining the main line task prediction results and branch line task prediction results output by the multi-task model.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the method according to any one of claims 1-9.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-9.