Multitask prediction method, electronic equipment and storage medium
Through the combination of focal convolution modules and low-rank convolution units, shared features and independent features in the multi-task model are extracted, which solves the problems of excessive parameters and training complexity in multi-task learning and achieves efficient and accurate multi-task prediction.
Patent Information
- Application Number
- CN202510757057.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-19
AI Technical Summary
In existing multi-task learning methods, hard sharing leads to excessive number of parameters, affecting prediction accuracy, while soft sharing leads to increased number of parameters, training complexity and prolonged time.
A combination of focal convolution modules, convolution module groups and multi-scale fusion modules is adopted to extract shared features and independent features between multiple tasks through low-rank convolution units, and convolution processing and fusion are performed in combination with low-rank parameter matrices to optimize the training process.
While reducing the number of model parameters, the accuracy and efficiency of multi-task prediction are improved, the conflict and imbalance problems between tasks are solved, and the efficiency of feature extraction is enhanced.
Smart Images

Figure CN120673225A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to a multi-task prediction method, electronic device, and storage medium. Background Art
[0002] Multi-Task Learning (MTL) is a machine learning method that aims to improve a model's generalization ability by simultaneously training multiple related tasks. Its core idea is to leverage shared information between tasks to improve the learning performance of each task. MTL is widely used in fields such as natural language processing, computer vision, and speech recognition.
[0003] Currently, there are two main approaches to multi-task model architecture: hard sharing and soft sharing. Hard sharing means that the model's main components share parameters when processing different tasks, learning shared features across different tasks, resulting in a relatively small number of model parameters. Soft sharing means that each task has independent parameters, and the model learns its own feature representation for each task.
[0004] However, the disadvantage of hard sharing is that a large number of parameters of the main body are shared, resulting in the performance of multiple tasks not being able to achieve the effect of a single task, thus affecting the accuracy of the task prediction results; the disadvantage of soft sharing is that the parameters increase exponentially with the number of tasks, which not only increases the training complexity but also causes the task prediction time to become longer. Summary of the Invention
[0005] In view of the above-mentioned defects or deficiencies in the prior art, the present application aims to provide a multi-task prediction method, electronic device and storage medium to solve the problems of low multi-task prediction accuracy and long task prediction time in the prior art.
[0006] This embodiment of the present application provides a multi-task prediction method, which includes:
[0007] Acquire the target image;
[0008] Input the target image into the trained multi-task model to obtain prediction results for each task;
[0009] The multi-task model includes a focal convolution module, at least one convolution module group, and a multi-scale fusion module corresponding to each task;
[0010] The focal convolution module is used to perform convolution processing on the target image and input the obtained features into the connected convolution module group;
[0011] The convolution module group is used to extract the initial shared features between multiple tasks from the input, determine the final shared features between the multiple tasks and the independent features of each task based on the initial shared features, and input the independent features of each task into the multi-scale fusion module corresponding to each task respectively, and input the final shared features into other connected convolution module groups so that other convolution module groups can extract the final shared features between the multiple tasks and the independent features of each task at another scale;
[0012] The multi-scale fusion module is used to fuse independent features of different scales corresponding to the task and determine the task prediction result based on the fusion result.
[0013] Optionally, the focal convolution module includes a slicing unit, a first splicing unit and a first low-rank convolution unit;
[0014] The slicing unit is configured to slice the target image and input each obtained slice into the first stitching unit;
[0015] The first splicing unit is configured to perform splicing processing on each slice and input a first splicing result obtained into the first low-rank convolution unit;
[0016] The first low-rank convolution unit is used to perform convolution processing and fusion on the input according to the internal low-rank parameter matrix and the original parameter matrix.
[0017] Optionally, the convolution module group includes a shared feature convolution module and a task feature convolution module, and the shared feature convolution module includes a second low-rank convolution unit;
[0018] The shared feature convolution module is used to extract initial shared features between multiple tasks from the input and input the initial shared features into the task feature convolution module;
[0019] The task feature convolution module is used to extract the independent features of each task and the final shared features between multiple tasks based on the initial shared features;
[0020] The second low-rank convolution unit is used to perform convolution processing and fusion on the input according to the internal low-rank parameter matrix and the original parameter matrix, so as to extract initial shared features between multiple tasks from the input.
[0021] Optionally, the task feature convolution module includes a third low-rank convolution unit, a fourth low-rank convolution unit, at least one residual module, a second splicing unit and an output unit;
[0022] The third low-rank convolution unit is used to perform convolution processing and fusion on the initial shared features according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result into the residual module;
[0023] The residual module is used to perform convolution and jump connection processing on the input convolution result, and input the obtained residual result into the second splicing unit;
[0024] The fourth low-rank convolution unit is configured to perform convolution processing and fusion on the initial shared features according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result into the second splicing unit;
[0025] The second splicing unit is used to splice the input convolution result and the residual result, and input the obtained second splicing result to the output unit;
[0026] The output unit is configured to extract the independent features of each task and the final shared features among multiple tasks from the second splicing result.
[0027] Optionally, the residual module includes a fifth low-rank convolution unit, a sixth low-rank convolution unit and a connection unit;
[0028] The fifth low-rank convolution unit is used to perform convolution processing and fusion on the input convolution result according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result to the sixth low-rank convolution unit and the connection unit;
[0029] The sixth low-rank convolution unit is used to perform convolution processing and fusion on the input convolution result according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result to the connection unit;
[0030] The connection unit is used to add all convolution results of the input.
[0031] Optionally, the output unit includes a shared feature output layer and a task feature output layer corresponding to each task;
[0032] The shared feature output layer is used to perform convolution processing and fusion on the second splicing result according to the internal shared low-rank parameter matrix and the original parameter matrix to obtain the final shared features between multiple tasks;
[0033] The task feature output layer is used to perform convolution processing and fusion on the second splicing result according to the internal original parameter matrix and the task low-rank parameter matrix of the corresponding task to obtain independent features of the corresponding task.
[0034] Optionally, the low-rank convolution unit includes a low-rank convolution layer, a normalization layer, and an activation layer;
[0035] The low-rank convolution layer is used to perform convolution processing on the input according to the internal first low-rank parameter matrix, the second low-rank parameter matrix and the original parameter matrix, and fuse the obtained convolution results, and input the obtained fusion results into the normalization layer;
[0036] The normalization layer is used to normalize the input and input the normalized result to the activation layer;
[0037] The activation layer is used to perform nonlinear transformation on the input.
[0038] Optionally, the training of the multi-task model includes:
[0039] Build the initial model and obtain pre-training samples and optimized training samples;
[0040] Training the initial model based on the pre-training samples to update all original parameter matrices in the initial model until a pre-training stop condition is reached, thereby obtaining a pre-training model;
[0041] Freeze all original parameter matrices in the pre-trained model, and train the pre-trained model based on the optimized training sample to update all low-rank parameter matrices in the initial model until the optimization training stop condition is reached, thereby obtaining the multi-task model.
[0042] An embodiment of the present application further provides an electronic device, comprising:
[0043] processor and memory;
[0044] The processor is configured to execute the steps of the multi-task prediction method provided in any embodiment of the present application by calling the program or instruction stored in the memory.
[0045] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program or instruction, wherein the program or instruction enables a computer to execute the steps of the multi-task prediction method provided in any embodiment of the present application.
[0046] In summary, the present application proposes a multi-task prediction method, which inputs a target image into a trained multi-task model, performs convolution processing on the target image through the focal convolution module in the model, extracts the final shared features between multiple tasks and the independent features of each task through the convolution module group in the model, and inputs the final shared features into other connected convolution module groups, so that other convolution module groups can extract shared features and independent features at another scale, and input the independent features of each task into the multi-scale fusion module corresponding to each task, and then fuses the independent features of the corresponding tasks at different scales through the multi-scale fusion module in the model, and outputs the corresponding task prediction results. This method can extract the shared features between multiple tasks and the independent features of each task, and then fuse the independent features of different scales output by different convolution module groups for task prediction, so that the tasks maintain sharing while supporting the independence of each task, solving the conflict and imbalance problems between multiple tasks, and improving the accuracy of multi-task prediction. In addition, by extracting shared features, the efficiency of feature extraction between tasks can also be improved, thereby improving the efficiency of multi-task prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0048] Figure 1 This is a flowchart of a multi-task prediction method provided by an embodiment of the present application;
[0049] Figure 2 This is an architectural diagram of a multi-tasking module provided in an embodiment of the present application;
[0050] Figure 3 This is an architecture diagram of a focal convolution module provided in an embodiment of the present application;
[0051] Figure 4 This is an architecture diagram of a low-rank convolution unit provided in an embodiment of the present application;
[0052] Figure 5 This is a processing diagram of a low-rank convolutional layer provided in an embodiment of the present application;
[0053] Figure 6 This is an architectural diagram of a convolution module group provided in an embodiment of the present application;
[0054] Figure 7This is an architecture diagram of a task feature convolution module provided in an embodiment of the present application;
[0055] Figure 8 This is an architecture diagram of a residual module provided in an embodiment of the present application;
[0056] Figure 9 This is a processing diagram of an output unit provided in an embodiment of the present application;
[0057] Figure 10 This is an architecture diagram of a multi-scale fusion module provided in an embodiment of the present application;
[0058] Figure 11 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.
[0060] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0061] As mentioned in the background technology, to address the problems in the prior art, this application proposes a multi-task prediction method. Figure 1 This is a flowchart of a multi-task prediction method provided by an embodiment of the present application. Figure 1 , the multi-task prediction method specifically includes:
[0062] S110: Acquire a target image.
[0063] S120: Input the target image into the trained multi-task model to obtain prediction results for each task.
[0064] The multi-task model includes a focal convolution module, at least one convolution module group, and a multi-scale fusion module corresponding to each task;
[0065] The focal convolution module is used to perform convolution processing on the target image and input the obtained features into the connected convolution module group;
[0066] A convolution module group is used to extract the initial shared features between multiple tasks from the input, determine the final shared features between multiple tasks and the independent features of each task based on the initial shared features, input the independent features of each task into the multi-scale fusion module corresponding to each task, and input the final shared features into other connected convolution module groups so that other convolution module groups can extract the final shared features between multiple tasks and the independent features of each task at another scale;
[0067] The multi-scale fusion module is used to fuse independent features of different scales corresponding to the task and determine the task prediction result based on the fusion result.
[0068] In an embodiment of the present application, the multi-task may include at least two of lane detection, 2D object detection, 3D object detection, feasible area detection, and region segmentation. The target image may be input into a trained multi-task model so that the multi-task model performs multi-task prediction based on the target image, obtaining task prediction results corresponding to each task.
[0069] Specifically, the multi-task model consists of a focal convolution module, at least one convolution module group, and a multi-scale fusion module corresponding to each task. If there is only one convolution module group, the focal convolution module is connected to the convolution module group, which in turn is connected to each multi-scale fusion module. If there are multiple convolution module groups, the focal convolution module is connected to the first convolution module group, which in turn is connected to the next convolution module group and each multi-scale fusion module.
[0070] For example, Figure 2 This is an architecture diagram of a multi-task module provided in an embodiment of the present application, such as Figure 2 As shown in the figure, taking the number of convolution module groups equal to 4 as an example, after the target image is input into the multi-task model, it first passes through the focal convolution module in the multi-task model. After processing by the focal convolution module, it will be sent to the convolution module group 1. After processing by the convolution module group 1, part of the result will be sent to the convolution module group 2, and part of the result will be sent to each multi-scale fusion module. After processing by the convolution module group 2, part of the result will be sent to the convolution module group 3, and part of the result will be sent to each multi-scale fusion module. After processing by the convolution module group 3, part of the result will be sent to the convolution module group 4, and part of the result will be sent to each multi-scale fusion module. After processing by the convolution module group 4, part of the result will be sent to each multi-scale fusion module.
[0071] Specifically, after the target image is input into the multi-task model, the focal convolution module can first perform convolution processing on it to obtain the down-sampled features corresponding to the target image, such as obtaining twice the down-sampled features, to achieve feature aggregation of the target image and focus on the local features in the target image that are related to the task.
[0072] In a specific embodiment, the focal convolution module includes a slicing unit, a first splicing unit, and a first low-rank convolution unit;
[0073] a slicing unit, configured to slice the target image and input each obtained slice into the first splicing unit;
[0074] A first splicing unit is used to splice the slices and input the obtained first splicing result into the first low-rank convolution unit;
[0075] The first low-rank convolution unit is used to perform convolution processing and fusion on the input according to the internal low-rank parameter matrix and the original parameter matrix.
[0076] Among them, the slicing unit is used to slice the image, the first splicing unit is used to splice the slices, and the first low-rank convolution unit is a convolution unit including a low-rank parameter matrix and an original parameter matrix, and is used to perform convolution processing based on the low-rank parameter matrix and the original parameter matrix.
[0077] It should be noted that the rank of the low-rank parameter matrix is much smaller than the rank of the original parameter matrix, that is, the number of parameters in the low-rank parameter matrix is much smaller than the number of parameters in the original parameter matrix. During the model construction process, the original parameter matrix and the low-rank parameter matrix can be constructed for the first low-rank convolution unit. The purpose of constructing the original parameter matrix and the low-rank parameter matrix is to freeze the original parameter matrix and optimize the low-rank parameter matrix during model training to reduce training resources without increasing any inference overhead. In addition, the extracted features can more effectively represent the key information of the input data.
[0078] For example, Figure 3 This is an architecture diagram of a focal convolution module provided in an embodiment of the present application, such as Figure 3 As shown, the slicing unit is connected to the first splicing unit, and the first splicing unit is connected to the first low-rank convolution unit.
[0079] Specifically, after the target image is input into the multi-task model, it will first pass through the slicing unit of the focal convolution module, which can slice the target image to obtain multiple slices, and then input the multiple slices into the first splicing unit.
[0080] Furthermore, the first splicing unit can splice all the slices to splice multiple slices into a tensor along the specified channel dimension, obtain a first splicing result, and input the first splicing result into the first low-rank convolution unit.
[0081] Furthermore, the first low-rank convolution unit can perform convolution processing on the input through the internal low-rank parameter matrix, and at the same time, perform convolution processing on the input through the internal original parameter matrix, and then fuse the convolution results obtained by the two convolution branches, and input the fused results to the connected convolution module group.
[0082] In one example, the low-rank convolution unit includes a low-rank convolution layer, a normalization layer, and an activation layer;
[0083] The low-rank convolution layer is used to perform convolution on the input according to the internal first low-rank parameter matrix, the second low-rank parameter matrix, and the original parameter matrix, and fuse the convolution results obtained, and input the obtained fusion results into the normalization layer;
[0084] The normalization layer is used to normalize the input and input the normalized result into the activation layer;
[0085] Activation layer, used to perform nonlinear transformation on the input.
[0086] Among them, the low-rank convolution layer is connected to the normalization layer, and the normalization layer is connected to the activation layer. For example, Figure 4 This is an architecture diagram of a low-rank convolution unit provided in an embodiment of the present application, such as Figure 4 As shown in the figure, the first low-rank convolution unit can adopt this architecture. The first splicing result is first input into the low-rank convolution layer. After convolution and fusion processing, it enters the normalization layer for normalization processing, and finally passes through the activation layer for nonlinear transformation.
[0087] Specifically, the low-rank convolution layer includes two low-rank parameter matrices and an original parameter matrix. The low-rank convolution layer contains two convolution branches. The input passes through the two convolution branches respectively. The first convolution branch uses the original parameter matrix to convolve the input. The second convolution branch first uses the first low-rank parameter matrix to convolve the input, and then continues to use the second low-rank parameter matrix to convolve on this result. The convolution results obtained by the two convolution branches are finally fused and input into the normalization layer.
[0088] For example, Figure 5 This is a processing diagram of a low-rank convolutional layer provided in an embodiment of the present application, such as Figure 5As shown in the figure, assuming the input is X, the batch size, height, width, and number of channels of X are represented by B, H, W, and C respectively. The first convolution branch on the left is the original convolution, whose input is C channels and output is C_out channels. The convolution kernel size is k, and the convolution is performed using the original parameter matrix. The number of parameters in the original parameter matrix is C×C_out×k×k. The second convolution branch on the right is a low-rank convolution, which contains two low-rank parameter matrices. The number of parameters in the first low-rank parameter matrix is C×k×r×k, where r is the rank of the matrix, and the number of parameters in the second low-rank parameter matrix is r×k×C_out×k. The results of the two convolution branches can be added to obtain the output X_out, where the batch size, height, width, and number of channels are B, H, W, and C_out respectively.
[0089] Since the rank of the low-rank parameter matrix is a parameter much smaller than C, assuming that r is 2, C and _C_out are 256 and 512 respectively, and k is 3, then the number of parameters required for the original convolution branch is: 256×512×3×3=1179648, and the number of parameters required for the low-rank convolution branch is: (256×3)×(2×3)+(2×3)×(512×3)=13824. It can be seen that the sum of the number of parameters of the first low-rank parameter matrix and the second low-rank parameter matrix is much lower than the number of parameters of the original parameter matrix, and the low-rank convolution branch can save 98.8% of the parameters.
[0090] After the low-rank convolution layer, the normalization layer can continue to perform normalization, and then the activation layer can continue to perform nonlinear transformation. The activation layer can use an activation function such as SiLU (Sigmoid-Weighted Linear Unit).
[0091] Through the above-mentioned low-rank convolution unit, the low-rank parameter matrix can be shared between multiple tasks, thereby enabling good interaction between multiple tasks, while reducing model parameters and improving the efficiency of model feature extraction. The first low-rank convolution unit, the second low-rank convolution unit, the third low-rank convolution unit, the fourth low-rank convolution unit, the fifth low-rank convolution unit or the sixth low-rank convolution unit in the embodiment of the present application can all adopt this type of structure.
[0092] The focal convolution module provided in the above embodiment slices and splices the input target image in sequence through the slicing unit and the first splicing unit, which can optimize the computational efficiency and enhance the feature diversity. Moreover, the convolution processing is performed through the original parameter matrix and the low-rank parameter matrix inside the first low-rank convolution unit, which can realize efficient model training, reduce the computational complexity of the model, and improve the generalization ability of the model.
[0093] In an embodiment of the present application, after the target image undergoes convolution processing by the focal convolution module, the focal convolution module can input the obtained features into the convolution module group connected to it. The convolution module group can first extract the initial shared features between multiple tasks from the input, and then process the initial shared features to extract the final shared features between multiple tasks and the independent features of each task, and then input the independent features of each task into the multi-scale fusion module corresponding to each task respectively, and finally input the shared features into the next convolution module group, so that the next convolution module group performs shared feature extraction and independent feature extraction on the input at other scales, that is, extracts the final shared features between multiple tasks and the independent features of each task at another scale.
[0094] It should be noted that, assuming that the convolution module group connected to the focal convolution module is the first convolution module group, for the last convolution module group, it can input the independent features of each task into the multi-scale fusion module corresponding to each task respectively, without the need to input the final shared features into other modules.
[0095] In a specific embodiment, the convolution module group includes a shared feature convolution module and a task feature convolution module, and the shared feature convolution module includes a second low-rank convolution unit;
[0096] The shared feature convolution module is used to extract the initial shared features between multiple tasks from the input and input the initial shared features into the task feature convolution module;
[0097] The task feature convolution module is used to extract the independent features of each task and the final shared features between multiple tasks based on the initial shared features;
[0098] The second low-rank convolution unit is used to perform convolution processing and fusion on the input according to the internal low-rank parameter matrix and the original parameter matrix to extract the initial shared features between multiple tasks from the input.
[0099] Among them, the shared feature convolution module is used to extract the initial shared features between multiple tasks, and the task feature convolution module is used to extract the independent features of each task and the final shared features between multiple tasks.
[0100] Specifically, the shared feature convolution module may include a convolution unit to perform convolution processing on the input to extract initial shared features between multiple tasks.
[0101] The second low-rank convolution unit is a convolution unit including a low-rank parameter matrix and an original parameter matrix, and is configured to perform convolution processing based on the low-rank parameter matrix and the original parameter matrix. The rank of the low-rank parameter matrix is much smaller than the rank of the original parameter matrix.
[0102] Specifically, the second low-rank convolution unit can convolve the input through the internal low-rank parameter matrix, and at the same time, convolve the input through the internal original parameter matrix, and then fuse the convolution results obtained by the two convolution branches to obtain the initial shared features between multiple tasks.
[0103] By performing convolution processing on the original parameter matrix and the low-rank parameter matrix inside the second low-rank convolution unit, efficient model training can be achieved, and the computational complexity of the model can be reduced, thereby improving the generalization ability of the model.
[0104] After the shared feature convolution module obtains the initial shared features, the initial shared features can be input into the task feature convolution module. The task feature convolution module may include a convolution unit, which further extracts the independent features of each task and the final shared features between multiple tasks from the initial shared features through the convolution unit.
[0105] For example, Figure 6 This is an architecture diagram of a convolution module group provided in an embodiment of the present application, such as Figure 6 As shown in Figure 1, after the input is processed by the shared feature convolution module, it is further processed by the task feature convolution module.
[0106] Among them, the role of independent features is to input into the multi-scale fusion module for task prediction, and the role of final shared features is to facilitate the next convolution module group to continue to extract independent features and shared features, thereby obtaining independent features of other scales, so as to capture the underlying details and high-level semantics at different scales and ensure the accuracy of task prediction.
[0107] In one example, the task feature convolution module includes a third low-rank convolution unit, a fourth low-rank convolution unit, at least one residual module, a second splicing unit, and an output unit;
[0108] The third low-rank convolution unit is used to perform convolution processing and fusion on the initial shared features according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result into the residual module;
[0109] The residual module is used to perform convolution and jump connection processing on the input convolution result, and input the obtained residual result into the second splicing unit;
[0110] The fourth low-rank convolution unit is used to perform convolution processing and fusion on the initial shared features according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result into the second splicing unit;
[0111] A second splicing unit is used to splice the input convolution result and the residual result, and input the obtained second splicing result to the output unit;
[0112] The output unit is used to extract the independent features of each task and the final shared features among multiple tasks from the second splicing result.
[0113] Among them, the third low-rank convolution unit is connected to each residual module, each residual module is connected to the second splicing unit, the fourth low-rank convolution unit is connected to the second splicing unit, and the second splicing unit is connected to the output unit.
[0114] For example, Figure 7 This is an architecture diagram of a task feature convolution module provided in an embodiment of the present application, such as Figure 7 As shown, the initial shared features output by the shared feature convolution module can be input to the third low-rank convolution unit and the fourth low-rank convolution unit at the same time. After the output of the third low-rank convolution unit passes through each residual module, it is spliced together with the output of the fourth low-rank convolution unit through the second splicing unit and finally input to the output unit.
[0115] The third low-rank convolution unit is a convolution unit including a low-rank parameter matrix and an original parameter matrix, and is configured to perform convolution processing based on the low-rank parameter matrix and the original parameter matrix. The fourth low-rank convolution unit is a convolution unit including a low-rank parameter matrix and an original parameter matrix, and is configured to perform convolution processing based on the low-rank parameter matrix and the original parameter matrix. The rank of the low-rank parameter matrix is much smaller than the rank of the original parameter matrix.
[0116] Specifically, the third low-rank convolution unit can convolve the initial shared features of the input through the internal low-rank parameter matrix, and at the same time, convolve the initial shared features of the input through the internal original parameter matrix, thereby fusing the convolution results obtained by the two convolution branches. Similarly, the processing process of the fourth low-rank convolution unit is the same.
[0117] Furthermore, the convolution result obtained by the third low-rank convolution unit passes through each residual module. The residual module can perform convolution and jump connection processing on the input convolution result, and input the obtained residual result into the second splicing unit.
[0118] In some optional embodiments, the residual module includes a fifth low-rank convolution unit, a sixth low-rank convolution unit and a connection unit;
[0119] The fifth low-rank convolution unit is used to perform convolution processing and fusion on the input convolution results according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution results to the sixth low-rank convolution unit and the connection unit;
[0120] The sixth low-rank convolution unit is used to perform convolution processing and fusion on the input convolution results according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution results to the connection unit;
[0121] The connection unit is used to add all the convolution results of the input.
[0122] Among them, the fifth low-rank convolution unit is connected to the sixth low-rank convolution unit and the connection unit, and the sixth low-rank convolution unit is connected to the connection unit.
[0123] For example, Figure 8 This is an architecture diagram of a residual module provided in an embodiment of the present application, such as Figure 8 As shown in the figure, the convolution result obtained by the third low-rank convolution unit is first processed by the fifth low-rank convolution unit. After the processing result is processed by the sixth low-rank convolution unit, it is fused with the processing result of the fifth low-rank convolution unit in the connection unit to achieve the purpose of jump connection.
[0124] Specifically, the fifth low-rank convolution unit can convolve the input through the internal low-rank parameter matrix, and at the same time, convolve the input through the internal original parameter matrix, thereby fusing the convolution results obtained by the two convolution branches. Furthermore, the sixth low-rank convolution unit will perform similar processing on the output of the fifth low-rank convolution unit.
[0125] Furthermore, the connection unit can add the result of processing the fifth low-rank convolution unit to the result of processing the sixth low-rank convolution unit to pass the output of the shallow network directly to the deep network, realizing the cross-layer flow of information and solving problems such as gradient disappearance / explosion, network degradation and feature reuse.
[0126] In the above optional implementation, through the fifth low-rank convolution unit, the sixth low-rank convolution unit and the connection unit, not only can the cross-layer flow between features be achieved and the feature utilization rate be improved, but also the model training efficiency can be further improved, the model calculation complexity can be further reduced, the model generalization ability can be improved, and the overfitting of the model can be reduced.
[0127] After convolution and jump connection processing by the residual module, the residual result obtained will further pass through the second splicing unit, and then the second splicing unit can splice the residual result with the convolution result output by the fourth low-rank convolution unit to obtain a second splicing result and input it into the output unit.
[0128] Furthermore, the output unit may process the second splicing result to obtain independent features of each task and final shared features among multiple tasks.
[0129] In some optional embodiments, the output unit includes a shared feature output layer and a task feature output layer corresponding to each task;
[0130] The shared feature output layer is used to perform convolution processing and fusion on the second splicing result based on the internal shared low-rank parameter matrix and the original parameter matrix to obtain the final shared features between multiple tasks;
[0131] The task feature output layer is used to perform convolution processing and fusion on the second splicing result according to the internal original parameter matrix and the task low-rank parameter matrix of the corresponding task to obtain the independent features of the corresponding task.
[0132] Among them, the output unit may include a shared low-rank parameter matrix (the number can be 2), an original parameter matrix (the number is 1), and a task low-rank parameter matrix of each task (the number of task low-rank parameter matrices for each task can be 2).
[0133] For example, Figure 9 This is a processing diagram of an output unit provided in an embodiment of the present application, such as Figure 9 As shown, taking the number of tasks as three as an example, the number of shared low-rank parameter matrices is 2, the number of task low-rank parameter matrices of each task is 2, and the number of original parameter matrices is 1.
[0134] refer to Figure 9 Specifically, the shared feature output layer contains two convolution branches. The first convolution branch can use the original parameter matrix to perform convolution processing on the second splicing result. The second convolution branch can first use the first shared low-rank parameter matrix to perform convolution processing on the second splicing result, and then use the second shared low-rank parameter matrix to perform convolution processing on this basis. Finally, the results of the two convolution branches are fused to obtain the final shared features between multiple tasks.
[0135] At the same time, for each task corresponding to the task feature output layer, it contains two convolution branches, refer to Figure 9 The first convolution branch can use the original parameter matrix to perform convolution processing on the second splicing result, and the second convolution branch can first use the low-rank parameter matrix of the first task to perform convolution processing on the second splicing result, and then use the second task low-rank parameter matrix to perform convolution processing on this basis. Finally, the results of the two convolution branches are fused to obtain the independent features of the task.
[0136] In the above embodiment, the output unit may include the original matrix shared by each task, as well as the shared low-rank parameter matrix and the task low-rank parameter matrix of each task. The shared feature output layer uses the original matrix and the shared low-rank parameter matrix to extract the final shared features between the tasks. The feature output layer of each task uses the original matrix and the corresponding task low-rank parameter matrix to extract the independent features of each task. While reducing the number of model parameters, it can resolve conflicts between tasks and improve the accuracy of multi-task prediction. Moreover, if the number of tasks needs to be expanded, it is only necessary to add a task feature output layer to the output unit without changing the overall structure of the model.
[0137] In an embodiment of the present application, the task feature convolution module extracts the final shared features between tasks and the independent features of each task through a third low-rank convolution unit, a fourth low-rank convolution unit, at least one residual module, a second splicing unit and an output unit, which can improve the accuracy of feature representation while reducing the number of model parameters, thereby improving the accuracy of task prediction.
[0138] After the task feature convolution module extracts the final shared features between tasks and the independent features of each task, if it is connected to other convolution module groups, it can send the final shared features to other convolution module groups, which will continue to extract independent features and shared features. The processing process can refer to the previous steps and send the independent features of each task to the multi-scale fusion module corresponding to each task. For the last convolution module group, it can only send the independent features of each task to the multi-scale fusion module corresponding to each task.
[0139] It should be noted that, since the resolution of the output is lower than the resolution of the input after the convolution module group extracts features from the input, for example, assuming that the step size of the shared feature convolution module in the convolution module group is 2, and the step size of the remaining convolution units in the convolution module group is 1, then after processing by one convolution module group, the output is a downsampled image that is twice the input; therefore, the number of convolution module groups can be determined according to the resolution of the input target image. If the resolution of the input target image is higher, the number of convolution module groups is greater.
[0140] For example, assuming that a multi-task model is applied to a vehicle-mounted system to perform 2D target detection, 3D target detection, and lane line detection for autonomous driving, the input of the multi-task model is the image captured by the vehicle's onboard camera. The number of convolutional module groups can be set based on the resolution of the image captured by the camera.
[0141] Specifically, after the independent features of different scales of each task enter the multi-scale fusion module corresponding to the task, the multi-scale fusion module can fuse the independent features of all scales, and then obtain the task prediction result through the fusion result.
[0142] For example, the multi-scale fusion module corresponding to each task may include an upsampling unit, a third splicing unit and a task head unit. Figure 10 This is an architecture diagram of a multi-scale fusion module provided in an embodiment of the present application, wherein the upsampling unit is connected to the third splicing unit, and the third splicing unit is connected to the task head unit.
[0143] refer to Figure 10 The upsampling unit can be used to upsample the independent features of the task output by each convolution module group to convert them into features of the same size.
[0144] For example, assuming that the number of convolution module groups is 4, and each convolution module group can downsample the input by 2 times, the independent features of a task output by the four convolution module groups are 4-fold, 8-fold, 16-fold and 32-fold downsampled feature maps, respectively. The upsampling unit can upsample the 8-fold, 16-fold and 32-fold downsampled feature maps to sample them to the size of 4 times the feature map.
[0145] Furthermore, the third splicing unit can fuse the feature maps processed by the upsampling unit, and then the task head unit can use convolution to predict the fused results to obtain the task prediction result.
[0146] The multi-task model provided in the embodiments of the present application can solve the problems of soft sharing and win sharing in existing multi-tasks, enabling the model to achieve sharing between tasks and support the independence of each task with fewer training parameters, resolving the problem of multi-task conflicts and enabling each task to achieve the same prediction effect as a single task. Furthermore, the low-rank convolutional layer in the multi-task model provided in the embodiments of the present application can also be applied to other modules or models containing convolution, and has strong portability, versatility, and ease of use.
[0147] In a specific embodiment, the training of the multi-task model includes the following steps:
[0148] Step 1: Build the initial model and obtain pre-training samples and optimized training samples;
[0149] Step 2: Train the initial model based on the pre-training samples to update all original parameter matrices in the initial model until the pre-training stop condition is reached to obtain the pre-training model;
[0150] Step 3: Freeze all original parameter matrices in the pre-trained model and train the pre-trained model based on the optimized training samples to update all low-rank parameter matrices in the initial model until the optimized training stop condition is reached to obtain a multi-task model.
[0151] Among them, the initial model can include a focal convolution module, at least one convolution module group, and a multi-scale fusion module corresponding to each task; the structure of each module can be referred to the above description and will not be repeated here.
[0152] Specifically, in the above step 1, an initial model including a focal convolution module, at least one convolution module group and various multi-scale fusion modules can be built, and a sample set can be obtained, and the sample set can be divided into pre-training samples and optimized training samples.
[0153] Furthermore, in step 2, the initial model can be trained using pre-training samples. During this training process, the loss function is calculated based on the output results of the initial model. Through the calculated loss value, all the original parameter matrices in the initial model are reversely adjusted, including the original parameter matrices in each low-rank convolution unit and the original parameter matrix in the output unit. This process is repeated until the pre-training stop condition is reached to obtain the pre-trained model.
[0154] Among them, the pre-training stopping condition can be that the number of iterations reaches a preset threshold, or the evaluation indicators of the initial model on the validation set (such as accuracy, precision, recall rate, or F1 score, etc.) reach a preset evaluation threshold, or the loss value of the initial model no longer decreases.
[0155] Furthermore, in step 3, after obtaining the pre-trained model, all original parameter matrices in the pre-trained model, including the original parameter matrices in each low-rank convolution unit and the original parameter matrix in the output unit, can be frozen, and then the pre-trained model can be trained using the optimized training samples. The loss function is calculated based on the output results of the pre-trained model, and all low-rank parameter matrices in the pre-trained model, including the low-rank parameter matrices in each low-rank convolution unit and the task low-rank parameter matrix in the output unit, are reversely adjusted through the calculated loss value. This process is repeated until the optimized training stop condition is reached to obtain a multi-task model.
[0156] Among them, the optimization training stopping condition can be that the number of iterations reaches a preset threshold, or the evaluation indicators of the pre-trained model on the validation set (such as accuracy, precision, recall rate, or F1 score, etc.) reach a preset evaluation threshold, or the loss value of the pre-trained model no longer decreases.
[0157] Through the above steps 1 to 3, model pre-training and optimization training can be achieved. During the pre-training process, some parameters are trained, and during the optimization training process, the trained parameters are frozen, and the low-rank parameter matrix is optimized to reduce the number of training parameters. While keeping the main body of the pre-trained model unchanged, only a small number of parameters need to be fine-tuned to adapt to each task, thereby improving the training efficiency of the model while ensuring the accuracy of the model prediction.
[0158] The multi-task prediction method provided in the embodiment of the present application inputs a target image into a trained multi-task model, performs convolution processing on the target image through the focal convolution module in the model, extracts the initial shared features between multiple tasks through the convolution module group in the model, determines the final shared features between multiple tasks and the independent features of each task based on the initial shared features, and inputs the final shared features into other connected convolution module groups so that the other convolution module groups can extract shared features and independent features at another scale, inputs the independent features of each task into the multi-scale fusion module corresponding to each task, and then fuses the independent features of the corresponding task at different scales through the multi-scale fusion module in the model, and outputs the corresponding task prediction results. This method can extract the shared features between multiple tasks and the independent features of each task, and then fuse the independent features of different scales output by different convolution module groups for task prediction, so that the tasks maintain sharing while also supporting the independence of each task, solving the conflict and imbalance problems between multiple tasks, and improving the accuracy of multi-task prediction. In addition, by extracting shared features, the efficiency of feature extraction between tasks can also be improved, thereby improving the efficiency of multi-task prediction.
[0159] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 11 As shown, the electronic device 400 includes one or more processors 401 and a memory 402 .
[0160] The processor 401 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.
[0161] The memory 402 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may execute the program instructions to implement the multi-task prediction method of any embodiment of the present application described above and / or other desired functions. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage medium.
[0162] In one example, electronic device 400 may further include an input device 403 and an output device 404, which are interconnected via a bus system and / or other connection mechanisms (not shown). Input device 403 may include, for example, a keyboard, a mouse, etc. Output device 404 may output various information to the outside, including warning information, braking force, etc. Output device 404 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.
[0163] Of course, to simplify, Figure 11 Only some of the components related to the present application in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 400 may further include any other appropriate components according to specific application scenarios.
[0164] In addition to the above methods and devices, embodiments of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to perform the steps of the multi-task prediction method provided by any embodiment of the present application.
[0165] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0166] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps of the multi-task prediction method provided by any embodiment of the present application.
[0167] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0168] It should be noted that the terms used in this application are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates an exception, the words "one", "an", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.
[0169] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0170] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. The above is only the preferred implementation method of this application. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of this application, they can also make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of this application.
Claims
1. A multi-task prediction method, characterized in that: include: Acquire the target image; Input the target image into the trained multi-task model to obtain prediction results for each task; The multi-task model includes a focal convolution module, at least one convolution module group, and a multi-scale fusion module corresponding to each task; The focal convolution module is used to perform convolution processing on the target image and input the obtained features into the connected convolution module group; The convolution module group is used to extract the initial shared features between multiple tasks from the input, determine the final shared features between the multiple tasks and the independent features of each task based on the initial shared features, and input the independent features of each task into the multi-scale fusion module corresponding to each task respectively, and input the final shared features into other connected convolution module groups so that other convolution module groups can extract the final shared features between the multiple tasks and the independent features of each task at another scale; The multi-scale fusion module is used to fuse independent features of different scales corresponding to the task and determine the task prediction result based on the fusion result.
2. The method according to claim 1, characterized in that The focal convolution module includes a slicing unit, a first splicing unit and a first low-rank convolution unit; The slicing unit is configured to slice the target image and input each obtained slice into the first stitching unit; The first splicing unit is configured to perform splicing processing on each slice and input a first splicing result obtained into the first low-rank convolution unit; The first low-rank convolution unit is used to perform convolution processing and fusion on the input according to the internal low-rank parameter matrix and the original parameter matrix.
3. The method according to claim 1, characterized in that The convolution module group includes a shared feature convolution module and a task feature convolution module, and the shared feature convolution module includes a second low-rank convolution unit; The shared feature convolution module is used to extract initial shared features between multiple tasks from the input and input the initial shared features into the task feature convolution module; The task feature convolution module is used to extract the independent features of each task and the final shared features between multiple tasks based on the initial shared features; The second low-rank convolution unit is used to perform convolution processing and fusion on the input according to the internal low-rank parameter matrix and the original parameter matrix, so as to extract initial shared features between multiple tasks from the input.
4. The method according to claim 3, characterized in that The task feature convolution module includes a third low-rank convolution unit, a fourth low-rank convolution unit, at least one residual module, a second splicing unit and an output unit; The third low-rank convolution unit is used to perform convolution processing and fusion on the initial shared features according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result into the residual module; The residual module is used to perform convolution and jump connection processing on the input convolution result, and input the obtained residual result into the second splicing unit; The fourth low-rank convolution unit is configured to perform convolution processing and fusion on the initial shared features according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result into the second splicing unit; The second splicing unit is used to splice the input convolution result and the residual result, and input the obtained second splicing result to the output unit; The output unit is configured to extract the independent features of each task and the final shared features among multiple tasks from the second splicing result.
5. The method according to claim 4, characterized in that The residual module includes a fifth low-rank convolution unit, a sixth low-rank convolution unit and a connection unit; The fifth low-rank convolution unit is used to perform convolution processing and fusion on the input convolution result according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result to the sixth low-rank convolution unit and the connection unit; The sixth low-rank convolution unit is used to perform convolution processing and fusion on the input convolution result according to the internal low-rank parameter matrix and the original parameter matrix, and input the obtained convolution result to the connection unit; The connection unit is used to add all convolution results of the input.
6. The method according to claim 4, characterized in that The output unit includes a shared feature output layer and a task feature output layer corresponding to each task; The shared feature output layer is used to perform convolution processing and fusion on the second splicing result according to the internal shared low-rank parameter matrix and the original parameter matrix to obtain the final shared features between multiple tasks; The task feature output layer is used to perform convolution processing and fusion on the second splicing result according to the internal original parameter matrix and the task low-rank parameter matrix of the corresponding task to obtain independent features of the corresponding task.
7. The method according to any one of claims 2 to 5, characterized in that: The low-rank convolution unit includes a low-rank convolution layer, a normalization layer, and an activation layer; The low-rank convolution layer is used to perform convolution processing on the input according to the internal first low-rank parameter matrix, the second low-rank parameter matrix and the original parameter matrix, and fuse the obtained convolution results, and input the obtained fusion results into the normalization layer; The normalization layer is used to normalize the input and input the normalized result to the activation layer; The activation layer is used to perform nonlinear transformation on the input.
8. The method according to any one of claims 3 to 5, characterized in that: The training of the multi-task model includes: Build the initial model and obtain pre-training samples and optimized training samples; Training the initial model based on the pre-training samples to update all original parameter matrices in the initial model until a pre-training stop condition is reached, thereby obtaining a pre-training model; Freeze all original parameter matrices in the pre-trained model, and train the pre-trained model based on the optimized training sample to update all low-rank parameter matrices in the initial model until the optimization training stop condition is reached, thereby obtaining the multi-task model.
9. An electronic device, characterized in that: The electronic device comprises: processor and memory; The processor is configured to execute the steps of the multi-task prediction method according to any one of claims 1 to 8 by calling the program or instructions stored in the memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, which enables a computer to execute the steps of the multi-task prediction method according to any one of claims 1 to 8.