Data annotation model training method and device, data annotation method and device and vehicle

By building a data labeling model, using the combination of non-homologous data sets and target prediction modules, the problem of flexibility and efficiency in training of multi-task learning models is solved, and the full feature learning of autonomous driving perception data is realized, and training efficiency and flexibility are improved.

CN120014343APending Publication Date: 2025-05-16CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510087513.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the training samples selected by the multi-task learning model during training have low flexibility and low training efficiency.

Method used

By building a data labeling model, using non-homologous data sets as training samples, according to the task type to which the elements whose real labels belong, the training samples are marked, feature information is input to the corresponding target prediction module for learning, and model parameters are updated.

Benefits of technology

In the case of non-homologous data sets being implemented, the data annotation module can learn all the features of autonomous driving perception data at one time, improving the flexibility of training samples, reducing training costs, and improving training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014343A_ABST
    Figure CN120014343A_ABST
Patent Text Reader

Abstract

The invention relates to a data annotation model training method, a data annotation method, a data annotation device and a vehicle, and the method comprises the steps: obtaining a training sample which comprises an image and a real label annotated on the image; inputting the training sample into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model; and inputting the feature information into a target prediction module corresponding to the target task type, and updating model parameters of the data annotation model based on a prediction result output by the target prediction module and the real label. Under the condition that the input training sample is a non-homologous data set, all features of the automatic driving perception data can be learned, flexible selection of the training sample is realized, the training cost is reduced, and the training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data annotation model training method, a data annotation method, a device and a vehicle. Background Art

[0002] The key elements of autonomous driving technology include perception, decision-making, and execution. Perception refers to helping the intelligent driving system perceive and understand lane lines, drivable areas, pedestrians, vehicles, and obstacles in the environment, so that the intelligent driving system can make accurate decisions and control the vehicle execution system to execute the decision content. Accurate perception can make vehicle control safer and more comfortable.

[0003] Perception includes the perception of vehicles, pedestrians, obstacles, drivable areas, lane lines and other elements. The data is complex and diverse, especially in 3D perception and motion planning, which rely on deep learning. Since perception includes the perception of elements of different natures, different learning tasks are needed to complete the learning of all element features.

[0004] In the related art, different elements in different perceptions are learned and inferred through a multi-task learning model. When training a multi-task learning model, in order to ensure that the multi-task learning model can learn all elements, all samples in the training sample set are required to include all elements and labels of the corresponding elements. This model training has low flexibility and low model training efficiency. Summary of the invention

[0005] One of the purposes of the present invention is to provide a data labeling method to solve the problems of low flexibility and low training efficiency in the selection of training samples during training of the multi-task learning model in the prior art; the second purpose is to provide a data labeling method; the third purpose is to provide a data labeling device; the fourth purpose is to provide a vehicle.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] In a first aspect, a data annotation model training method includes:

[0008] Acquire a training sample, wherein the training sample includes an image and a true label marked on the image;

[0009] Inputting the training sample into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model;

[0010] The feature information is input into a target prediction module corresponding to the target task type, and the model parameters of the data annotation model are updated based on the prediction result output by the target prediction module and the true label.

[0011] According to the above technical means, a data labeling model is constructed based on multi-task learning, and non-homologous data sets are used as training samples. When the data labeling model is trained, during the forward propagation of the input training samples, the task type corresponding to the elements labeled with real labels in the training samples is identified, and the target prediction module into which the training samples flow is determined. Then, the training samples are detected based on the target prediction module, and the model parameters of the data labeling model are updated based on the prediction results output by the target prediction module to complete the training of the data labeling model. That is, in the model training process, for non-homologous training samples, the task type corresponding to the elements labeled with real labels is input into the corresponding prediction module branch instead of inputting into all prediction modules, thereby avoiding the inability of some prediction modules to learn the elements without real labels, and realizing that in the case of non-homologous data sets, the data labeling module can also learn all the features of the autonomous driving perception data at one time, and the training samples can be flexibly selected to reduce the training cost and improve the training efficiency.

[0012] Further, the step of inputting the feature information into a target prediction module corresponding to the target task type, and updating the model parameters of the data annotation model based on the prediction result output by the target prediction module and the true label, includes:

[0013] Determining a target task type according to a task type label associated with the feature information, wherein, when preprocessing the image, adding the task type label according to the task type to which the element in the image that has been annotated with the real label belongs;

[0014] Determine a corresponding target prediction module according to the target task type;

[0015] The feature information is input into the target prediction module, and the model parameters of the data annotation model are updated based on the prediction result output by the target prediction module and the true label.

[0016] According to the above technical means, by adding task type labels, the task types to which training samples in non-homologous data sets belong are distinguished, which facilitates the identification of training samples and then facilitates the flow of training samples into corresponding prediction modules, so that non-homologous data sets can be input into the model and can also satisfy the model's learning of all features of all tasks.

[0017] Further, updating the model parameters of the data annotation model based on the prediction result output by the target prediction module and the true label includes:

[0018] Determine the training loss between the prediction result output by the target prediction module and the true label based on the loss function corresponding to the target prediction module;

[0019] Update a first model parameter of the target prediction module according to the training loss of the target prediction module;

[0020] The second model parameters of the feature extraction module are updated according to the training losses of all the target prediction modules, and the model parameters of the data annotation model include the first model parameters and the second model parameters.

[0021] According to the above technical means, the first model parameters of the prediction module and the second model parameters of the feature extraction module are updated separately based on the training loss, while the shared feature extraction module is updated in a cumulative manner. This method meets the learning requirements of learning all features of multiple task types at one time in the scenario where non-homologous data sets are used as training samples, and the training effect is good.

[0022] Further, updating the first model parameter of the target prediction module according to the training loss of the target prediction module includes:

[0023] Calculating a first gradient of a training loss of the target prediction module with respect to the first model parameter;

[0024] Calculating a loss weighted value of the target prediction module according to the first gradient magnitude;

[0025] The first model parameter is updated according to the loss weight value and the first gradient.

[0026] According to the above technical means, in multi-task learning, the update of the first model parameters of each prediction module is balanced through the loss weight value, ensuring that the learning speed of the multi-task data labeling model is close and reaches simultaneous convergence.

[0027] Further, updating the second model parameter of the feature extraction module according to the training losses of all the target prediction modules includes:

[0028] Respectively calculating the second gradients of the training losses of all the target prediction modules with respect to the second model parameters;

[0029] According to all the second gradients, obtaining a total gradient of the second model parameter;

[0030] The second model parameters are updated according to the total gradient.

[0031] According to the above technical means, the corresponding second model parameters of all prediction modules are updated through the training losses, ensuring that these shared parts can serve the needs of multiple tasks at the same time, and the model can share knowledge between different tasks to improve the overall performance.

[0032] Furthermore, the feature extraction module includes a backbone network and a neck network, and the feature extraction module based on the data annotation model extracts feature information of the image, including:

[0033] In the backbone network of the data annotation model, feature extraction is performed on the image to obtain features of different scales;

[0034] In the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image.

[0035] According to the above technical means, in the neck network of the data annotation model, the different scale features extracted by the backbone network are fused based on the dense feature fusion pyramid module, which can obtain richer feature information, improve the model learning efficiency, and increase the model training speed.

[0036] Furthermore, in the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image, including:

[0037] In the neck network of the data annotation model, the following fusion process is performed on the different scale features extracted by the backbone network based on the dense feature fusion pyramid module:

[0038] Based on the top-down path, the current features of the previous layer of the backbone network are sequentially fused into the current features of the next layer until all features are fused into the last layer of features;

[0039] Based on the bottom-up path, the current features of the next layer and the input features of the previous layer are fused into the current features of the previous layer in sequence, until all features are fused into the top layer features;

[0040] The feature information of the image is obtained according to the fused current features of each layer.

[0041] According to the above technical means, in the process of bottom-up path fusion, in addition to the fusion of the front and back layer features, interlayer connections (also jump connections) are added. The interlayer connections connect the original feature nodes of the same layer to the output feature nodes, thereby fusing more features without increasing more costs and improving the fusion feature level.

[0042] Further, obtaining feature information of the image according to the fused current features of each layer includes:

[0043] Using the fused current features of each layer as input features of each layer, and repeating the top-down path and the bottom-up path to fuse the features of each layer;

[0044] When the number of repetitions reaches N, feature information of the image is obtained according to the current features of each layer after fusion, and N is greater than or equal to 1.

[0045] According to the above technical means, since each bidirectional (top-down and bottom-up) path is regarded as a feature network layer and then the same layer is repeated multiple times, information fusion can be enhanced and higher-level feature fusion can be achieved.

[0046] In a second aspect, a data annotation method includes:

[0047] Inputting the acquired image into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model;

[0048] Inputting the feature information into prediction modules corresponding to each task type respectively, and obtaining all the annotation information of the image based on the prediction results output by all the prediction modules;

[0049] Wherein, the data annotation model is obtained through the training method described in the first aspect above.

[0050] According to the above technical means, independent image processing task models are merged into a single model that uses a shared backbone network to extract image features and multiple task prediction modules to output results. This increases the diversity of output attributes for the automatic labeling system, while reducing the total number of model parameters and the computing resources occupied, thereby improving the model's reasoning speed.

[0051] Furthermore, the feature extraction module includes a backbone network and a neck network, and the feature extraction module based on the data annotation model extracts feature information of the image, including:

[0052] In the backbone network of the data annotation model, feature extraction is performed on the image to obtain features of different scales;

[0053] In the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image.

[0054] Furthermore, in the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image, including:

[0055] In the neck network of the data annotation model, the following fusion process is performed on the different scale features extracted by the backbone network based on the dense feature fusion pyramid module:

[0056] Based on the top-down path, the current features of the previous layer of the backbone network are sequentially fused into the current features of the next layer until all features are fused into the last layer of features;

[0057] Based on the bottom-up path, the current features of the next layer and the input features of the previous layer are fused into the current features of the previous layer in sequence, until all features are fused into the top layer features;

[0058] The feature information of the image is obtained according to the fused current features of each layer.

[0059] A training device for a data annotation model, comprising:

[0060] An acquisition unit, configured to acquire a training sample, wherein the training sample includes an image and a real label marked on the image;

[0061] A feature processing unit, used for inputting the training sample into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model;

[0062] A training unit is used to input the feature information into a target prediction module corresponding to the target task type, and update the model parameters of the data annotation model based on the prediction result output by the target prediction module and the true label.

[0063] A data labeling device, comprising:

[0064] A feature extraction unit, used to input the acquired image into a data annotation model, and extract feature information of the image based on a feature extraction module of the data annotation model;

[0065] The prediction unit is used to input the feature information into the prediction modules corresponding to each task type respectively, and obtain all the annotation information of the image based on the prediction results output by all the prediction modules.

[0066] An electronic device comprising a memory and a processor;

[0067] The memory is used to store computer programs / instructions; the processor is used to implement the method described in the first aspect / second aspect according to the computer programs / instructions stored in the memory.

[0068] A vehicle, comprising: a memory, a processor;

[0069] The memory is used to store computer programs / instructions; the processor is used to implement the method described in the second aspect according to the computer programs / instructions stored in the memory.

[0070] A computer-readable storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a processor, it is used to implement the method described in the first aspect / second aspect above.

[0071] A computer program product comprises a computer program, wherein the computer program is used to implement the method described above when executed by a processor.

[0072] Beneficial effects of the present invention:

[0073] (1) The present invention provides a multi-task learning model. During the training process of the model, the training sample set can be a non-homologous data set. For non-homologous training samples, the training samples are input into the prediction module corresponding to the task type to which the elements annotated with the real labels belong for learning. By learning multiple sets of training samples, the learning of all task features can be completed. In the case of non-homologous data sets, the data annotation module can also learn all the features of the autonomous driving perception data at one time. The training samples can be flexibly selected, which reduces the training cost and improves the training efficiency.

[0074] (2) In the present invention, a backbone network is used to extract image features, and a neck network is used to fuse the features. During the fusion process, based on a dense feature fusion pyramid network, features with richer information, more depth information, and higher levels can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 A schematic diagram of the architecture of a data annotation model provided in an embodiment of the present application;

[0076] Figure 2 A flowchart of a data annotation model training method provided in one embodiment of the present application;

[0077] Figure 3 A flowchart of a data annotation model training method provided in another embodiment of the present application;

[0078] Figure 4 A flowchart of a data annotation model training method provided in yet another embodiment of the present application;

[0079] Figure 5 A schematic diagram of the structure of a shared backbone network provided in an embodiment of the present application;

[0080] Figure 6 A schematic diagram of the structure of the swin-transformer provided in the embodiment of the present application;

[0081] Figure 7 A schematic diagram of the fusion path of the dense feature fusion pyramid module provided in an embodiment of the present application;

[0082] Figure 8 A schematic diagram of a data annotation method according to an embodiment of the present invention;

[0083] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0084] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.

[0085] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0086] In the field of autonomous driving, perception systems usually use deep learning to learn and infer various elements in perception data. Since the perception system of autonomous driving needs to perceive multiple elements, such as pedestrians, obstacles, vehicles, lane lines, drivable areas, etc., these elements can be learned in the following two ways:

[0087] For example, one way is to build multiple models separately, such as vehicle detection model, pedestrian detection module, obstacle detection model, drivable area segmentation model, and lane segmentation model. When training the model, you can annotate the real labels of 2D images by major categories, build non-homologous data sets (training sample sets), and train each model separately. This way of building multiple models does not require that all samples in the training sample set include all requirements and labels of corresponding elements, but the output efficiency of this model detection result is high.

[0088] Another way is to build a multi-task learning model, which can complete the learning and reasoning of multiple tasks such as vehicle detection, pedestrian detection, obstacle detection, drivable area segmentation, and lane line segmentation. This way of building a multi-task learning model can improve the output efficiency of the model detection results, but when training this multi-task learning model, in order to ensure that the multi-task learning model can learn all elements, all samples in the training sample set are required to include all elements and the labels of the corresponding elements (that is, the same source data set is used as the training sample). This model training has low flexibility and low model training efficiency.

[0089] It should be noted that the above-mentioned homologous dataset refers to a training sample set, each training sample includes all elements, and each element has a corresponding true label. Taking the perception image as a training sample as an example, all training samples include pedestrians, vehicles, obstacles, drivable areas, lane lines, etc., and these elements are annotated in the image. Non-homologous datasets refer to a training sample set, in which the training samples may include some elements and / or some elements have corresponding true labels (only labels of a single task or some tasks are annotated in the image). Taking the image as a training sample as an example, training sample 1 includes pedestrians, vehicles, obstacles, drivable areas, lane lines, pedestrians’ true labels and lane lines’ true labels; training sample 2 includes vehicles, drivable areas, lane lines and vehicles’ true labels, drivable areas’ true labels, lane lines’ true labels, etc. That is, in non-homologous datasets, it is not required that all training samples in the training sample set include all elements and labels of corresponding elements.

[0090] This application proposes a training method for a data annotation model based on multi-task learning. When the training sample set is a non-homologous data set, it can also meet the learning of all tasks corresponding to multiple elements. The requirements for the training sample set are relatively flexible. As long as some real labels are marked in the sample, it can be used as a training sample, which reduces the cost of obtaining training samples and improves the efficiency of model training.

[0091] Figure 1 The schematic diagram of the architecture of the data annotation model provided in the embodiment of the present application is as follows: Figure 1 As shown, the data annotation model constructed in the embodiment of the present application includes an encoder and multiple decoders, wherein the multiple decoders share one encoder, and the multiple decoders are respectively used to process different specific tasks.

[0092] Exemplarily, the encoder includes a backbone network backbone and a neck network neck, wherein the backbone network backbone is used to extract features of an input image to obtain features of different scales, and the neck network neck is used to fuse features of different scales output from the backbone network backbone to obtain feature information of the image.

[0093] For example, in the field of autonomous driving, three specific decoders can be set according to the elements of the perception data, such as a target detection decoder, a drivable area segmentation decoder, and a lane segmentation decoder. Among them, there are three target detection decoders, which are used to detect vehicles, pedestrians, and obstacles respectively. Therefore, the data annotation model provided in this embodiment includes five decoders, namely, three target detection decoders, one drivable area segmentation decoder, and one lane segmentation decoder.

[0094] In this embodiment, there are no complex and redundant shared modules between different decoders, which reduces computational consumption and allows the network to be easily trained end-to-end.

[0095] Figure 2 A flow chart of a data annotation model training method provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, this embodiment trains the above data annotation model, including:

[0096] S201, obtaining a training sample, where the training sample includes an image and a true label marked on the image;

[0097] For example, if the data annotation model is applied to the annotation of perception data of an autonomous driving system, the training samples are images. If it is applied to other fields or systems, the training samples can be other types of data, which are equivalent to images.

[0098] As an example, training samples are obtained from a data set that has been collected for autonomous driving system development. The data set is collected by vehicles in various cities across the country, and the data set includes panoramic scenes captured by multiple cameras installed around the vehicle.

[0099] The number of data sets can be adjusted dynamically according to demand or training conditions. Taking a data set with 72w frames of images and corresponding multi-task annotations as an example, it can be divided into a training set of 70w images, a validation set of 1w images, and a test set of 1w images to complete the training and testing of the data annotation model.

[0100] It should be noted that the above datasets can be non-homologous datasets, that is, the image can be annotated with only one type of task label, or with multiple types of task labels, or the image can include one or more of the element features. Among them, the multiple tasks include target detection tasks and segmentation tasks. The target detection tasks include three tasks: pedestrians, vehicles, and obstacles. The segmentation tasks include two tasks: lane line segmentation and drivable area segmentation. The target detection task includes pedestrian labels, obstacle labels, and vehicle labels (including all types of vehicles, such as cars, buses, trucks, etc.); the segmentation task includes lane line labels and drivable area labels.

[0101] The above dataset has geographical, environmental, and weather diversity, so algorithms trained on it are robust enough to be transferred to new environments.

[0102] As an example, after obtaining the training samples, the training samples are preprocessed. The preprocessing includes image cropping, format conversion, filtering, etc. For example, during the training and testing process, the size of the 3840*2160 original image is scaled to 960*540 according to a certain ratio, and then cropped to a size of 960*512 as the input of the model.

[0103] As an example, since the training samples are non-homologous data sets, when preprocessing the training samples, task type labels are added according to the task types of the elements in the image that have been annotated with real labels, so that each training sample is associated with a task type. For example, if the real labels annotated in training sample 1 include pedestrians, then a pedestrian detection label is added to training sample 1; if the real labels annotated in training sample 2 include lane lines, then a lane line segmentation label is added to training sample 2; if the real labels annotated in training sample 3 include obstacles and drivable areas, then an obstacle detection label and a drivable area segmentation label are added to training sample 3. By adding task type labels, the learning tasks that need to be performed corresponding to non-homologous data can be distinguished, so that the model can distinguish the branch to which each training sample belongs.

[0104] S202, inputting the training sample into the data annotation model, and extracting feature information of the image based on the feature extraction module of the data annotation model;

[0105] In this step, the training samples obtained above (specifically, the training set) are input into the constructed data annotation model. The feature extraction module of the data annotation model will perform feature extraction on the training samples, extracting all or part of the features in the training samples (for example, some features that have been marked with real labels).

[0106] Exemplarily, the feature extraction module in this step includes an encoder, and the feature extraction module can extract feature information of the image in a variety of ways, which are not listed here one by one.

[0107] S203: Input the feature information into a target prediction module corresponding to the target task type, and update the model parameters of the data annotation model based on the prediction result output by the target prediction module.

[0108] In this step, the data annotation module includes a shared encoder and multiple decoders, where each decoder is used as a prediction module to learn a specific task and then output the prediction result of the specific task. The data annotation model is a multi-task learning model, so it includes multiple prediction modules.

[0109] Since the training samples are non-homologous data sets, the input training samples may only be labeled with labels of some task types, so these training samples cannot be input into the prediction modules corresponding to all task types in sequence. For example, if only pedestrians are labeled on the image, the image cannot be input into the prediction module corresponding to the segmentation task, or the image cannot be input into the prediction module corresponding to the obstacle detection task or vehicle detection task. Otherwise, the image cannot be predicted based on these prediction modules, which will affect the training results of the model.

[0110] In the embodiment of the present application, the feature information is input into the prediction module corresponding to the target task type to which it belongs, so as to predict the feature information, and then update the model parameters based on the prediction results, and the extracted feature information will not be input into the prediction modules of other task types. For example, if only pedestrians are marked in training sample 1, the characteristic information extracted from training sample 1 will be input into the prediction module corresponding to the pedestrian detection task; if lane lines are marked in training sample 2, the feature information extracted from training sample 2 will be input into the prediction module corresponding to the lane line segmentation task; if obstacles and drivable areas are marked in training sample 3, the feature information extracted from training sample 3 will be input into the prediction module corresponding to the obstacle detection task and the prediction module corresponding to the drivable area segmentation task respectively. It should be noted that the prediction module responsible for the output of the target task type is the target prediction module.

[0111] In an optional implementation, the feature information is input into a target prediction module corresponding to the target task type, and the model parameters of the data annotation model are updated based on the prediction result output by the target prediction module, including:

[0112] A1: Determine the target task type based on the task type label associated with the feature information;

[0113] A2: Determine the corresponding target prediction module according to the target task type;

[0114] As in step S201 above, when preprocessing the image (training sample), a task type label is added to the image according to the task type to which the elements in the image that have been annotated with real labels belong. Therefore, the feature information of the image is associated with the task type label. During the model training process, the target task type corresponding to the image can be determined based on the task type label associated with the feature information, so that the feature information of the image directly flows into the target prediction module corresponding to the target task type.

[0115] For example, if pedestrians are marked in the image, the task type label corresponding to the image is a pedestrian detection task; if lane lines are marked in the image, the task type label corresponding to the image is a lane line segmentation task.

[0116] A3: Input the feature information into the target prediction module, and update the model parameters of the data annotation model based on the prediction results output by the target prediction module.

[0117] The feature information is input into the corresponding target prediction module, which can then learn and predict based on the feature information, output the prediction results, and update the model parameters of the data annotation model according to the prediction results, so that the data annotation model tends to converge, completing the learning of the data annotation model.

[0118] In this embodiment, a data annotation model is constructed based on multi-task learning, and non-homologous data sets are used as training samples. When the data annotation model is trained, during the forward propagation of the input training samples, the task type corresponding to the elements labeled with real labels in the training samples is identified to determine the target prediction module into which the training samples flow. Then, the training samples are detected based on the target prediction module, and the model parameters of the data annotation model are updated based on the prediction results output by the target prediction module to complete the training of the data annotation model. That is, during the model training process, for non-homologous training samples, the task type corresponding to the elements labeled with real labels is input into the corresponding prediction module branch instead of inputting into all prediction modules. This avoids the situation where some prediction modules cannot learn the elements without real labels, and in the case of non-homologous data sets, the data annotation module can also learn all the features of the autonomous driving perception data at one time. The training samples can be flexibly selected to reduce the training cost and improve the training efficiency.

[0119] Figure 3 A flowchart of a data annotation model training method provided in another embodiment of the present application is shown in FIG. Figure 3 As shown, based on the above embodiment, the specific implementation process of S203 is described in detail, and S203 includes:

[0120] S301, determining the training loss between the prediction result output by the target prediction module and the true label based on the loss function corresponding to the target prediction module;

[0121] Based on the above embodiments, it can be seen that the data annotation model is a multi-task learning model, which includes multiple prediction modules (head). In the field of autonomous driving, based on the type of perception data, three tasks are determined, of which target detection tasks include three, so five prediction modules are required, that is, five decoders, so the multi-task loss includes five parts, three detection losses and two segmentation losses.

[0122] As an example, based on the task type, three detection loss functions, one lane segmentation loss function, and one drivable area segmentation loss function are configured.

[0123] Exemplarily, the detection loss is constructed according to the classification loss, target loss and bounding box loss as follows:

[0124]

[0125] in, and It is the focal loss, which is used to reduce the loss of easy-to-classify samples, so that the network focuses on difficult-to-classify samples. For penalty classification, Used to penalize the confidence of the prediction box. yes Represents the distance, overlap, scale similarity, and aspect ratio between the predicted box and the true value.

[0126] Exemplarily, the drivable area segmentation loss function is constructed according to Tversky loss and Focal loss, as shown in the following formula:

[0127]

[0128] in, for Tversky's loss; It is the focal loss.

[0129] For example, based on the consideration of sparse lane categories, a lane segmentation loss function is constructed according to Tversky loss, Focal loss and IOU loss, as shown in the following formula:

[0130]

[0131] in, for Tversky's loss; For Focal loss; is the IOU loss.

[0132] In the segmentation loss, Tversky loss can solve the problem of category imbalance well, and Focal loss is used to reduce the classification error between pixels.

[0133] As an optional implementation, the above Tversky loss can be calculated by the following formula:

[0134]

[0135] Focal loss can be calculated by the following formula:

[0136]

[0137] In the above formula, TP P (c) To detect the correct positive training samples, FNP (c) is the negative training sample for detection error, FP P (c) is a positive training sample for detecting errors, p n (c) Predict the probability that pixel n belongs to class c, g n (c) The true value pixel n belongs to class c, C is the number of classes, and N is the pixel of the input image.

[0138] As an optional implementation, the above TP P (c) FN P (c) and FP P (c) can be calculated by the following formulas:

[0139]

[0140] In the above way, the loss function corresponding to each prediction module is configured in the data standard model. During the data annotation model training process, the target prediction module is determined according to the training samples, and after the training samples are input into the target prediction module, the training loss between the prediction result output by the target prediction model and the true label is calculated through the loss function.

[0141] Since the training samples are non-homologous data sets, after the training samples are input into the data annotation model, all prediction modules are target prediction modules. However, the prediction modules do not process all training samples, but only predict the corresponding training samples according to the required task types corresponding to the real labels annotated by the training samples.

[0142] Therefore, the target prediction module includes at least one of the prediction modules for pedestrian detection tasks, vehicle detection tasks, obstacle detection tasks, lane segmentation tasks, and drivable area segmentation tasks. Each prediction module calculates the training loss between the prediction result and the corresponding true label according to the corresponding loss function.

[0143] S302, updating a first model parameter of the target prediction module according to the training loss of the target prediction module;

[0144] S303, updating the second model parameters of the feature extraction module according to the training losses of all target prediction modules;

[0145] It should be noted that the model parameters of the data annotation model include the first model parameters and the second model parameters, wherein the first model parameters are the model parameters in each prediction module, and the second model parameters are the model parameters in the backbone network (feature extraction module) shared by each prediction module, such as the model parameters of the backbone and neck networks. The above-mentioned first model parameters and second model parameters are only used to distinguish the model parameters of the prediction module (head) from the model parameters of the shared backbone network (backbone and neck), and no other restrictions are made.

[0146] In this embodiment, for the model parameters (first model parameters) of the prediction module corresponding to each task, the corresponding first model parameters are updated based on the training loss of each prediction module. As for the model parameters (second model parameters) of the shared backbone network, since they are shared by all prediction modules, the corresponding second model parameters are updated through the training loss of all prediction modules, ensuring that these shared parts can serve the needs of multiple tasks at the same time, and the model can share knowledge between different tasks to improve overall performance.

[0147] As an example, the implementation method of updating the first model parameter of the target prediction module according to the training loss of the target prediction module may be:

[0148] First, the first gradient of the training loss of the target prediction module with respect to the first model parameter is calculated based on back propagation; then the loss weighted value of the target prediction module is calculated according to the first gradient size; and then the first model parameter is updated according to the loss weighted value and the first gradient.

[0149] As an example, the loss weighted value of the target prediction module can be calculated based on the gradient size using the following formula:

[0150]

[0151] Where n is the task vector, i is the target task type, is the first gradient; the loss weight value of all prediction modules = 1.

[0152] It is understandable that the loss weighted value reflects the loss size of the target prediction module relative to all prediction modules. In this embodiment, the first gradient of the prediction module is weighted. When the loss weighted value is larger, the first gradient scaling is larger, which increases the update of the first model parameters of the corresponding prediction module; when the loss weighted value is smaller, the first gradient scaling is smaller, which reduces the update of the first model parameters of the corresponding prediction module. The update of the first model parameters of each prediction module is balanced by the loss weighted value to ensure that the learning speed of the multi-task data annotation model is close and converges simultaneously.

[0153] As an example, the implementation method of updating the second model parameters of the feature extraction module according to the training losses of all target prediction modules may be:

[0154] The second gradients of the training losses of all target prediction modules with respect to the second model parameters are calculated based on back propagation respectively; the total gradient of the second model parameters is obtained according to all the second gradients; and the second model parameters are updated according to the total gradient.

[0155] Exemplarily, the data standard model includes 5 prediction modules, and all 5 prediction modules may be target prediction modules. During the prediction module training process, the second gradient of the second model parameter of the shared backbone network can also be calculated based on the training loss. Since the shared backbone network is shared by all prediction modules, the second model parameter of the shared backbone network is updated based on the training loss of all prediction modules.

[0156] As an example, all the second gradients are accumulated, the accumulated values ​​are used as the total gradients of the second model parameters with respect to the shared backbone network, and then the second model parameters are updated based on the total gradients.

[0157] In this example, the gradient accumulation-based method helps to achieve effective parameter sharing in multi-task learning, enabling the model to transfer and share knowledge between different tasks, thereby improving the generalization ability and efficiency of the model.

[0158] Exemplarily, the gradient update of the shared backbone network is as follows:

[0159]

[0160] Among them, θ sh is the total gradient after the shared backbone network update, θ sh is the total gradient before the shared backbone network is updated, The second gradient corresponding to the training loss of the vehicle detection task; The second gradient corresponding to the training loss of the pedestrian detection task; The second gradient corresponding to the training loss of the obstacle detection task; The second gradient corresponding to the training loss of the lane segmentation task; The second gradient corresponding to the training loss of the drivable area segmentation task.

[0161] Based on the above, in the embodiments of the present application, after calculating the training loss, the model parameters are updated based on gradient descent in the prediction module or the feature extraction module.

[0162] As an optional implementation, the pytorch deep learning framework is used, and the above data annotation model is built based on the OpenMMLabDetection code. In the process of the experiment, the alternating optimization algorithm is used to gradually train the model. In each step, the model can focus on one or more related tasks without considering irrelevant tasks. First, only the prediction module (such as the detection head) corresponding to the encoder and detection tasks is trained. Then, the prediction modules (such as the segmentation head) corresponding to the two segmentation tasks are trained while the encoder and detection head are fixed. Finally, the entire network is jointly trained for all these tasks to achieve better output effects.

[0163] Based on the above model training method, the first model parameters of the prediction module and the second model parameters of the feature extraction module are updated separately based on the training loss, while the shared feature extraction module is updated in an additive manner. This method meets the requirements of learning all features of multiple task types at one time in the scenario where non-homologous data sets are used as training samples, and the training effect is good.

[0164] Figure 4 A flowchart of a data annotation model training method provided in another embodiment of the present application is shown in FIG. Figure 4 As shown, based on the above embodiment, the specific implementation process of S202 is described in detail, and S202 includes:

[0165] S401, extracting features of the image in the backbone network of the data annotation model to obtain features of different scales;

[0166] It is understandable that if Figure 5 As shown, the feature extraction module is composed of a shared backbone network, which includes a backbone network backbone and a neck network neck, wherein the backbone network is used to extract the features of the image in the input training sample; the backbone network and the neck network are connected, and the features extracted by the backbone network are input into the neck network, and the neck network performs feature fusion.

[0167] In the field of autonomous driving, the images captured by perception have high resolution and involve relatively intensive visual tasks. Therefore, in this step, swin-transformer is designed as the backbone network.

[0168] Optionally, based on the backbone network of the swin-transformer structure, features are extracted in the following ways:

[0169] Based on the sliding window, the image in the input training sample is divided into blocks to obtain multiple window areas; based on the self-attention mechanism of the ransformer, the features in each window area are modeled, the features in each window area are extracted, and the local features are fused into global features through window area fusion.

[0170] In this step, the sliding window size is small, with a fixed and smaller size as the input sequence, and a fixed and smaller size as the input sequence. The attention mechanism is used for modeling, focusing on global information. In multi-task learning, the target detection task requires high-quality global and local information to accurately locate and identify, and the image segmentation task also has high requirements for pixel classification and local and global cognition. Therefore, the embodiment of the present application can extract both global features and local features through the above method.

[0171] Optionally, swin-transformer processes the features of the image hierarchically, using a pyramid structure similar to CNN, and realizes multi-scale feature extraction from local to global by reducing the resolution of the feature map layer by layer. This hierarchical structure makes swin-transformer more flexible and robust when processing images of different scales.

[0172] The following is a specific example to illustrate the process of extracting features from an image using the backbone network in the embodiment of the present application:

[0173] Figure 6 The schematic diagram of the structure of swin-transformer is shown in Figure 6 As shown, the input RGB image is first segmented into non-overlapping patches (window areas) through the patch segmentation module. Each patch is regarded as a "marker". The features in each patch are flattened from left to right and from top to bottom in the original RGB image, and a one-dimensional vector is output. For example, using a 4×4 patch size, the feature dimension of each patch is 4×4×3=48. A linear embedding layer is used on this raw value feature to project it to an arbitrary dimension (defined as C).

[0174] A swin-transformer block with self-attention calculation is applied on these patch labels, and the number of labels contained in these swin-transformer blocks is Together with a linear embedding layer as the first stage.

[0175] The backbone network consists of multiple layers, and the multiple layers are hierarchical representations of the backbone network. During the feature extraction process, the hierarchical representation of the backbone network is generated in the following ways:

[0176] As the network goes deeper, the number of markers is reduced by patch merging layers. The first patch merging layer concatenates the features of each set of 2×2 adjacent patches and applies a linear layer on the concatenated 4C-dimensional features. This reduces the number of markers by 2×2=4 (downsampling the resolution by a factor of 2), and the swin-transformer block is then applied to transform the features, keeping the resolution at The patch merging and feature transformation of these modules are used as the second stage. This process is repeated twice as the third and fourth stages, and the output feature sizes are The first to fifth stages produce five levels of representation, respectively, with the same feature map resolution as a typical convolutional network.

[0177] S402, in the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on the dense feature fusion pyramid module to obtain feature information of the image.

[0178] It should be noted that the backbone network processes the features of the image in a hierarchical manner, wherein each layer of features includes features of different scales.

[0179] In this step, the neck network is used to fuse the features of different scales generated by the backbone network to achieve feature fusion of different scales and depths. Specifically, the neck network fuses the features represented by different levels in the backbone network.

[0180] In the neck network, fusing features of different scales can improve the performance of the multi-task learning model. For example, shallow features have higher resolution and contain more position information and texture information, but have lower semantics and more noise, while deep features have stronger semantic information, but have low resolution and poor detail perception. Fusion of features at different levels can make the extracted feature information richer and the feature expression more comprehensive.

[0181] As an example, in order to enhance information fusion, a dense feature fusion pyramid module is set up to fuse features of different layers. Specifically, the dense feature fusion pyramid module is a top-down and bottom-up fusion structure, which can fuse features of each layer in the backbone network from top to bottom and bottom to top, and can densely fuse features of multiple levels.

[0182] As an example, in the neck network of the data annotation model, the following fusion process is performed on the different scale features extracted by the backbone network based on the dense feature fusion pyramid module:

[0183] B1: Based on the top-down path, the current features of the previous layer of the backbone network are sequentially fused into the current features of the next layer until all features are fused into the last layer of features;

[0184] like Figure 7 As shown, taking the backbone network including 5 levels of representation as an example, that is, from top to bottom, it includes the first layer feature S1, the second layer feature S2, the third layer feature S3, the fourth layer feature S4 and the fifth layer feature S5.

[0185] The input feature P1 of the first layer is fused into the input feature P2 of the second layer, and the fused feature is used as the current feature P2' of the second layer. Then the current feature P2' of the second layer is fused into the input feature P3 of the third layer, and the fused feature is used as the current feature P3' of the third layer. Then the current feature P3' of the third layer is fused into the input feature P4 of the fourth layer, and the fused feature is used as the current feature P4' of the fourth layer. Then the current feature P4' of the fourth layer is fused into the input feature P5 of the fifth layer, and the fused feature is used as the current feature P5' of the fifth layer. A top-down path fusion is achieved.

[0186] B2: Based on the bottom-up path, the current features of the next layer and the input features of the previous layer are fused into the current features of the previous layer in sequence, until all features are fused into the top layer features;

[0187] Continue with Figure 7 For example, from bottom to top, the network integrates features of each layer from S5 to S1.

[0188] As an example, in the process of bottom-up path fusion, the current features of the next layer can be fused into the current features of the previous layer in sequence until all features are fused into the top layer features. The specific fusion method is similar to the above step B1, and please refer to the above description for details.

[0189] As another example, in order to fuse more features without increasing more costs and improve the fusion feature level, in the process of bottom-up path fusion, in addition to the fusion of the previous and next layer features, an interlayer connection (also a jump connection) is added. The interlayer connection connects the original feature nodes of the same layer to the output feature nodes.

[0190] like Figure 7 As shown, the current features of the next layer and the input features of the previous layer are sequentially fused into the current features of the previous layer until all features are fused into the top layer features. Specifically, it can be:

[0191] The current feature P5' of the fifth layer and the input feature P4 of the fourth layer are fused together into the current feature P4' of the fourth layer, and the fused feature is used as the current feature P4' of the fourth layer (in the figure, P4 to P4' are directly connected, which is the above-mentioned interlayer connection). Then the current feature P4'' of the fourth layer and the input feature P3 of the third layer are fused together into the current feature P3' of the third layer, and the fused feature is used as the current feature P3'' of the third layer. Then the current feature P3'' of the third layer and the input feature P2 of the second layer are fused together into the current feature P2' of the second layer, and the fused feature is used as the current feature P2'' of the second layer. Then the current feature P2'' of the second layer is fused into the input feature P1 of the first layer, and the fused feature is used as the current feature P1'' of the first layer. A one-time fusion of the bottom-up path is achieved.

[0192] It should be noted that in this embodiment, in the process of fusing the different scale features extracted by the backbone network based on the dense feature fusion pyramid module, the node with only one input edge is deleted. If a node has only one input edge and no feature fusion, its contribution to the output feature in the fusion is small, so deleting the node has little effect on feature fusion.

[0193] B3: Obtain the feature information of the image based on the current features of each layer after fusion.

[0194] In some examples, each bidirectional (top-down and bottom-up) path is considered as a feature network layer, and the same layer is repeated multiple times to enhance information fusion to achieve higher-level feature fusion. For example:

[0195] Step B3 includes: taking the current features of each layer after fusion as the input features of each layer, and repeating the top-down path and the bottom-up path to fuse the features of each layer; when the number of repetitions reaches N, obtaining the feature information of the image according to the current features of each layer after fusion, N is greater than or equal to 1.

[0196] In this example, repeated top-down and bottom-up fusion is based on the lateral path fusion after the first fusion.

[0197] As mentioned above, steps B1 and B2 are the first top-down and bottom-up fusion processes. Figure 7 , after a top-down and bottom-up fusion, the input features of each layer are all fused features, that is, the first layer input features are P1'', the second layer input features are P2'', the third layer input features are P3'', the fourth layer input features are P4'', and the fifth layer input features are P5''. Then, the top-down and bottom-up fusion are performed through the above-mentioned methods, and so on, and repeated many times to obtain the final output result. The specific process will not be repeated here.

[0198] In an optional implementation, since different input features have different resolutions, they usually contribute unequally to the output features. In order to better fuse multi-layer features, an adaptive weighted fusion mechanism is designed. For example, all features before fusion are concatenated in the feature channel direction, and a 1*1 convolution network is added with softmax to obtain the weight of each feature position. The final fused feature is output based on the weight and the weighted feature after convolution. The formula for fused feature output is as follows

[0199]

[0200] Among them, O is the output fused feature, W is the weight of each position, and Pi is the input feature of each layer;

[0201] As an example, the weight of each position can be calculated by the following formula:

[0202]

[0203] In this embodiment, the weights include channel weights and spatial weights. After 1*1 convolution, the sum is performed in the channel direction, and the obtained one-dimensional vector is obtained by softmax to obtain the channel weight. The sum is performed in the spatial direction to obtain a matrix of the output size, and the spatial weight is obtained by softmax summation.

[0204] In this embodiment, in the neck network of the data annotation model, the different scale features extracted by the backbone network are fused based on the dense feature fusion pyramid module, so as to obtain richer feature information, improve the model learning efficiency, and increase the model training speed.

[0205] Figure 8 A flow chart of a data annotation method provided in an embodiment of the present application is shown as follows: Figure 8 As shown, based on the data annotation model trained in all the above embodiments, the process of annotating data by the data annotation model is described, and the method includes:

[0206] S801, inputting the acquired image into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model;

[0207] In this step, the data annotation model is obtained by the above training method. The data annotation model is applied to the vehicle to annotate the images collected by the vehicle's perception system.

[0208] As an example, the feature extraction module includes a backbone network and a neck network. The feature extraction module based on the data annotation model extracts feature information of the image, including:

[0209] In the backbone network of the data annotation model, feature extraction is performed on the image to obtain features of different scales;

[0210] In the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on the dense feature fusion pyramid module to obtain the feature information of the image.

[0211] As an example, in the neck network of the data annotation model, the following fusion process is performed on the different scale features extracted by the backbone network based on the dense feature fusion pyramid module:

[0212] Based on the top-down path, the current features of the previous layer of the backbone network are fused into the current features of the next layer in sequence until all features are fused into the last layer of features;

[0213] Based on the bottom-up path, the current features of the next layer and the input features of the previous layer are fused into the current features of the previous layer in sequence, until all features are fused into the top layer features;

[0214] The feature information of the image is obtained based on the current features of each layer after fusion.

[0215] In the embodiment of the present application, the specific process of extracting features of the image by the backbone network and the neck network is the same as that described above. Figure 4 The embodiments shown are the same, and the specific process refers to the above embodiments, which will not be repeated here.

[0216] S802, inputting the feature information into the prediction modules corresponding to each task type respectively, and obtaining all the annotation information of the image based on the prediction results output by all the prediction modules;

[0217] In this step, the data annotation model includes multiple encoders, and the multiple encoders can support the prediction of multiple task types. Therefore, the extracted feature information is input into the prediction modules corresponding to each task type respectively, and the annotation of all feature information of the image can be output. Through one data annotation model, the annotation of all features of the image can be completed, and the annotation efficiency is high.

[0218] In this embodiment, independent image processing task models are merged into a single model that uses a shared backbone network to extract image features and outputs results from multiple task prediction modules. This increases the diversity of output attributes for the automatic labeling system, while reducing the total number of model parameters and computing resources occupied, thereby improving the model's reasoning speed.

[0219] The present application also provides a data annotation model training device, including:

[0220] An acquisition unit, used for acquiring training samples, where the training samples include images and true labels marked on the images;

[0221] A feature processing unit, used to input the training sample into the data annotation model, and extract feature information of the image based on the feature extraction module of the data annotation model;

[0222] The training unit is used to input feature information into a target prediction module corresponding to the target task type, and update the model parameters of the data annotation model based on the prediction results and true labels output by the target prediction module.

[0223] The training device for the data annotation model provided in the embodiment of the present application can execute the above method embodiment. Its specific implementation principles and technical effects can be found in the above method embodiment, and this embodiment will not be repeated here.

[0224] The present application also provides a data labeling device, including:

[0225] A feature extraction unit, used to input the acquired image into the data annotation model, and extract feature information of the image based on the feature extraction module of the data annotation model;

[0226] The prediction unit is used to input the feature information into the prediction modules corresponding to each task type respectively, and obtain all the annotation information of the image based on the prediction results output by all the prediction modules.

[0227] Optionally, the data labeling device may be applied in a vehicle.

[0228] The data labeling device provided in the embodiment of the present application can execute the above method embodiment. Its specific implementation principles and technical effects can be found in the above method embodiment, and this embodiment will not be repeated here.

[0229] like Fig. 9 As shown, the embodiment of the present application further provides an electronic device, including: a processor 901, a memory 902; optionally, the electronic device 900 further includes a communication component 903. The processor 901, the memory 902 and the communication component 903 are connected via a CAN bus.

[0230] The memory 902 is used to store computer programs / instructions; the processor 901 is used to execute the computer programs / instructions stored in the memory to implement the methods involved in the above embodiments.

[0231] The electronic device 900 also includes a communication interface, wherein the processor is used to provide computing power and control capabilities, and can be a GPU, CPU, NPU, MCU, FPGA, etc. The storage device includes an internal memory and a non-volatile memory. The non-volatile memory stores a computer program that implements the above method. The internal memory provides an environment for program startup and operation. The communication interface is used to communicate with an external terminal by wire or wireless.

[0232] As an example, the electronic device may be a vehicle.

[0233] The present invention also provides a computer-readable storage medium / computer program product, in which computer control instructions are stored / the computer program product includes computer control instructions, and when the computer control instructions are executed by a processor, they are used to implement the methods involved in the above-mentioned embodiments.

[0234] The above embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Any equivalent substitution or change made by a person skilled in the art based on the present invention is within the protection scope of the present invention.

Claims

1. A data annotation model training method, characterized in that: include: Acquire a training sample, wherein the training sample includes an image and a true label marked on the image; Inputting the training sample into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model; The feature information is input into a target prediction module corresponding to the target task type, and the model parameters of the data annotation model are updated based on the prediction result output by the target prediction module and the true label.

2. The method according to claim 1, characterized in that Inputting the feature information into a target prediction module corresponding to the target task type, and updating the model parameters of the data annotation model based on the prediction result output by the target prediction module and the true label, including: Determining a target task type according to a task type label associated with the feature information, wherein, when preprocessing the image, adding the task type label according to the task type to which the element in the image that has been annotated with the real label belongs; Determine a corresponding target prediction module according to the target task type; The feature information is input into the target prediction module, and the model parameters of the data annotation model are updated based on the prediction result output by the target prediction module and the true label.

3. The method according to claim 1, characterized in that The updating of the model parameters of the data annotation model based on the prediction result output by the target prediction module and the true label includes: Determine the training loss between the prediction result output by the target prediction module and the true label based on the loss function corresponding to the target prediction module; Update a first model parameter of the target prediction module according to the training loss of the target prediction module; The second model parameters of the feature extraction module are updated according to the training losses of all the target prediction modules, and the model parameters of the data annotation model include the first model parameters and the second model parameters.

4. The method according to claim 3, characterized in that The updating of the first model parameter of the target prediction module according to the training loss of the target prediction module comprises: Calculating a first gradient of a training loss of the target prediction module with respect to the first model parameter; Calculating a loss weighted value of the target prediction module according to the first gradient magnitude; The first model parameter is updated according to the loss weight value and the first gradient.

5. The method according to claim 3, characterized in that: The updating of the second model parameter of the feature extraction module according to the training losses of all the target prediction modules comprises: Respectively calculating the second gradients of the training losses of all the target prediction modules with respect to the second model parameters; According to all the second gradients, obtaining a total gradient of the second model parameter; The second model parameters are updated according to the total gradient.

6. The method according to any one of claims 1 to 5, characterized in that: The feature extraction module includes a backbone network and a neck network. The feature extraction module based on the data annotation model extracts feature information of the image, including: In the backbone network of the data annotation model, feature extraction is performed on the image to obtain features of different scales; In the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image.

7. The method according to claim 6, characterized in that In the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image, including: In the neck network of the data annotation model, the following fusion process is performed on the different scale features extracted by the backbone network based on the dense feature fusion pyramid module: Based on the top-down path, the current features of the previous layer of the backbone network are sequentially fused into the current features of the next layer until all features are fused into the last layer of features; Based on the bottom-up path, the current features of the next layer and the input features of the previous layer are fused into the current features of the previous layer in sequence, until all features are fused into the top layer features; The feature information of the image is obtained according to the fused current features of each layer.

8. The method according to claim 7, characterized in that The obtaining feature information of the image according to the fused current features of each layer includes: Using the fused current features of each layer as input features of each layer, and repeating the top-down path and the bottom-up path to fuse the features of each layer; When the number of repetitions reaches N, feature information of the image is obtained according to the current features of each layer after fusion, and N is greater than or equal to 1.

9. A data labeling method, characterized in that: include: Inputting the acquired image into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model; Inputting the feature information into prediction modules corresponding to each task type respectively, and obtaining all the annotation information of the image based on the prediction results output by all the prediction modules; Wherein, the data annotation model is obtained through the training method described in any one of claims 1 to 6.

10. The method according to claim 9, characterized in that The feature extraction module includes a backbone network and a neck network. The feature extraction module based on the data annotation model extracts feature information of the image, including: In the backbone network of the data annotation model, feature extraction is performed on the image to obtain features of different scales; In the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image.

11. The method according to claim 10, characterized in that In the neck network of the data annotation model, the features of different scales extracted by the backbone network are fused based on a dense feature fusion pyramid module to obtain feature information of the image, including: In the neck network of the data annotation model, the following fusion process is performed on the different scale features extracted by the backbone network based on the dense feature fusion pyramid module: Based on the top-down path, the current features of the previous layer of the backbone network are sequentially fused into the current features of the next layer until all features are fused into the last layer of features; Based on the bottom-up path, the current features of the next layer and the input features of the previous layer are fused into the current features of the previous layer in sequence, until all features are fused into the top layer features; The feature information of the image is obtained according to the fused current features of each layer.

12. A training device for a data annotation model, characterized in that: include: An acquisition unit, configured to acquire a training sample, wherein the training sample includes an image and a real label marked on the image; A feature processing unit, used for inputting the training sample into a data annotation model, and extracting feature information of the image based on a feature extraction module of the data annotation model; A training unit is used to input the feature information into a target prediction module corresponding to the target task type, and update the model parameters of the data annotation model based on the prediction result output by the target prediction module and the true label.

13. A data labeling device, characterized in that: include: A feature extraction unit, used to input the acquired image into a data annotation model, and extract feature information of the image based on a feature extraction module of the data annotation model; The prediction unit is used to input the feature information into the prediction modules corresponding to each task type respectively, and obtain all the annotation information of the image based on the prediction results output by all the prediction modules.

14. An electronic device, characterized in that: Including memory, processor; The memory is used to store computer programs / instructions; the processor is used to implement the method as claimed in any one of claims 1 to 8, or to implement the method as claimed in any one of claims 9 to 11 according to the computer programs / instructions stored in the memory.

15. A vehicle, characterized in that: include: Memory, processor; The memory is used to store computer programs / instructions; The processor is configured to implement the method according to any one of claims 9 to 11 according to the computer program / instructions stored in the memory.

16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program / instruction, and the computer program / instruction is used to implement the method according to any one of claims 1 to 11 when executed by a processor.