Driving data labeling method, model training method, program product and equipment
Through the pre-trained driving data annotation model, the problem of limited labeling accuracy and inability to label multi-dimensionality in the prior art is solved, and the training accuracy of the autonomous driving model is improved.
Patent Information
- Application Number
- CN202510472884.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing autonomous driving video annotation method depends on the accuracy of the event start position, resulting in limited labeling accuracy and the inability to label multi-dimensionality, making it difficult to improve the training accuracy of the autonomous driving model.
The current driving video is analyzed through the pre-trained driving data annotation model, and the vehicle status information, event positioning information and event text description information are obtained, and the video is marked in multiple dimensions based on this information.
Improve the labeling accuracy, provide more dimensions of data labeling, and enhance the training accuracy of the autonomous driving model.
Smart Images

Figure CN119992430A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a driving data labeling method, a driving data labeling model training method, a computer program product, and an electronic device. Background Art
[0002] In the related technology, when generating intensive annotations for autonomous driving videos, the video is mainly analyzed in stages. The starting position of the event is first determined, and then text description information is generated based on the starting position of the event. That is, the accuracy of the generation of the text description information is too dependent on the accuracy of the event starting position determined in the previous stage, resulting in limited overall annotation accuracy. In addition, this annotation method can only annotate the event starting position and text description, but not other dimensions. It cannot provide more dimensional data for the training of the autonomous driving model, and it is difficult to improve the training accuracy of the autonomous driving model.
[0003] In view of this, how to perform multi-dimensional annotation of autonomous driving video data and improve the annotation accuracy is a problem that technical personnel in this field need to solve. Summary of the invention
[0004] The present application provides a driving data labeling method, a driving data labeling model training method, a computer program product and an electronic device, which are helpful in improving the labeling accuracy during use, providing more dimensional data for the training of the autonomous driving model, and improving the training accuracy of the autonomous driving model.
[0005] This application provides a driving data annotation method, including: Get the current driving video of the vehicle; Input the current driving video into a pre-trained driving data annotation model to obtain vehicle status information, event location information corresponding to each event in the current driving video, and event text description information; The current driving video is labeled according to the vehicle state information and the event location information and event text description information corresponding to each event in the current driving video; wherein the pre-trained driving data labeling model is obtained by model training based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.
[0006] This application also provides a driving data annotation model training method, including: Acquire multiple sample driving videos and real vehicle motion states corresponding to each sample driving video; Performing model training based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video to obtain a trained driving data annotation model; Among them, the driving data annotation model is used to analyze the current driving video of the vehicle, obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.
[0007] The present application also provides a driving data annotation device, comprising: A first acquisition module, used to acquire a current driving video of the vehicle; An analysis module, used to input the current driving video into a pre-trained driving data annotation model to obtain vehicle status information, event location information corresponding to each event in the current driving video, and event text description information; The labeling module is used to label the current driving video according to the vehicle state information and the event location information and event text description information corresponding to each event in the current driving video; wherein the pre-trained driving data labeling model is obtained by the training module through model training based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.
[0008] The present application also provides a driving data annotation model training device, comprising: The second acquisition module is used to acquire multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video; A training module, used for performing model training based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video, to obtain a trained driving data annotation model; Among them, the driving data annotation model is used to analyze the current driving video of the vehicle, obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.
[0009] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned driving data labeling methods, or implements the steps of any of the above-mentioned driving data labeling model training methods.
[0010] The present application also provides an electronic device, comprising: Memory for storing computer programs; A processor is used to implement the steps of any of the above-mentioned driving data labeling methods when executing a computer program, or to implement the steps of any of the above-mentioned driving data labeling model training methods.
[0011] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any one of the above-mentioned driving data labeling methods are implemented, or the steps of any one of the above-mentioned driving data labeling model training methods are implemented.
[0012] It can be seen from the above technical solution that the beneficial effects of the present invention are: The driving data annotation method provided by the present application is pre-trained based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video, to obtain a driving data annotation model, and in the process of annotating the driving video, obtain the current driving video, and input the current driving video into the driving data annotation model for analysis, to obtain vehicle status information, event location information and event text description information corresponding to each event of the current driving video, and annotate the current driving video based on the vehicle status information, event location information and event text description information corresponding to each event of the current driving video. When the present application annotates and analyzes the current driving video, the event location information and event text description information of each event can be obtained by analyzing the current driving video at the same time. The event text description information is not generated by relying on the event location information, which is conducive to improving the annotation accuracy, and the vehicle status information can also be obtained, so that more dimensional annotation information can be obtained, which is conducive to providing more dimensional data for the training of the autonomous driving model, so as to better improve the training accuracy of the autonomous driving model. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 A flowchart of a driving data labeling method provided in an embodiment of the present application; Figure 2 A structural block diagram of a driving data annotation model provided in an embodiment of the present application; Figure 3 A schematic diagram of a process for generating annotation information for a current driving video based on a driving data annotation model provided in an embodiment of the present application; Figure 4A flowchart of a driving data annotation model training method provided in an embodiment of the present application; Figure 5 A training architecture diagram of a driving data annotation model provided in an embodiment of the present application; Figure 6 A flowchart of another driving data annotation model training method provided in an embodiment of the present application; Figure 7 A flow chart of a vehicle state estimation module provided in an embodiment of the present application for estimating a vehicle motion state and determining a state loss value; Figure 8 A schematic diagram of a process for determining a predicted speed and a predicted steering angle provided in an embodiment of the present application; Fig. 9 A schematic diagram of the structure of a driving data labeling device provided in an embodiment of the present application; Fig.10 A schematic diagram of the structure of a driving data annotation model training device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0016] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0017] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0018] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the driving data labeling method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0019] The embodiment of the present application provides a driving data annotation method, please refer to Figure 1, the method is described in detail in combination with the execution flow of the driving data labeling method. The method includes the following S110 to S130.
[0020] S110: Acquire the current driving video of the vehicle.
[0021] It should be noted that a large number of sample driving videos and the real vehicle motion state corresponding to each sample driving video can be pre-acquired in the present application. These sample driving videos are all embodied perspective driving videos of the vehicle, that is, driving videos with the vehicle as the first-person perspective. The sample driving video includes multiple real events, and each real event is also annotated with corresponding annotation information, and the annotation information includes the real positioning information and real text description information of the corresponding event. Among them, the real positioning information in the annotation information can be the real start and end time of the event, and the real text description information can include event behavior description information and / or event cause analysis information. The real vehicle motion state in the present application can be determined based on the positioning data of the GPS (Global Positioning System), and the real vehicle motion state corresponding to the sample driving video is determined according to the positioning data. In the present application, a large number of sample driving videos and the corresponding real vehicle motion states that have been pre-acquired can be used to train the initial model to obtain a driving data annotation model.
[0022] In practical applications, a current driving video of the vehicle is obtained, which is also a driving video from the embodied perspective of the vehicle during the driving process of the vehicle, that is, a current driving video from the first-person perspective of the vehicle.
[0023] S120: Input the current driving video into a pre-trained driving data annotation model to obtain vehicle state information, event location information corresponding to each event in the current driving video, and event text description information.
[0024] It can be understood that each time a current driving video is obtained, the current driving video is input into the trained driving data annotation model. Since the driving data annotation model is obtained by model training based on a large number of sample driving videos and the corresponding real vehicle motion states, and each real event in each sample driving video is annotated with corresponding annotation information, the annotation information includes the real positioning information and real text description information of the corresponding event. Therefore, the current driving video is analyzed through the driving data annotation model, and the vehicle state information corresponding to the current driving video, the event positioning information and event text description information corresponding to each event in the current driving video can be output.
[0025] S130: Annotate the current driving video according to the vehicle state information and the event location information and event text description information respectively corresponding to each event in the current driving video; wherein the pre-trained driving data annotation model is obtained by model training based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.
[0026] In the present application, after obtaining the vehicle status information of the current driving video and the event location information and event text description information corresponding to each event of the current driving video, the current driving video can be annotated according to the vehicle status information and the event location information and event text description information corresponding to each event of the current driving video, thereby completing the driving data annotation of the current driving video.
[0027] It should be noted that in this application, the current driving video is analyzed through the driving data annotation model, and the event location information and event text description information corresponding to each event can be obtained at the same time. In this application, there is no need to perform annotation in stages, and the event text description information is not obtained based on the event location information. Therefore, the accuracy of the event location information will not affect the accuracy of the event text description information, which is conducive to improving the overall annotation accuracy. This application can not only obtain the event location information and event text description information of each event, but also obtain the vehicle status information, that is, it can annotate data of more dimensions, thereby obtaining annotation information of more dimensions, which is conducive to providing more dimensional data for the training of the autonomous driving model, so as to better improve the training accuracy of the autonomous driving model.
[0028] It can be seen that in this application, a driving data annotation model is obtained in advance based on multiple sample driving videos and the real vehicle motion state training corresponding to each sample driving video, and the current driving video is obtained in the process of annotating the driving video, and the current driving video is input into the driving data annotation model for analysis to obtain vehicle status information, event location information corresponding to each event of the current driving video, and event text description information, and the current driving video is annotated according to the vehicle status information, event location information corresponding to each event of the current driving video, and event text description information. When the present application annotates and analyzes the current driving video, the event location information and event text description information of each event can be obtained by analyzing the current driving video at the same time. The event text description information is not generated by relying on the event location information, which is conducive to improving the annotation accuracy, and the vehicle status information can also be obtained, so that more dimensional annotation information can be obtained, which is conducive to providing more dimensional data for the training of the autonomous driving model, so as to better improve the training accuracy of the autonomous driving model.
[0029] Based on the above embodiments, the embodiments of the present application further illustrate and optimize the technical solution, as follows: In one embodiment, Figure 2 As shown, the driving data annotation model in the present application may include a motion feature extraction module, a coding and decoding module, a vehicle state estimation module, an event location module and a text generation module.
[0030] That is, in the process of training the driving data annotation model, a large number of pre-acquired sample driving videos and the corresponding real vehicle motion states are used to train the motion feature extraction module, codec module, vehicle state estimation module, event location module and text generation module in the initial model to obtain the trained motion feature extraction module, codec module, vehicle state estimation module, event location module and text generation module, thereby obtaining a trained driving data annotation model.
[0031] like Figure 3 As shown, the process of inputting the current driving video into the pre-trained driving data annotation model in the above S120 to obtain vehicle status information, event location information corresponding to each event of the current driving video, and event text description information may include the following S210 to S250.
[0032] S210: Input the current driving video into a pre-trained driving data annotation model, and extract motion features from the current driving video through a motion feature extraction module to obtain various current motion features.
[0033] It should be noted that, in the embodiment of the present application, after the current driving video is acquired, the current driving video may be input into a motion feature extraction model of a driving data annotation model to extract motion features, thereby obtaining a plurality of current motion features.
[0034] In practical applications, the driving data annotation model may further include a preprocessing module, into which the current driving video may be first input, and the current driving video may be preprocessed by the preprocessing module. For example, the current driving video may be processed into a video with a preset number of frames per second (30 frames), and then N frames of images may be evenly extracted from the processed video, and the N frames of images with equal intervals may be input into a motion feature extraction module for motion feature extraction, thereby obtaining multiple current motion features.
[0035] The motion feature extraction module in the embodiment of the present invention may include a trained initial motion feature extraction unit and a feature expansion unit, wherein the feature expansion unit may include multiple convolutional neural networks of different scales, wherein the initial motion feature extraction unit may perform feature extraction on N frames of images in the current driving video to obtain N first motion features, and in order to further improve the accuracy of data annotation, the N first motion features may be input into the feature expansion unit for expansion to obtain multiple current motion features corresponding to each scale, that is, each scale of the convolutional neural network will output multiple current motion features.
[0036] S220: performing encoding and decoding processing on each current motion feature through the encoding and decoding module to obtain each current coded motion feature and each current predicted event representation feature.
[0037] It should be noted that after obtaining each current motion feature, the current motion feature can be input into a codec module for processing. The codec module includes an encoder and a decoder, wherein the encoder can encode each current motion feature to obtain multiple current encoded motion features, and the decoder can use the cross-attention mechanism to interact with the current encoded motion feature and the pre-stored initialization representation features of each preset event to obtain the representation features of each current predicted event.
[0038] S230: Processing each current encoded motion feature through a vehicle state estimation module to obtain vehicle state information.
[0039] The vehicle state estimation module in the present application can obtain corresponding vehicle state information by performing state estimation on each current encoded motion feature output by the encoder in the codec module, wherein the vehicle state information may include speeds and corresponding steering angles at multiple moments. For example, the preset number of frames can be determined according to the preset video length, and the speed and steering angle corresponding to each frame in the preset number of frames can be obtained.
[0040] S240: Performing event location processing on each predicted event representation feature through an event location module to obtain event location information corresponding to each event in the current driving video.
[0041] It should be noted that the decoder in the encoding and decoding module outputs each current predicted event characterization feature, and the event location module can perform event location processing on each predicted event characterization feature to obtain event location information corresponding to each event in the current driving video, wherein the event location information may include the event start time and the event end time.
[0042] S250: Performing text generation processing on each current predicted event representation feature through a text generation module to obtain event text description information corresponding to each event.
[0043] It can be understood that the text generation module in the present application also performs text generation processing based on the characterization features of each current predicted event output by the decoder to obtain event text description information corresponding to each event, and the event text description information includes event behavior description information and / or event cause analysis information.
[0044] When the driving data annotation model in the present application analyzes and annotates the current driving video, a plurality of current motion features are obtained through the motion feature extraction module, and then the plurality of current motion features are processed by the encoder to obtain a plurality of current coded motion features, which are output to the vehicle state estimation module on the one hand and to the decoder on the other hand. The vehicle state estimation module can obtain vehicle state information by performing state estimation on the plurality of current coded motion features; the encoder obtains a plurality of current predicted event representation features by interacting the plurality of current coded motion features with the plurality of pre-stored initialization representation features, and then performs event positioning processing on each predicted event representation feature through the event positioning module to obtain event positioning information corresponding to each event in the current driving video, and performs text generation processing on each current predicted event representation feature through the text generation module to obtain event text description information corresponding to each event. Thus, vehicle state information, event positioning information and event text description information can be obtained to realize multi-dimensional annotation of driving data.
[0045] That is, when the present application performs annotation analysis on the current driving video, it can encode and decode the multiple current motion features corresponding to the current driving video to obtain multiple current encoded motion features and multiple current predicted event representation features. The current multiple predicted event representation features in the present application are obtained by comprehensively considering the interaction of multiple current encoded motion features and multiple initialization representation features, and the event location information and event text description information of each event in the present application are obtained by analyzing multiple predicted event representation features. There is no progressive relationship between event location and text description, that is, the event text description information is not generated based on the event location information. Since the multiple predicted event representation features obtained by the decoder in the present application have high accuracy, the event location information and event text description information of each event obtained according to the multiple predicted event representation features have high accuracy, which is conducive to improving the overall annotation accuracy. In addition, while obtaining the event location information and event text description information in the present application, each current encoded motion feature can also be analyzed through the vehicle state estimation module to obtain more accurate vehicle state information, which can further improve the accuracy of data annotation, and can also provide more dimensional and more accurate annotation data for the training of the autonomous driving model to better improve the training accuracy of the autonomous driving model.
[0046] It should be noted that the training process of the driving data annotation model involved in the driving data annotation method provided in the above embodiment can refer to the training method of the driving data annotation model introduced in the next embodiment, and this embodiment will not be repeated here.
[0047] Based on the above embodiments, the present application also provides a driving data annotation model training method. Figure 4 , the method includes the following S310 to S320.
[0048] S310: Acquire multiple sample driving videos and the actual vehicle motion state corresponding to each sample driving video.
[0049] It should be noted that in the embodiment of the present application, multiple sample driving videos can be pre-acquired, in which the real event location information and real text description information of each real event are annotated, and the real vehicle status information corresponding to the sample driving video is also obtained.
[0050] In practical applications, multiple original driving videos (videos with the vehicle as the first-person perspective) and positioning data (such as GPS data) corresponding to each original driving video can be obtained in advance, and the multiple original driving videos can be preprocessed to obtain various sample driving videos.
[0051] The original driving video is annotated with real location information and real text description information of each real event, the sample driving video includes N frames of images distributed at equal intervals, the real location information includes the start and end time of the real event, and the event text description information includes description information and / or event cause analysis information;
[0052] It should be noted that in practical applications, the preprocessing process may include denoising the original driving video, etc. In order to further increase the number of samples and thus improve the accuracy of model training, in this application, sample expansion can be performed based on each denoised original driving video, and then each expanded initial sample driving video can be obtained. For example, according to the real text description information corresponding to each real event in each original driving video, the target original video data containing the turning action in the real text description information can be screened out from all the original video data, and then each target original video data is flipped to obtain the flipped target original video data, and the vehicle action in the event text description information of the event corresponding to the turning action in the flipped target original video data is modified to the corresponding reverse action (for example, the event text description information corresponding to the original event is left turn, and the modified reverse action is right turn). Then, each original video data and each flipped target original video data are processed into a video with the same frame rate (for example, 30 frames / s), and a preset number (N) of frame images are obtained from each video at equal intervals to form each sample video data.
[0053] S320: Performing model training based on multiple sample driving videos and the actual vehicle motion state corresponding to each sample driving video to obtain a trained driving data annotation model.
[0054] Among them, the driving data annotation model is used to analyze the current driving video of the vehicle, obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.
[0055] It should be noted that in the embodiment of the present application, the model can be trained according to each sample driving video annotated with the real positioning information and real text description information of each real event and the real vehicle motion state corresponding to each sample driving video. The final trained driving data annotation model is obtained through training. Then, when the driving data of the running vehicle is annotated, the trained driving data standard model can be directly used to analyze the current driving video of the vehicle to obtain the current vehicle state information, the event positioning information and event text description information corresponding to each event of the current driving video, and the current driving video is annotated with this information, so that multi-dimensional annotation of the vehicle driving video can be achieved, and the event positioning information and event text description information of each event obtained are independent of each other, and the event text description information is not obtained by relying on the event positioning information, so that the accuracy of driving data annotation can be improved, and multi-dimensional data can be annotated, which is conducive to providing more dimensional annotation data for the training of the autonomous driving model to improve the training accuracy of the autonomous driving model.
[0056] The following will further explain and introduce the training process of the driving data annotation model. Please refer to Figure 5 and Figure 6 ,in, Figure 5 A flowchart of a training process of a driving data annotation model provided in an embodiment of the present application, Figure 6 A training architecture diagram of a driving data annotation model provided in an embodiment of the present application, the training process of the driving data annotation model in the above S320 may include the following S410 to S460.
[0057] S410: For each iterative training, N frames of images in the current sample driving video are input into the motion feature extraction module to obtain various motion features; the current sample driving video is annotated with real positioning information and real text description information of each real event; wherein N is an integer not less than 2.
[0058] It should be noted that the driving data annotation model can be initialized before training. Each module in the initialized driving data annotation model is an initial module, for example, an initial motion feature extraction module, an initial encoding and decoding module, an initial vehicle state estimation module, an initial event location module and an initial text generation module. After multiple iterations, the current motion feature extraction module, encoding and decoding module, vehicle state estimation module, event location module and text generation module can be obtained.
[0059] In the embodiment of the present invention, one iterative training is taken as an example to explain in detail. A plurality of sample driving videos are obtained in advance. For the current round of iterative training, N frames of images in the current sample video can be The motion feature extraction module is used to extract motion features from N frames of images to obtain multiple motion features, among which f1 to f N They are the 1st frame image to the Nth frame image respectively.
[0060] In one embodiment, S410 inputs N frames of images in the current sample driving video into a motion feature extraction module to obtain various motion features, which may include: performing motion feature extraction on the N frames of images in the current sample driving video through an initial motion feature extraction unit in the motion feature extraction module to obtain N initial motion features.
[0061] The N initial motion features are respectively expanded by convolutional neural networks of multiple different scales in the motion feature extraction module to obtain multiple motion features corresponding to the convolutional neural network of each scale; wherein each convolutional neural network of each scale outputs multiple motion features corresponding thereto.
[0062] It should be noted that the motion feature extraction module in the present application includes an initial motion feature extraction unit and a feature expansion unit, wherein, in order to improve the accuracy of feature extraction, the initial motion feature extraction unit can be a stream model in a pre-trained I3D (Two-Stream Inflated 3D ConvNets, a method of inflating a 2D network into a 3D network) model, and the feature expansion unit can include multiple convolutional neural networks of different scales, for example, including i convolutional neural networks of different scales. In order to reduce the amount of calculation while ensuring the accuracy of the final text generation and the accuracy of event positioning, 4 convolutional neural networks of different scales can be selected in actual applications, that is, i can be 4.
[0063] It is understandable that in the present application, the initial motion feature extraction unit can be used to extract N frames of images in the current driving video. Perform feature extraction to obtain N initial motion features ,in, to Represent the 1st to Nth initial motion features respectively, and then input the N initial motion features into the feature expansion unit for expansion to obtain multiple motion features corresponding to each scale ,in, represents the jth motion feature of the i-th convolutional layer. Each scale of the convolutional neural network will output multiple motion features, that is, the motion features output by the i-th convolutional layer are ; Expressed as a relation: ,in: i∈{1,2,3,4}, , Kernel represents the convolution kernel, stride represents the step size, the value of stride can be m, k represents the scale of the convolution kernel, k can be 2, 4, 8, 16, and p represents sequence zero padding.
[0064] S420: Processing each motion feature through a coding and decoding module to obtain each coded motion feature and a coded representation feature corresponding to each predicted event.
[0065] In practical applications, the encoding and decoding module may include an encoder and a decoder; the implementation process of S320 may include: All motion features are encoded by the encoder to obtain various encoded motion features at different scales; The decoder uses the cross-attention mechanism to interactively process the initialization representation features and the encoded motion features corresponding to each predicted event, and obtains the encoded representation features corresponding to each predicted event.
[0066] It should be noted that the encoding and decoding module in the present application may include an encoder and a decoder, wherein the encoder in the present application may be an encoder of a Transformer architecture, and the various motion features output by the motion feature extraction module may be processed by the current encoder. Perform feature encoding processing to obtain the corresponding coded motion features ,in, represents the jth encoded motion feature of the i-th convolutional layer.
[0067] In the embodiment of the present application, multiple fixed events can be preset. For example, Q fixed events can be preset, and Q prediction events can be randomly initialized at the decoder input end in accordance with the DETR (Detection Transformer, a Transformer-based target detection model) mode to obtain Q initialization representation features. ,in, to They represent the 1st to Qth initialization representation features respectively. During the model training process, each encoding representation feature output by the encoder can be received. The decoder uses the cross-attention mechanism to interactively process each initialization representation feature and each encoding motion feature, and obtains the encoded encoding representation features corresponding to each predicted event at the end. ,in, to Represent the 1st to Qth coding representation features respectively.
[0068] Among them, Cross-Attention is to calculate attention on two different sequences to process the semantic relationship between the two sequences. For example, in the translation task, it is necessary to align the source language sentence and the target language sentence, and cross-attention is needed to calculate the attention weight between the two sentences.
[0069] The cross-attention mechanism is a special form of multi-head attention that splits the input tensor into two parts and ,in, Represents a set of real numbers, and then takes one part as the query set X1, and the other part as the key value set X2. The output is a tensor of size d1×d2. For each row vector, its attention weight for all row vectors is given.
[0070] Specifically, and , then the calculation of cross attention is as follows: ,in, and is the learned projection matrix, d1, d2, d k is the dimension of the key-value set (also the dimension of the query set), n is the sequence length, , K, and V are query matrix, key matrix, and value matrix respectively. is the activation function.
[0071] In this application, the input X1 of the cross attention can be each encoded motion feature , input X2 can be used to initialize the characterization feature , so that the encoding representation features can be obtained through the above cross attention mechanism .
[0072] S430: Processing each encoded motion feature through a vehicle state estimation module to obtain a predicted vehicle motion state, and determining a state loss value based on the predicted vehicle motion state and the actual vehicle motion state corresponding to the current sample driving video.
[0073] It should be noted that in the embodiment of the present application, after encoding each motion feature through an encoder to obtain each encoded motion feature, the state of each encoded motion feature can be estimated through the current vehicle state estimation module to obtain a predicted vehicle motion state, which may include multiple predicted speeds and predicted steering angles corresponding to each predicted speed. Since the corresponding real vehicle motion state is obtained when obtaining each sample driving video data in the present application, the state loss value corresponding to the vehicle state estimation module in this iterative training can be obtained based on the each real speed and the corresponding real steering angle in the real vehicle motion state corresponding to the current sample driving video.
[0074] In one embodiment, if Figure 7 As shown, the implementation process of S430 may include S510 to S520.
[0075] S510: Estimate each encoded motion feature at different scales through the vehicle state estimation module to obtain the predicted speed and predicted steering angle corresponding to each of the M frames in the current sample driving video; M is an integer greater than 0 and less than or equal to N.
[0076] In practical applications, the real vehicle motion state of the current sample driving video can be obtained in advance through the GPS data corresponding to the current sample driving video, wherein the GPS data records the vehicle speed during the entire video of the current sample driving video. The steering angle can be obtained by calculating the heading angle change between adjacent frames in the GPS data, thereby obtaining the real speed and the corresponding real steering angle. For example, the real vehicle motion state of the current sample driving video is , where v1, v2, …v M-1 、v M are the actual vehicle speeds corresponding to the first to the Mth frames in the current sample driving video, s1, s2, …s M-1 , 0 are the real steering angles corresponding to the 1st frame to the Mth frame. Since the GPS data is recorded according to the time interval rather than frame by frame, M≤N.
[0077] Since the length of each sample driving video is the same, the specific value of the M frame can be pre-set in the vehicle state estimation module. When performing state estimation, the predicted speed and predicted steering angle corresponding to each of the M frames in the current sample driving video can be obtained, where M≤N.
[0078] It should be noted that in order to improve the prediction speed and the estimation accuracy of the predicted steering angle, please refer to Figure 8The above process of obtaining the predicted speed and predicted steering angle corresponding to each frame in the M frames in the current sample driving video in the present application can be implemented through the following S610 to S640.
[0079] S610: Performing feature enhancement on each coded motion feature at each scale through a first bidirectional long short-term memory network in a vehicle state estimation module to obtain enhanced coded motion features at each scale.
[0080] It can be understood that in order to more accurately estimate the predicted speed and the predicted steering angle in the present application, the first bidirectional long short-term memory network (Bi-directional Long Short-Term Memory, BiLSTM) can be used to enhance the features of each encoded motion feature at each scale to obtain each enhanced encoded motion feature at each scale, which is expressed as follows: ,in, represents the first bidirectional long short-term memory network, i∈{1,2,3,4}, represents the j-th enhanced coded motion feature at the i-th scale, so that N at each scale can be obtained i An enhanced encoded motion feature.
[0081] S620: Downsampling each enhanced coded motion feature at each scale by a linear interpolation method to obtain each downsampled coded motion feature at each scale; wherein the number of each downsampled coded motion feature at any scale is M.
[0082] Furthermore, in order to improve the calculation speed and reduce resource consumption, the present application uses a linear interpolation method to downsample each enhanced coded motion feature at each scale, and the sampling can be specifically downsampled to M, for example: ,in, to are the encoded motion features after sampling at the i-th scale.
[0083] S630: Perform maximum pooling processing on all downsampled coded motion features to obtain M maximum pooled coded motion features.
[0084] It is understandable that in order to further retain key information and significant features and reduce computational complexity in this application, the encoded motion features after downsampling at each scale may be subjected to maximum pooling processing MaxPooling, which may be achieved, for example, by the following relationship: ,in, , that is, the coded motion features after each maximum pooling are obtained as ,in, to They represent the encoded motion features after the 1st to the Mth maximum pooling respectively.
[0085] S640: Utilize the first multi-layer perceptron network to convert each maximum pooled encoded motion feature into a scalar value, and obtain a predicted speed and a predicted steering angle corresponding to each frame in the M frames.
[0086] In practical applications, the encoded motion features after each maximum pooling can be converted into is converted into a scalar value, which includes speed v and steering angle s, wherein the first multi-layer perceptron network includes and . It can be obtained through the following relationship: ; .
[0087] in, Indicates the predicted speed corresponding to each frame in M frames, It represents the predicted steering angles corresponding to the first M-1 frames in the M frames. The steering angle of the Mth frame can be considered as 0, so that the predicted speed and predicted steering angle corresponding to each frame in the M frames can be obtained quickly and accurately.
[0088] S520: Determine a state loss value based on the predicted speed and predicted steering angle corresponding to each frame in the M frames and the real speed and real steering angle corresponding to each frame in the M frames in the real vehicle motion state corresponding to the current sample driving video.
[0089] In the embodiment of the present application, the state loss value can be calculated according to each predicted speed and predicted steering angle, and each real speed and real steering angle corresponding to the current sample driving video. In order to improve the calculation accuracy of the state loss value, the mean variance loss (that is, the state loss value) can be calculated according to the following relationship: ) calculation: ,in, represents the actual vehicle speed in the i-th frame, represents the predicted speed of the i-th frame, represents the actual steering angle of the i-th frame, represents the predicted steering angle of the i-th frame.
[0090] S440: Perform positioning analysis on each encoded representation feature through the event positioning module to obtain predicted positioning information corresponding to each predicted event, and combine the real positioning information of each real event of the current sample driving video to obtain a positioning loss value.
[0091] It should be noted that the various encoding representation features output by the decoder will be input into the event location module and the text generation module, wherein the event location module performs location analysis based on the various encoding representation features, and can obtain predicted event location information corresponding to each predicted event, wherein the current sample driving video carries labels of the real location information of each real event, and therefore the location loss value can be further calculated based on each predicted event location information and each real event location information.
[0092] In one embodiment, the process of S440 may include: The event location module locates and analyzes the encoding representation features corresponding to each predicted event to obtain the predicted start and end time and predicted event category corresponding to each predicted event.
[0093] It should be noted that in order to enhance the event monitoring effect in this application, an event type allocation strategy is introduced in the process of training the event location module. Event type allocation further assists event time location, which is beneficial to improve the accuracy of predicted event location.
[0094] That is, the event location module in the present application includes a second multi-layer perceptron network based on The time positioning unit and the classification unit can further map the encoding representation features corresponding to the Q predicted events into a two-dimensional vector through the second multi-layer perceptron network in the current event positioning module. , , get the predicted start time corresponding to each predicted event and forecast end time ; For example, it can be obtained through the following relationship: ;in, Indicates the first Encoding representation features, Indicates the first The predicted start time corresponding to the encoding representation feature, Indicates the first The predicted end time corresponding to the encoded representation feature.
[0095] In order to make full use of the given information of the event to enhance the event detection effect, in the embodiment of the present invention, while detecting the time area, the event is also classified, and the semantic information of the category is used to constrain the detection of the starting position of the event. Therefore, in this application, the classification unit in the current event location module can also perform category prediction on the encoding representation features corresponding to each predicted event, and obtain the predicted event category corresponding to each predicted event.
[0096] Furthermore, according to the predicted start and end time and predicted event category corresponding to each predicted event, combined with the real start and end time and real category of each real event in the current sample driving video, the positioning loss value is obtained.
[0097] It can be understood that the label of the real positioning information corresponding to the current sample driving video may include the start and end times of the real events corresponding to P real events respectively. The set of P real events can be recorded as ,in, represents the jth real event, , represents the starting time of the jth real event, represents the end time of the jth real event, , represents the driving behavior description of the jth real event, Represents the behavioral cause description of the jth real event.
[0098] The above process of obtaining the positioning loss value may include: The predicted start time and the predicted end time corresponding to each predicted event are matched with the real start time and the real end time corresponding to each real event in the current sample driving video to obtain a predicted event corresponding to each real event, so as to obtain each mapping event pair.
[0099] It should be noted that, in the embodiment of the present application, the predicted event set Q can be matched with each event in the real event set P, that is, the predicted events that match the P real events are found from the Q predicted events. For example, the Hungarian algorithm can be used to map each real event in P to a predicted event in Q at the minimum cost: , the jth real event in P The predicted events mapped in Q are ,but and is a mapping event pair, so that multiple mapping event pairs can be obtained.
[0100] For each true event and predicted event in a mapped event pair, the first loss value of the true event and the predicted event is calculated.
[0101] That is, for the mapping event pair and , the first loss value of the mapping event pair can be calculated , so that the first loss value corresponding to each mapping event pair can be obtained.
[0102] After obtaining each first loss value, the time loss value can be obtained according to the first loss value corresponding to each mapping event pair; for example, the total time loss value L can be obtained by adding each first loss value. loc : .
[0103] For each predicted event, a second loss value is calculated according to the predicted event category corresponding to the predicted event and the category label corresponding to the predicted event; wherein corresponding category labels are respectively assigned to each predicted event in advance.
[0104] It is understandable that in the embodiment of the present application, a set of multiple event categories can be predetermined, such as {left turn, right turn, acceleration, deceleration, parking, other}, and each predicted event can be classified in advance to determine the corresponding category label. For example, the Q predicted events can be classified by the large language model GPT-4, and the operation instruction can be: "Please classify the following driving behaviors according to the description. The optional categories are {left turn, right turn, acceleration, deceleration, parking, other}, and the driving behavior description is '[text description]'", thereby obtaining the category label of each predicted event, which is a pseudo label y.
[0105] In the above process, it has been introduced that in this application, the classification unit can be used to perform category prediction on the encoding representation features corresponding to each predicted event, and obtain the predicted event category corresponding to each predicted event. For example, a multi-layer perceptron network can also be used to predict the predicted event. The encoding characteristics of Perform category prediction and obtain a label x of a predicted event category corresponding to the predicted event.
[0106] It should be noted that the second loss value corresponding to the cross entropy loss can be further calculated based on the label x of the predicted event category corresponding to each predicted event and the pseudo label y (which can be used as the true label) corresponding to the predicted event. The relationship between the cross entropy loss CrossEntropyLoss is as follows: ,in, Indicates the cross entropy loss function operation on x and y. Indicates The predicted event corresponds to The labels of the categories (that is, the labels of the predicted event categories), Indicates The first predicted event The true label corresponding to each category, C represents the number of categories, and Q represents the total number of predicted events.
[0107] The category loss value is obtained according to the second loss value corresponding to each predicted event.
[0108] After obtaining the second loss value corresponding to each predicted event, the second loss values can be summed to obtain the total category loss value L cls .
[0109] According to the time loss value L loc And the category loss value L cls , and get the positioning loss value.
[0110] In the present application embodiment, L loc +L cls As the overall loss of event location detection, that is, the sum of the two is used as the location loss value, which not only ensures the optimal correspondence between real events and predicted events, but also enhances the semantic similarity between the two, thereby achieving the purpose of minimizing the overall loss. The model trained based on this can improve the accuracy of event location when performing event location.
[0111] S450: Perform text processing on each encoded representation feature through a text generation module to obtain predicted text description information corresponding to each predicted event, and combine the real text description information of each real event to obtain a text loss value.
[0112] It should be noted that in the embodiment of the present application, the current text module in this iterative training process performs text processing on each encoded representation feature output by the encoder, and the predicted text description information corresponding to each predicted event can be obtained. According to the above introduction, by matching each predicted event with each real event, each mapping event pair is obtained, and each mapping event pair includes a real event and a predicted event corresponding to the real event. For each predicted event, the corresponding real event can be determined according to the corresponding mapping event pair, and then the real text description information corresponding to each real event corresponding to the current sample driving video can be obtained. The real text representation of the real event. Then, combined with the real text description information of each real event, the text loss value can be obtained.
[0113] In one embodiment, the process implemented in S450 may include:
[0114] The second bidirectional long short-term memory network in the text generation module performs text generation processing on the encoded representation features corresponding to each predicted event, so as to obtain predicted text description information corresponding to each predicted event.
[0115] It can be understood that in the embodiment of the present application, the second bidirectional long short-term memory network BiLSTM based on the attention mechanism in the current text generation module performs text generation processing on the encoded representation features corresponding to each predicted event, wherein the maximum length limit of the text generation (e.g., 50 words) can also be preset. For example, the j-th predicted event in the set Q is mapped to the i-th real event in the set P, and the text generation process can be implemented by the following relationship: ,in, represents the encoding representation feature corresponding to the j-th predicted event, represents the prediction text description information sequence corresponding to the jth prediction event. Based on this relation, a prediction text description information sequence unique to each prediction event can be obtained.
[0116] According to the real events corresponding to each predicted event, according to the predicted text description information of the predicted event and the real text description information corresponding to the real event, a third loss value corresponding to the predicted event is determined.
[0117] According to the mapping event pairs obtained in the above introduction, the real text description information of the real event corresponding to each predicted event can be obtained. , so that the generated prediction text description information can be The corresponding real text description information , calculate the loss value to obtain the corresponding third loss value. For example, the third loss value can be calculated by the following relationship : ,in, Express and Perform cross entropy loss function calculation.
[0118] The text loss value is obtained according to the third loss value corresponding to each predicted event.
[0119] It can be understood that after obtaining the third loss value corresponding to each predicted event, each third loss value can be summed to obtain the overall text loss value L text .
[0120] S460: Based on the state loss value, the positioning loss value and the text loss value, the various parameters in the motion feature extraction module, the encoding and decoding module, the vehicle state estimation module, the event positioning module and the text generation module are updated, and a trained driving data annotation model is obtained when the training end conditions are met.
[0121] It should be noted that, in the embodiment of the present invention, after obtaining the state loss value, the positioning loss value and the text loss value, the respective loss values can be summed to obtain the total loss value of this iterative training, that is, the total loss value is: L mse +L loc +L cls +L text , and then it can be further determined whether the training end conditions are met. If not, the gradient calculation and back propagation of the neural network can be performed according to the total loss value to update the parameters of each network in each module in the driving data annotation model, thereby obtaining each updated module and performing the next round of iterative training. If the training end conditions are met, the training is terminated to obtain a trained driving data annotation model. Among them, the training end conditions may include that the total loss value is less than a preset loss value or the number of iterations reaches a maximum number of iterations. In the present application, during the training process of the driving data annotation model, the loss functions of each stage are integrated to update the parameters of each module and the corresponding network in the driving data annotation model, thereby improving the accuracy of the parameter update, thereby obtaining a more accurate driving data annotation model.
[0122] It can be seen that in this application, a driving data annotation model is obtained in advance based on multiple sample driving videos and the real vehicle motion state training corresponding to each sample driving video, and the current driving video is obtained in the process of annotating the driving video, and the current driving video is input into the driving data annotation model for analysis to obtain vehicle status information, event location information corresponding to each event of the current driving video, and event text description information, and the current driving video is annotated according to the vehicle status information, event location information corresponding to each event of the current driving video, and event text description information. When the present application annotates and analyzes the current driving video, the event location information and event text description information of each event can be obtained by analyzing the current driving video at the same time. The event text description information is not generated by relying on the event location information, which is conducive to improving the annotation accuracy, and the vehicle status information can also be obtained, so that more dimensional annotation information can be obtained, which is conducive to providing more dimensional data for the training of the autonomous driving model, so as to better improve the training accuracy of the autonomous driving model.
[0123] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0124] Based on the above embodiments, the embodiments of the present application also provide a driving data labeling device, please refer to Fig. 9 , the device comprises: A first acquisition module 11 is used to acquire a current driving video of the vehicle; The analysis module 12 is used to input the current driving video into a pre-trained driving data annotation model to obtain vehicle status information, event location information corresponding to each event in the current driving video, and event text description information; The labeling module 13 is used to label the current driving video according to the vehicle state information and the event location information and event text description information corresponding to each event in the current driving video; wherein the pre-trained driving data labeling model is obtained by the training module through model training based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.
[0125] In one embodiment, the driving data annotation model includes a motion feature extraction module, a coding and decoding module, a vehicle state estimation module, an event location module and a text generation module.
[0126] The analysis module 12 includes: A first extraction unit is used to input the current driving video into a pre-trained driving data annotation model, and extract motion features from the current driving video through a motion feature extraction module to obtain various current motion features; A coding and decoding unit, used for coding and decoding each current motion feature through a coding and decoding module to obtain each current coded motion feature and each current predicted event representation feature; A state estimation unit, used to process each current encoded motion feature through a vehicle state estimation module to obtain vehicle state information; A positioning unit, used to perform event positioning processing on the representation features of each predicted event through an event positioning module to obtain event positioning information corresponding to each event in the current driving video; The text generation unit is used to perform text generation processing on each current predicted event representation feature through a text generation module to obtain event text description information corresponding to each event.
[0127] For the description of the features in the embodiment corresponding to the driving data labeling device in the embodiment of the present application, reference can be made to the relevant description of the embodiment corresponding to the driving data labeling method, which will not be repeated here one by one.
[0128] Based on the above embodiments, this application also provides a driving data annotation model training device, please refer to Fig.10 , the device comprises: The second acquisition module 21 is used to acquire a plurality of sample driving videos and the real vehicle motion state corresponding to each sample driving video; A training module 22, used for performing model training based on a plurality of sample driving videos and a real vehicle motion state corresponding to each sample driving video, to obtain a trained driving data annotation model; Among them, the driving data annotation model is used to analyze the current driving video of the vehicle, obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.
[0129] In one embodiment, the driving data annotation model includes a motion feature extraction module, a coding and decoding module, a vehicle state estimation module, an event location module and a text generation module.
[0130] Training module 22, including: A second extraction unit is used for inputting N frames of images in the current sample driving video into the motion feature extraction module for each iterative training to obtain various motion features; the current sample driving video is annotated with real positioning information and real text description information of each real event; wherein N is an integer not less than 2; A first processing unit, configured to process each motion feature through a coding and decoding module to obtain each coded motion feature and a coded representation feature corresponding to each predicted event; a second processing unit, configured to process each encoded motion feature through a vehicle state estimation module to obtain a predicted vehicle motion state, and determine a state loss value based on the predicted vehicle motion state and a real vehicle motion state corresponding to a current sample driving video; A first analysis unit is used to perform positioning analysis on each encoding representation feature through an event positioning module to obtain predicted positioning information corresponding to each predicted event, and to obtain a positioning loss value by combining the real positioning information of each real event of the current sample driving video; A third processing unit is used to perform text processing on each encoded representation feature through a text generation module to obtain predicted text description information corresponding to each predicted event, and to obtain a text loss value by combining the real text description information of each real event; The updating unit is used to update various parameters in the motion feature extraction module, the encoding and decoding module, the vehicle state estimation module, the event positioning module and the text generation module based on the state loss value, the positioning loss value and the text loss value, and obtain a trained driving data annotation model when the training end conditions are met.
[0131] In one embodiment, the second extraction unit comprises: An extraction subunit is used to extract motion features from N frames of images in the current sample driving video through an initial motion feature extraction unit in the motion feature extraction module to obtain N initial motion features; The expansion subunit is used to perform feature expansion on N initial motion features through multiple convolutional neural networks of different scales in the motion feature extraction module to obtain multiple motion features corresponding to the convolutional neural networks of each scale; wherein each convolutional neural network of each scale outputs multiple motion features corresponding to it.
[0132] In one embodiment, the encoding and decoding module includes an encoder and a decoder.
[0133] The first processing unit comprises: A first processing subunit is used to encode all motion features through an encoder to obtain various encoded motion features at different scales; The second processing sub-unit is used to interactively process the initialization representation features and the encoded motion features corresponding to each prediction event through the decoder using a cross-attention mechanism to obtain the encoded representation features corresponding to each prediction event.
[0134] In one embodiment, the second processing unit includes: The state estimation subunit is used to estimate various coded motion features at different scales through the vehicle state estimation module to obtain the predicted speed and predicted steering angle corresponding to each of the M frames in the current sample driving video; M is an integer greater than 0 and less than or equal to N; The first loss value determination subunit is used to determine the state loss value based on the predicted speed and predicted steering angle corresponding to each frame in the M frames and the actual speed and actual steering angle corresponding to each frame in the M frames in the actual vehicle motion state corresponding to the current sample driving video.
[0135] In one embodiment, the state estimation subunit includes: A feature enhancement subunit is used to perform feature enhancement on each coded motion feature at each scale through a first bidirectional long short-term memory network in a vehicle state estimation module to obtain each enhanced coded motion feature at each scale; A downsampling processing subunit is used to downsample each enhanced coded motion feature at each scale by a linear interpolation method to obtain each downsampled coded motion feature at each scale; wherein the number of each downsampled coded motion feature at any scale is M; The maximum pooling processing subunit is used to perform maximum pooling processing on all downsampled coded motion features to obtain M maximum pooled coded motion features; The first conversion subunit is used to convert each maximum pooled encoded motion feature into a scalar value by using a first multi-layer perceptron network to obtain a predicted speed and a predicted steering angle corresponding to each frame in the M frames.
[0136] In one embodiment, the first analysis unit includes: The positioning analysis subunit is used to perform positioning analysis on the coding representation features corresponding to each predicted event through the event positioning module to obtain the predicted start and end time and predicted event category corresponding to each predicted event; The second loss value determination subunit is used to obtain a positioning loss value according to the predicted start and end time and predicted event category corresponding to each predicted event, combined with the real start and end time and real category of each real event in the current sample driving video.
[0137] In one embodiment, the positioning analysis subunit includes: The second conversion subunit is used to map the encoding representation features corresponding to each predicted event into a two-dimensional vector through the second multi-layer perceptron network in the event location module to obtain the predicted start time and predicted end time corresponding to each predicted event; The category prediction subunit is used to perform category prediction on the encoding representation features corresponding to each predicted event through the classification unit in the event positioning module to obtain the predicted event category corresponding to each predicted event.
[0138] In one embodiment, the second loss value determination subunit includes: A matching subunit is used to match the predicted start time and the predicted end time corresponding to each predicted event with the real start time and the real end time corresponding to each real event in the current sample driving video, so as to obtain a predicted event corresponding to each real event, so as to obtain each mapping event pair; A first calculation subunit, configured to calculate a first loss value of the real event and the predicted event for each mapping event pair; A first determining subunit, configured to obtain a time loss value according to first loss values respectively corresponding to each mapping event pair; A second calculation subunit is used to calculate a second loss value for each predicted event according to the predicted event category corresponding to the predicted event and the category label corresponding to the predicted event; wherein corresponding category labels are respectively assigned to each preset component in advance; A third determination subunit is used to obtain a category loss value according to the second loss values corresponding to each prediction event; The fourth determination subunit is used to obtain a positioning loss value according to the time loss value and the category loss value.
[0139] In one embodiment, the third processing unit includes: A first generating subunit is used to perform text generation processing on the encoding representation features corresponding to each predicted event respectively through a second bidirectional long short-term memory network in the text generation module to obtain predicted text description information corresponding to each predicted event respectively; a fifth determining subunit, configured to determine, according to the real events respectively corresponding to each predicted event, the predicted text description information of the predicted event and the real text description information corresponding to the real event, a third loss value corresponding to the predicted event; The sixth determination subunit is used to obtain a text loss value according to the third loss value corresponding to each predicted event.
[0140] In one embodiment, the device further comprises: A pre-acquisition module, used for pre-acquisition of a plurality of original driving videos, where the original driving videos are annotated with real location information and real text description information of each real event; An expansion module, used to expand samples based on the original driving video to obtain each expanded initial sample driving video; The preprocessing module is used to preprocess each initial sample driving video to obtain each sample driving video; the sample driving video includes N frames of images distributed at equal intervals, where N is an integer not less than 2.
[0141] For the description of the features in the embodiment corresponding to the driving data labeling model training device in the embodiment of the present application, please refer to the relevant description of the embodiment corresponding to the driving data labeling model training method, which will not be repeated here.
[0142] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned driving data labeling method embodiments, or to implement the steps of any of the above-mentioned driving data labeling model training methods.
[0143] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned driving data labeling method embodiments when running, or implement the steps of any of the above-mentioned driving data labeling model training methods.
[0144] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0145] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned driving data labeling method embodiments, or implements the steps of any of the above-mentioned driving data labeling model training methods.
[0146] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned driving data labeling method embodiments are implemented.
[0147] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0148] The above is a detailed introduction to a driving data labeling method, program product, electronic device and storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A driving data labeling method, characterized in that: include: Get the current driving video of the vehicle; Inputting the current driving video into a pre-trained driving data annotation model to obtain vehicle state information, event location information corresponding to each event in the current driving video, and event text description information; The current driving video is labeled according to the vehicle state information and the event location information and event text description information corresponding to each event in the current driving video; wherein the pre-trained driving data labeling model is obtained by model training based on multiple sample driving videos and the real vehicle motion state corresponding to each of the sample driving videos.
2. The driving data labeling method according to claim 1, characterized in that: The driving data annotation model includes a motion feature extraction module, a coding and decoding module, a vehicle state estimation module, an event location module and a text generation module; The step of inputting the current driving video into a pre-trained driving data annotation model to obtain vehicle state information, event location information corresponding to each event in the current driving video, and event text description information includes: Inputting the current driving video into a pre-trained driving data annotation model, and extracting motion features from the current driving video through the motion feature extraction module to obtain various current motion features; Performing encoding and decoding processing on each of the current motion features through the encoding and decoding module to obtain each current coded motion feature and each current predicted event representation feature; Processing each of the current encoded motion features through the vehicle state estimation module to obtain vehicle state information; Performing event location processing on each of the current predicted event representation features through the event location module to obtain event location information corresponding to each event in the current driving video; The text generation module performs text generation processing on each of the current predicted event representation features to obtain event text description information corresponding to each of the events.
3. A driving data annotation model training method, characterized in that: include: Acquire multiple sample driving videos and real vehicle motion states corresponding to each of the sample driving videos; Performing model training based on the plurality of sample driving videos and the real vehicle motion state corresponding to each of the sample driving videos to obtain a trained driving data annotation model; Among them, the driving data annotation model is used to analyze the current driving video of the vehicle, obtain vehicle status information, event location information and event text description information corresponding to each event of the current driving video, and annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event of the current driving video.
4. The driving data annotation model training method according to claim 3, characterized in that: The driving data annotation model includes a motion feature extraction module, a coding and decoding module, a vehicle state estimation module, an event location module and a text generation module; The model training is performed based on the plurality of sample driving videos and the real vehicle motion state corresponding to each of the sample driving videos to obtain a trained driving data annotation model, including: For each iterative training, N frames of images in the current sample driving video are input into the motion feature extraction module to obtain various motion features; the current sample driving video is annotated with real positioning information and real text description information of each real event; wherein N is an integer not less than 2; Processing each of the motion features through the encoding and decoding module to obtain each coded motion feature and a coded representation feature corresponding to each predicted event; Processing each of the encoded motion features through the vehicle state estimation module to obtain a predicted vehicle motion state, and determining a state loss value based on the predicted vehicle motion state and a real vehicle motion state corresponding to the current sample driving video; Performing positioning analysis on each of the encoded representation features through the event positioning module to obtain predicted positioning information corresponding to each of the predicted events, and combining the real positioning information of each real event of the current sample driving video to obtain a positioning loss value; Performing text processing on each of the encoded representation features through the text generation module to obtain predicted text description information corresponding to each of the predicted events, and combining the real text description information of each real event to obtain a text loss value; Based on the state loss value, the positioning loss value and the text loss value, each parameter in the motion feature extraction module, the encoding and decoding module, the vehicle state estimation module, the event positioning module and the text generation module is updated, and a trained driving data annotation model is obtained when the training end conditions are met.
5. The driving data annotation model training method according to claim 4, characterized in that: The N frames of images in the current sample driving video are input into the motion feature extraction module to obtain various motion features, including: Extracting motion features from N frames of images in the current sample driving video by an initial motion feature extraction unit in the motion feature extraction module to obtain N initial motion features; The N initial motion features are respectively expanded by convolutional neural networks of multiple different scales in the motion feature extraction module to obtain multiple motion features corresponding to each convolutional neural network of the scale; wherein each convolutional neural network of the scale outputs multiple motion features corresponding to it.
6. The driving data annotation model training method according to claim 5, characterized in that: The coding and decoding module includes an encoder and a decoder; The encoding and decoding module processes each of the motion features to obtain each coded motion feature and a coded representation feature corresponding to each predicted event, including: All the motion features are encoded by an encoder to obtain various encoded motion features at different scales; The decoder uses a cross-attention mechanism to interactively process the initialization representation features corresponding to each predicted event and each encoded motion feature to obtain the encoded representation features corresponding to each predicted event.
7. The driving data annotation model training method according to claim 6, characterized in that: The method further comprises: processing each of the encoded motion features through the vehicle state estimation module to obtain a predicted vehicle motion state, and determining a state loss value based on the predicted vehicle motion state and a real vehicle motion state corresponding to the current sample driving video, including: The vehicle state estimation module estimates each of the encoded motion features at different scales to obtain a predicted speed and a predicted steering angle corresponding to each of the M frames in the current sample driving video; M is an integer greater than 0 and less than or equal to N; The state loss value is determined based on the predicted speed and predicted steering angle corresponding to each frame in the M frames and the real speed and real steering angle corresponding to each frame in the M frames in the real vehicle motion state corresponding to the current sample driving video.
8. The driving data annotation model training method according to claim 7, characterized in that: The vehicle state estimation module estimates each of the encoded motion features at different scales to obtain a predicted speed and a predicted steering angle corresponding to each of the M frames in the current sample driving video, including: Performing feature enhancement on each of the coded motion features at each scale by using a first bidirectional long short-term memory network in the vehicle state estimation module to obtain enhanced coded motion features at each scale; Downsampling each enhanced coded motion feature at each scale by a linear interpolation method to obtain each downsampled coded motion feature at each scale; wherein the number of each downsampled coded motion feature at any scale is M; Performing maximum pooling processing on all downsampled coded motion features to obtain M maximum pooled coded motion features; The first multi-layer perceptron network is used to convert each of the maximum pooled encoded motion features into a scalar value to obtain a predicted speed and a predicted steering angle corresponding to each frame in the M frames.
9. The driving data annotation model training method according to claim 5, characterized in that: The event positioning module performs positioning analysis on each of the encoded representation features to obtain predicted positioning information corresponding to each of the predicted events, and combines the real positioning information of each real event of the current sample driving video to obtain a positioning loss value, including: The event location module locates and analyzes the coding representation features corresponding to each of the predicted events to obtain the predicted start and end time and the predicted event category corresponding to each of the predicted events; According to the predicted start and end time and the predicted event category corresponding to each predicted event, combined with the real start and end time and the real category of each real event in the current sample driving video, a positioning loss value is obtained.
10. The driving data annotation model training method according to claim 9, characterized in that: The event location module performs location analysis on the coding representation features corresponding to each of the predicted events to obtain the predicted start and end time and the predicted event category corresponding to each of the predicted events, including: Mapping the encoding representation features corresponding to each of the predicted events into a two-dimensional vector through the second multi-layer perceptron network in the event location module to obtain the predicted start time and predicted end time corresponding to each of the predicted events; The classification unit in the event location module performs category prediction on the encoding representation features corresponding to each of the predicted events to obtain the predicted event category corresponding to each of the predicted events.
11. The driving data annotation model training method according to claim 10, characterized in that: The positioning loss value is obtained according to the predicted start and end time and predicted event category corresponding to each predicted event, combined with the real start and end time and real category of each real event in the current sample driving video, include: Matching the predicted start time and the predicted end time corresponding to each predicted event with the real start time and the real end time corresponding to each real event in the current sample driving video, to obtain a predicted event corresponding to each real event, so as to obtain each mapping event pair; For each real event and predicted event in the mapped event pair, calculating a first loss value between the real event and the predicted event; Obtaining a time loss value according to the first loss value respectively corresponding to each of the mapping event pairs; For each predicted event, a second loss value is calculated according to the predicted event category corresponding to the predicted event and the category label corresponding to the predicted event; wherein corresponding category labels are respectively assigned to each predicted event in advance; Obtaining a category loss value according to the second loss values respectively corresponding to the predicted events; A positioning loss value is obtained according to the time loss value and the category loss value.
12. The driving data annotation model training method according to claim 11, characterized in that: The text generation module performs text processing on each of the encoded representation features to obtain predicted text description information corresponding to each predicted event, and combines the real text description information of each real event to obtain a text loss value, including: Performing text generation processing on the encoding representation features corresponding to each of the predicted events through a second bidirectional long short-term memory network in the text generation module to obtain predicted text description information corresponding to each of the predicted events; According to the real events respectively corresponding to each of the predicted events, according to the predicted text description information of the predicted event and the real text description information corresponding to the real event, determining a third loss value corresponding to the predicted event; A text loss value is obtained according to the third loss value corresponding to each of the predicted events.
13. The driving data annotation model training method according to any one of claims 3 to 12, characterized in that: Also includes: Acquire multiple original driving videos in advance, where the original driving videos are annotated with real location information and real text description information of each real event; Performing sample expansion based on the original driving video to obtain expanded initial sample driving videos; Each of the initial sample driving videos is preprocessed to obtain each sample driving video; the sample driving video includes N frames of images distributed at equal intervals, wherein N is an integer not less than 2.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the driving data labeling method according to claim 1 or 2 are implemented, or the steps of the driving data labeling model training method according to any one of claims 3 to 13 are implemented.
15. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the driving data labeling method according to claim 1 or 2, or the steps of the driving data labeling model training method according to any one of claims 3 to 13, when executing the computer program.
Citation Information
Patent Citations
Video dense description method and device and medium
CN113312980A
Dense video description method, device and system and storage medium
CN115861868A
Driving behavior recognition model training method, recognition method, device and equipment
CN117076711A
Automatic labeling method and device for automatic driving data, computer equipment and medium
CN117557828A
Video labeling method and device, computer equipment and computer readable storage medium
CN117710845A