A driving data annotation method, model training method, program product, and equipment

By extracting vehicle state and event information through a driving data annotation model, the problem of annotation depending on the starting position of events in existing technologies is solved, multi-dimensional data annotation is achieved, and the training accuracy of autonomous driving models is improved.

CN119992430BActive Publication Date: 2025-10-31SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510472884.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-10-31
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing methods for annotating autonomous driving videos mainly rely on the accuracy of the event start location, which limits the accuracy of the annotation, makes it impossible to provide multi-dimensional data, and makes it difficult to improve the training accuracy of autonomous driving models.

Method used

By acquiring the vehicle's current driving video, and using a pre-trained driving data annotation model, vehicle status information, event location information, and event text description information are extracted. The model is trained using multiple sample driving videos and real vehicle motion states to achieve multi-dimensional annotation.

Benefits of technology

It improves annotation accuracy, provides more dimensions of data, enhances the training accuracy of autonomous driving models, ensures that event text description information is not dependent on event location information for generation, and enhances the independence and accuracy of annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992430B_ABST
    Figure CN119992430B_ABST
Patent Text Reader

Abstract

This application discloses a driving data annotation method, model training method, program product, and device, relating to the field of autonomous driving technology. The method involves pre-training a driving data annotation model based on multiple sample driving videos and the corresponding real vehicle motion states for each sample video. During the annotation process, the current driving video is acquired and input into the driving data annotation model for analysis, yielding vehicle state information, event location information corresponding to each event in the current driving video, and event text description information. The current driving video is then annotated based on the vehicle state information, the event location information corresponding to each event in the current driving video, and the event text description information. This method solves the problems of insufficient data acquisition dimensions and low annotation accuracy in related technologies, achieving the technical effect of acquiring more dimensional data and improving annotation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular to a driving data annotation method, a driving data annotation model training method, a computer program product, and an electronic device. Background Technology

[0002] In related technologies, when generating intensive annotations for autonomous driving videos, the main approach is to analyze the video in stages. First, the starting position of the event is determined, and then text description information is generated based on the starting position. In other words, the accuracy of the generated text description information depends too much on the accuracy of the starting position of the event determined in the previous stage, which limits the overall annotation accuracy. Furthermore, this annotation method can only annotate the starting position of the event and the text description, and cannot annotate other dimensions. It cannot provide more dimensional data for the training of autonomous driving models, making it difficult to improve the training accuracy of autonomous driving models.

[0003] Therefore, how to perform multi-dimensional annotation of autonomous driving video data and improve the accuracy of annotation is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a driving data annotation method, a driving data annotation model training method, a computer program product, and an electronic device, which are beneficial for improving annotation accuracy, providing more dimensional data for the training of autonomous driving models, and improving the training accuracy of autonomous driving models during use.

[0005] This application provides a driving data annotation method, including:

[0006] Obtain the vehicle's current driving video;

[0007] The current driving video is input into a pre-trained driving data annotation model to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video.

[0008] The current driving video is labeled based on vehicle status information, event location information, and event text description information corresponding to each event in the current driving video. The pre-trained driving data labeling model is obtained by training the model based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.

[0009] This application also provides a method for training a driving data annotation model, including:

[0010] Acquire multiple sample driving videos and the corresponding real vehicle motion state for each sample driving video;

[0011] The model is trained based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video to obtain the trained driving data annotation model.

[0012] The driving data annotation model is used to analyze the current driving video of the vehicle to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and to annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.

[0013] This application also provides a driving data annotation device, including:

[0014] The first acquisition module is used to acquire the current driving video of the vehicle;

[0015] The analysis module is used to input the current driving video into a pre-trained driving data annotation model to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video.

[0016] The annotation module is used to annotate the current driving video based on vehicle status information and event location information and event text description information corresponding to each event in the current driving video. The pre-trained driving data annotation model is obtained by the training module based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.

[0017] This application also provides a driving data annotation model training device, including:

[0018] The second acquisition module is used to acquire multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.

[0019] The training module is used to train the model based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video, so as to obtain the trained driving data annotation model.

[0020] The driving data annotation model is used to analyze the current driving video of the vehicle to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and to annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.

[0021] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described driving data annotation methods, or the steps of any of the above-described driving data annotation model training methods.

[0022] This application also provides an electronic device, including:

[0023] Memory, used to store computer programs;

[0024] A processor for executing a computer program to implement the steps of any of the above-described driving data annotation methods, or to implement the steps of any of the above-described driving data annotation model training methods.

[0025] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described driving data annotation methods, or implements the steps of any of the above-described driving data annotation model training methods.

[0026] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows:

[0027] The driving data annotation method provided in this application pre-trains a driving data annotation model based on multiple sample driving videos and the corresponding real vehicle motion states for each sample driving video. During the annotation process, the current driving video is acquired and input into the driving data annotation model for analysis. This yields vehicle state information, event location information corresponding to each event in the current driving video, and event text description information. The current driving video is then annotated based on these vehicle state information, event location information, and event text description information. This application, when annotating and analyzing the current driving video, can simultaneously obtain event location information and event text description information for each event. The event text description information is not generated dependent on event location information, which helps improve annotation accuracy. Furthermore, vehicle state information can be obtained, thus acquiring more dimensional annotation information. This provides more dimensional data for training autonomous driving models, thereby improving the training accuracy of autonomous driving models. Attached Figure Description

[0028] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 A flowchart illustrating a driving data annotation method provided in an embodiment of this application;

[0030] Figure 2 A structural block diagram of a driving data annotation model provided in this application embodiment;

[0031] Figure 3 This application provides a schematic diagram of a process for generating annotation information from a current driving video based on a driving data annotation model, as provided in an embodiment of the present application.

[0032] Figure 4 A flowchart illustrating a driving data annotation model training method provided in this application embodiment;

[0033] Figure 5 A training architecture diagram of a driving data annotation model provided in this application embodiment;

[0034] Figure 6 A flowchart illustrating another driving data annotation model training method provided in this application embodiment;

[0035] Figure 7 A flowchart illustrating how a vehicle state estimation module performs vehicle motion state estimation and determines state loss values, as provided in an embodiment of this application.

[0036] Figure 8 A schematic diagram illustrating a process for determining predicted speed and predicted steering angle, provided for an embodiment of this application;

[0037] Figure 9 This is a schematic diagram of the structure of a driving data annotation device provided in an embodiment of this application;

[0038] Figure 10 This is a schematic diagram of the structure of a driving data annotation model training device provided in an embodiment of this application. Detailed Implementation

[0039] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0040] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0041] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0042] The specific application environment architecture or specific hardware architecture on which the driving data annotation method depends is described here.

[0043] The embodiments of this application provide a driving data annotation method, please refer to... Figure 1 The method is described in detail below, taking into account the execution flow of the driving data annotation method. The method includes the following steps S110 to S130.

[0044] S110: Obtain the current driving video of the vehicle.

[0045] It should be noted that this application can pre-acquire a large number of sample driving videos and the corresponding real vehicle motion states for each sample driving video. These sample driving videos are all vehicle-centric first-person perspective driving videos, which are driving videos with the vehicle as the first-person viewpoint. Each sample driving video includes multiple real events, and each real event is also labeled with corresponding annotation information, including the real location information and real text description information of the corresponding event. The real location information in the annotation information can be the real start and end time of the event, and the real text description information can include event behavior description information and / or event cause analysis information. The real vehicle motion state in this application can be determined based on GPS (Global Positioning System) positioning data. The real vehicle motion state corresponding to the sample driving video is determined based on the positioning data. In this application, the pre-acquired large number of sample driving videos and their corresponding real vehicle motion states can be used to train the initial model to obtain the driving data annotation model.

[0046] In practical applications, the current driving video of the vehicle is obtained. This current driving video is also the driving video from the vehicle's own perspective during the vehicle's movement, which is the current driving video from the vehicle's first-person perspective.

[0047] S120: Input the current driving video into the pre-trained driving data annotation model to obtain vehicle status information, event location information corresponding to each event in the current driving video, and event text description information.

[0048] Understandably, each time a current driving video is acquired, it is input into a trained driving data annotation model. Since this driving data annotation model is trained based on a large number of sample driving videos and their corresponding real vehicle motion states, and each real event in each sample driving video is labeled with corresponding annotation information, including the real location information and real text description information of the corresponding event, the driving data annotation model can analyze the current driving video and output the vehicle state information corresponding to the current driving video, the event location information and event text description information corresponding to each event in the current driving video.

[0049] S130: The current driving video is labeled based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video; wherein, the pre-trained driving data labeling model is obtained by training the model based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.

[0050] In this application, after obtaining the vehicle status information of the current driving video and the event location information and event text description information corresponding to each event in the current driving video, the current driving video can be annotated according to the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video, thereby completing the annotation of driving data for the current driving video.

[0051] It should be noted that this application analyzes the current driving video using a driving data annotation model, simultaneously obtaining the event location information and event text description information for each event. This application eliminates the need for staged annotation, and the event text description information is not derived from the event location information. Therefore, the accuracy of the event location information does not affect the accuracy of the event text description information, thus improving the overall annotation accuracy. This application not only obtains the event location information and event text description information for each event but also vehicle state information, enabling annotation of data from more dimensions. This provides more dimensional annotation information, which is beneficial for providing more data for the training of autonomous driving models, thereby improving the training accuracy of autonomous driving models.

[0052] Therefore, this application pre-trains a driving data annotation model based on multiple sample driving videos and the corresponding real vehicle motion states for each sample driving video. During the annotation process, the current driving video is acquired and input into the driving data annotation model for analysis. This yields vehicle state information, event location information, and event text description information corresponding to each event in the current driving video. The current driving video is then annotated based on these vehicle state information, event location information, and event text description information. This application, when annotating and analyzing the current driving video, can simultaneously obtain event location information and event text description information for each event. The event text description information is not generated dependent on event location information, which improves annotation accuracy. Furthermore, vehicle state information can be obtained, thus acquiring more dimensional annotation information. This provides more dimensional data for training the autonomous driving model, thereby improving the training accuracy of the autonomous driving model.

[0053] Based on the above embodiments, the present application embodiments further illustrate and optimize the technical solution, as follows: In one embodiment, as follows... Figure 2 As shown, the driving data annotation model in this application may include a motion feature extraction module, an encoding / decoding module, a vehicle state estimation module, an event localization module, and a text generation module.

[0054] In other words, during the training of the driving data annotation model, a large number of pre-acquired sample driving videos and corresponding real vehicle motion states are used to train the motion feature extraction module, encoding and decoding module, vehicle state estimation module, event localization module, and text generation module in the initial model, thereby obtaining the trained motion feature extraction module, encoding and decoding module, vehicle state estimation module, event localization module, and text generation module, thus obtaining the trained driving data annotation model.

[0055] like Figure 3 As shown, the process of inputting the current driving video into the pre-trained driving data annotation model in S120 above to obtain vehicle status information, event location information corresponding to each event in the current driving video, and event text description information may include the following S210 to S250.

[0056] S210: Input the current driving video into the pre-trained driving data annotation model, and extract the motion features of the current driving video through the motion feature extraction module to obtain various current motion features.

[0057] It should be noted that, in this embodiment of the application, after obtaining the current driving video, the current driving video can be input into the motion feature extraction model of the driving data annotation model to extract motion features and obtain multiple current motion features.

[0058] In practical applications, the driving data annotation model may also include a preprocessing module. The current driving video can be input into the preprocessing module and preprocessed, for example, the current driving video can be processed into a video with a preset number of frames per second (30 frames). Then, N frames of images are extracted from the processed video at equal intervals. The N frames of images with equal intervals are input into the motion feature extraction module for motion feature extraction, thereby obtaining multiple current motion features.

[0059] The motion feature extraction module in this embodiment of the invention may include a trained initial motion feature extraction unit and a feature expansion unit. The feature expansion unit may include multiple convolutional neural networks of different scales. The initial motion feature extraction unit can extract features from N frames of images in the current driving video to obtain N first motion features. In order to further improve the accuracy of data annotation, the N first motion features can be input into the feature expansion unit for expansion to obtain multiple current motion features corresponding to each scale. That is, the convolutional neural network at each scale will output multiple current motion features.

[0060] S220: The encoding and decoding module performs encoding and decoding processing on each current motion feature to obtain each current encoded motion feature and each current predicted event representation feature.

[0061] It should be noted that after obtaining each current motion feature, the current motion feature can be input into the encoding and decoding module for processing. The encoding and decoding module includes an encoder and a decoder. The encoder can encode each current motion feature to obtain multiple current encoded motion features. The decoder can use a cross-attention mechanism to interact with the current encoded motion feature and the initialization representation features of each preset event to obtain the representation features of each current predicted event.

[0062] S230: The vehicle state estimation module processes each currently encoded motion feature to obtain vehicle state information.

[0063] The vehicle state estimation module in this application can obtain the corresponding vehicle state information by estimating the state of each current encoded motion feature output by the encoder in the encoding and decoding module. The vehicle state information can include the speed and corresponding steering angle at multiple times. For example, a preset number of frames can be determined according to the preset video length, and the speed and steering angle corresponding to each frame in the preset number of frames can be obtained.

[0064] S240: The event localization module performs event localization processing on the representation features of each predicted event to obtain the event localization information corresponding to each event in the current driving video.

[0065] It should be noted that the decoder in the encoding and decoding module outputs the representation features of each currently predicted event, and the event localization module can perform event localization processing on each prediction event representation feature to obtain event localization information corresponding to each event in the current driving video. The event localization information may include the event start time and the event end time.

[0066] S250: The text generation module performs text generation processing on the representation features of each currently predicted event to obtain the event text description information corresponding to each event.

[0067] It is understood that the text generation module in this application also performs text generation processing based on the representation features of each currently predicted event output by the decoder to obtain event text description information corresponding to each event. This event text description information includes event behavior description information and / or event cause analysis information.

[0068] The driving data annotation model in this application, when analyzing and annotating the current driving video, obtains multiple current motion features through a motion feature extraction module. Then, an encoder processes these current motion features to obtain multiple currently encoded motion features. These multiple currently encoded motion features are output to both a vehicle state estimation module and a decoder. The vehicle state estimation module estimates the vehicle state information by analyzing these multiple currently encoded motion features. The encoder interacts with multiple pre-stored initialization features to obtain multiple currently predicted event representation features. An event localization module then performs event localization processing on each predicted event representation feature to obtain event localization information corresponding to each event in the current driving video. Finally, a text generation module performs text generation processing on each currently predicted event representation feature to obtain event text description information corresponding to each event. Thus, vehicle state information, event localization information, and event text description information can be obtained, enabling multi-dimensional annotation of driving data.

[0069] In other words, when annotating and analyzing the current driving video, this application can encode and decode multiple current motion features corresponding to the current driving video to obtain multiple current encoded motion features and multiple current predicted event representation features. The current multiple predicted event representation features in this application are obtained by comprehensively considering the interaction of multiple current encoded motion features and multiple initial representation features. Furthermore, the event location information and event text description information of each event in this application are obtained based on the analysis of multiple predicted event representation features. There is no progressive relationship between event location and text description; that is, the event text description information is not generated dependent on event location information. Since the multiple predicted event representation features obtained by the decoder in this application have high accuracy, the accuracy of obtaining the event location information and event text description information of each event based on these multiple predicted event representation features is high, which is beneficial to improving the overall annotation accuracy. In addition, while obtaining the event location information and event text description information, this application can also analyze each current encoded motion feature through the vehicle state estimation module to obtain more accurate vehicle state information, which can further improve the data annotation accuracy and provide more dimensional and more accurate annotation data for the training of autonomous driving models, thereby better improving the training accuracy of autonomous driving models.

[0070] It should be noted that the training process of the driving data annotation model involved in the driving data annotation method provided in the above embodiments can refer to the training method of the driving data annotation model introduced in the next embodiment, and will not be repeated here.

[0071] Based on the above embodiments, this application also provides a driving data annotation model training method, please refer to... Figure 4 The method includes the following steps S310 to S320.

[0072] S310: Obtain multiple sample driving videos and the corresponding real vehicle motion state for each sample driving video.

[0073] It should be noted that in this embodiment of the application, multiple sample driving videos can be obtained in advance. These sample driving videos are marked with real event location information and real text description information of each real event, and the real vehicle status information corresponding to the sample driving video is also obtained.

[0074] In practical applications, multiple original driving videos (videos from the vehicle's first-person perspective) and corresponding location data (such as GPS data) can be pre-acquired, and the multiple original driving videos can be pre-processed to obtain individual sample driving videos.

[0075] The original driving videos are labeled with real location information and real text description information for each real event. The sample driving videos include N equally spaced frames of images. The real location information includes the start and end times of the real events, and the event text description information includes description information and / or event cause analysis information.

[0076] It should be noted that in practical applications, the preprocessing process may include noise reduction of the original driving video. To further increase the number of samples and thus improve the model training accuracy, this application can expand the samples based on each denoised original driving video to obtain expanded initial sample driving videos. For example, based on the real text description information corresponding to each real event in each original driving video, target original video data containing steering actions in the real text description information can be selected from all original video data. Then, each target original video data is flipped to obtain flipped target original video data. The vehicle action in the event text description information corresponding to the turning action in the flipped target original video data is modified to the corresponding reverse action (e.g., the original event text description information is left turn, and the modified reverse action is right turn). Then, each original video data and each flipped target original video data are processed into videos with the same frame rate (e.g., 30 frames / s), and a preset number (N) of frame images are obtained from each video at equal intervals to form each sample video data.

[0077] S320: The model is trained based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video to obtain the trained driving data annotation model.

[0078] The driving data annotation model is used to analyze the current driving video of the vehicle to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and to annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.

[0079] It should be noted that in this embodiment, the model can be trained based on each sample driving video labeled with real location information and real text description information of each real event, as well as the real vehicle motion state corresponding to each sample driving video. The final trained driving data annotation model is obtained through training. Then, when annotating driving data for a moving vehicle, this trained driving data standard model can be directly used to analyze the current driving video of the vehicle, obtaining the current vehicle state information, the event location information and event text description information corresponding to each event in the current driving video, and using this information to annotate the current driving video. This enables multi-dimensional annotation of the vehicle driving video. Furthermore, the event location information and event text description information obtained for each event are independent of each other; the event text description information is not dependent on the event location information. This improves the accuracy of driving data annotation and allows for multi-dimensional data annotation, which is beneficial for providing more dimensional annotated data for the training of autonomous driving models, thereby improving the training accuracy of autonomous driving models.

[0080] The training process of the driving data annotation model will be further explained and introduced below. Please refer to [link / reference]. Figure 5 and Figure 6 ,in, Figure 5 This is a flowchart illustrating the training process of a driving data annotation model provided in an embodiment of this application. Figure 6 The training architecture diagram of a driving data annotation model provided in this application embodiment is shown. The training process of the driving data annotation model in S320 above may include the following S410 to S460.

[0081] S410: For each training iteration, input N frames of images from the current sample driving video into the motion feature extraction module to obtain various motion features; the current sample driving video is labeled with the real location information and real text description information of each real event; where N is an integer not less than 2.

[0082] It should be noted that the driving data annotation model can be initialized before training. Each module in the initialized driving data annotation model is an initial module, such as the initial motion feature extraction module, the initial encoding and decoding module, the initial vehicle state estimation module, the initial event localization module, and the initial text generation module. After multiple iterations, the current motion feature extraction module, encoding and decoding module, vehicle state estimation module, event localization module, and text generation module can be obtained.

[0083] This invention provides a detailed explanation using a single iteration of training as an example. Multiple sample driving videos are pre-acquired. For the current round of iteration training, N frames from the current sample video can be... The input is fed into the current motion feature extraction module, which extracts motion features from N frames of images to obtain multiple motion features, where f1 to f N These are images from frame 1 to frame N.

[0084] In one embodiment, the process of inputting N frames of the current sample driving video into the motion feature extraction module to obtain various motion features in step S410 may include: extracting motion features from the N frames of the current sample driving video through the initial motion feature extraction unit in the motion feature extraction module to obtain N initial motion features.

[0085] The N initial motion features are augmented by multiple convolutional neural networks of different scales in the motion feature extraction module, resulting in multiple motion features corresponding to each scale of the convolutional neural network; wherein, each scale of the convolutional neural network outputs multiple motion features corresponding to it.

[0086] It should be noted that the motion feature extraction module in this application includes an initial motion feature extraction unit and a feature expansion unit. In order to improve the accuracy of feature extraction, the initial motion feature extraction unit can be a streaming model in a pre-trained I3D (Two-Stream Inflated 3D ConvNets, a method of inflating a 2D network into a 3D network) model. The feature expansion unit can include multiple convolutional neural networks of different scales, such as i convolutional neural networks of different scales. In order to reduce the amount of computation while ensuring the accuracy of the final text generation and the accuracy of event localization, four convolutional neural networks of different scales can be selected in practical applications, that is, i can be 4.

[0087] It is understood that in this application, the initial motion feature extraction unit can be used to extract N frames of images from the current driving video. Feature extraction is performed to obtain N initial motion features. ,in, to These represent the first to Nth initial motion features, respectively. Then, the N initial motion features are input into the feature augmentation unit for augmentation, resulting in multiple motion features corresponding to each scale. ,in, Let j represent the j-th motion feature of the i-th convolutional layer. Each scale of the convolutional neural network outputs multiple motion features; that is, the motion features output by the i-th convolutional layer are... ;

[0088] Expressed in relational form as follows:

[0089] ,in:

[0090] i∈{1,2,3,4}, Kernel represents the convolution kernel, stride represents the stride (the stride value can be m), k represents the scale of the convolution kernel (k can be 2, 4, 8, 16), and p represents zero padding of the sequence.

[0091] S420: The encoding and decoding modules process each motion feature to obtain each encoded motion feature and the encoded representation feature corresponding to each predicted event.

[0092] In practical applications, the encoding / decoding module may include an encoder and a decoder; the implementation process of the S320 may include:

[0093] The encoder encodes all motion features to obtain coded motion features at different scales.

[0094] The decoder uses a cross-attention mechanism to interactively process the initial representation features and encoded motion features corresponding to each predicted event, thereby obtaining the encoded representation features corresponding to each predicted event.

[0095] It should be noted that the encoding / decoding module in this application may include an encoder and a decoder. The encoder in this application may be a Transformer architecture encoder, capable of processing the motion features output by the motion feature extraction module. Feature encoding is performed to obtain the corresponding encoded motion features. ,in, This represents the j-th encoded motion feature associated with the i-th convolutional layer.

[0096] In this embodiment, multiple fixed events can be preset, for example, Q fixed events can be preset. Q prediction events can be randomly initialized at the decoder input following the DETR (Detection Transformer, a Transformer-based object detection model) pattern to obtain Q initialized representation features. ,in, to Let represent the 1st to the Qth initialization features respectively. During model training, the decoder can receive each encoded representation feature output by the encoder. The decoder uses a cross-attention mechanism to interact with each initialization representation feature and each encoded motion feature, and at the end, it obtains the encoded representation feature corresponding to each predicted event. ,in, to These represent the first to the Qth encoded features, respectively.

[0097] Cross-attention, in particular, computes attention across two different sequences to handle the semantic relationships between them. For example, in translation tasks where source and target language sentences need to be aligned, cross-attention is used to compute the attention weights between the two sentences.

[0098] Cross-attention is a special form of multi-head attention that splits the input tensor into two parts. and ,in, Let d1 represent the set of real numbers, then take one part as the query set X1 and the other part as the key-value set X2. The output is a tensor of size d1×d2, which gives the attention weight of each row vector to all row vectors.

[0099] Specifically, let and The calculation of cross attention is as follows:

[0100] ,in, and These are the projection matrices for learning, d1, d2, d... k Where n is the dimension of the key-value set (and also the dimension of the query set), and n is the sequence length. K and V are the query matrix, key matrix, and value matrix, respectively. This is the activation function.

[0101] In this application, the input X1 of the cross-attention can be any encoded motion feature. Input X2 can be used to initialize the representation features. Thus, the encoded representation features can be obtained through the aforementioned cross-attention mechanism. .

[0102] S430: The vehicle state estimation module processes each coded motion feature to obtain the predicted vehicle motion state, and determines the state loss value based on the predicted vehicle motion state and the real vehicle motion state corresponding to the current sample driving video.

[0103] It should be noted that, in this embodiment of the application, after encoding each motion feature through the encoder to obtain each encoded motion feature, the current vehicle state estimation module can be used to estimate the state of each encoded motion feature to obtain the predicted vehicle motion state. The predicted vehicle motion state may include multiple predicted speeds and the predicted steering angle corresponding to each predicted speed. Since this application obtains the corresponding real vehicle motion state when acquiring each sample driving video data, the state loss value corresponding to the vehicle state estimation module in this iteration training can be obtained based on each real speed and the corresponding real steering angle in the real vehicle motion state corresponding to the current sample driving video.

[0104] In one embodiment, such as Figure 7 As shown, the implementation process of S430 may include S510 to S520.

[0105] S510: The vehicle state estimation module estimates the coded motion features at different scales to obtain the predicted speed and predicted steering angle for each of the M frames in the current sample driving video; M is an integer greater than 0 and less than or equal to N.

[0106] In practical applications, the actual vehicle motion state of the current sample driving video can be obtained in advance using GPS data corresponding to the current sample driving video. The GPS data records the vehicle speed throughout the entire video. The steering angle can be obtained by calculating the change in heading angle between adjacent frames in the GPS data, thus yielding the actual speed and the corresponding actual steering angle. For example, the actual vehicle motion state of the current sample driving video is... Among them, v1, v2, ... v M-1 v M Let s1, s2, ... sm be the actual vehicle speeds corresponding to frames 1 through M in the current sample driving video. M-1 0 represents the actual turning angle corresponding to frames 1 through M. Since GPS data is recorded at time intervals rather than frame by frame, M ≤ N.

[0107] Since each sample driving video has the same length, the specific value of this M-frame can be pre-set in the vehicle state estimation module, and then the vehicle state estimation module can encode motion features at different scales. When performing state estimation, the predicted speed and predicted steering angle for each of the M frames in the current sample driving video can be obtained. Where M ≤ N.

[0108] It should be noted that, in order to improve the accuracy of the predicted speed and predicted steering angle, please refer to... Figure 8The process described above in this application for obtaining the predicted speed and predicted steering angle corresponding to each of the M frames in the current sample driving video can be implemented by the following steps S610 to S640.

[0109] S610: The first bidirectional long short-term memory network in the vehicle state estimation module is used to enhance the features of each encoded motion feature at each scale, so as to obtain the enhanced encoded motion features at each scale.

[0110] Understandably, in order to more accurately estimate the predicted speed and predicted turning angle, this application first uses a first bidirectional long short-term memory (BiLSTM) network to enhance the features of each encoded motion feature at each scale, resulting in the enhanced encoded motion features at each scale, which can be expressed by the following formula:

[0111] ,in, Let i represent the first bidirectional long short-term memory network, i∈{1,2,3,4}. This represents the encoded motion feature enhanced at the i-th scale, thus allowing us to obtain N at each scale. i An enhanced encoded motion feature.

[0112] S620: The enhanced coded motion features at each scale are downsampled using linear interpolation to obtain the downsampled coded motion features at each scale; wherein, the number of downsampled coded motion features at any scale is M.

[0113] Furthermore, to improve computational speed and reduce resource consumption, this application employs linear interpolation to downsample each enhanced coded motion feature at each scale. Specifically, it can downsample to M features, for example: ,in, to Let be the coded motion features after sampling at the i-th scale.

[0114] S630: Perform max pooling on all downsampled encoded motion features to obtain M max-pooled encoded motion features.

[0115] Understandably, in order to further preserve key information and salient features and reduce computational complexity, this application can perform max pooling on the encoded motion features after downsampling at each scale. For example, this can be achieved through the following relationship:

[0116] ,in, That is, the encoded motion features obtained after each max pooling are: ,in, to These represent the encoded motion features after the first to the Mth max pooling operations, respectively.

[0117] S640: The first multilayer perceptron network is used to convert the encoded motion features after max pooling into scalar values ​​to obtain the predicted velocity and predicted turning angle for each frame in the M frames.

[0118] In practical applications, a first-layer perceptron network (MLP) can be used to encode the motion features after each max pooling step. Converted to scalar values, which include velocity v and steering angle s, wherein the first multilayer perceptron network includes and It can be obtained through the following relationship:

[0119] ;

[0120] .

[0121] in, This represents the prediction velocity corresponding to each frame in the M frames. This represents the predicted steering angles corresponding to the first M-1 frames in the M-frame set. The steering angle of the M-th frame can be considered as 0, thus enabling the rapid and accurate acquisition of the predicted velocity and predicted steering angle for each frame in the M-frame set.

[0122] S520: Determine the state loss value based on the predicted speed and predicted steering angle corresponding to each frame in the M frames, and the real speed and real steering angle corresponding to each frame in the real vehicle motion state corresponding to the current sample driving video.

[0123] In this embodiment, the state loss value can be calculated based on each predicted speed and predicted steering angle, as well as each actual speed and actual steering angle corresponding to the current sample driving video. To improve the accuracy of the state loss value calculation, the mean-variance loss (i.e., the state loss value) can be calculated according to the following formula. Calculation of )

[0124] ,in, This represents the actual vehicle speed in the i-th frame. This represents the predicted velocity of the i-th frame. This represents the actual turning angle in the i-th frame. This represents the predicted steering angle for the i-th frame.

[0125] S440: The event localization module performs localization analysis on each encoded representation feature to obtain the predicted localization information corresponding to each predicted event. Combined with the real localization information of each real event in the current sample driving video, the localization loss value is obtained.

[0126] It should be noted that the various encoded representation features output by the decoder are input to the event localization module and the text generation module. The event localization module performs localization analysis based on each encoded representation feature to obtain the predicted event localization information corresponding to each predicted event. The current sample driving video carries labels of the real localization information of each real event, so the localization loss value can be further calculated based on the localization information of each predicted event and the localization information of each real event.

[0127] In one embodiment, the process of S440 may include:

[0128] The event localization module performs localization analysis on the encoded representation features corresponding to each predicted event to obtain the prediction start and end time and prediction event category corresponding to each predicted event.

[0129] It should be noted that, in order to enhance the event monitoring effect, this application introduces an event type allocation strategy during the training of the event localization module. By allocating event types, the event time localization is further assisted, which helps to improve the accuracy of predicted event localization.

[0130] That is, the event localization module in this application includes a second multilayer perceptron network. Furthermore, the temporal localization unit and classification unit can be used to map the encoded representation features corresponding to the Q predicted events into two-dimensional vectors through the second multilayer perceptron network in the current event localization module. , This yields the prediction start time corresponding to each prediction event. and predicted end time For example, it can be obtained through the following relation:

[0131] ;in, In Q, the first Each encoded representation feature In Q, the first The prediction start time corresponding to each encoded representation feature In Q, the first The prediction end time corresponding to each encoded representation feature.

[0132] In this embodiment of the invention, to fully utilize the information provided by events and enhance the event detection effect, events are classified while detecting time regions, and the semantic information of the categories is used to constrain the detection of the starting position of the events. Therefore, in this application, the classification unit in the current event localization module can also perform category prediction on the encoded representation features corresponding to each predicted event to obtain the predicted event category corresponding to each predicted event.

[0133] Furthermore, based on the predicted start and end times and predicted event categories corresponding to each predicted event, and combined with the actual start and end times and actual categories of each real event in the current sample driving video, the localization loss value is obtained.

[0134] Understandably, the labels of the real location information corresponding to the current sample driving video can include the start and end times of P real events, and the set of P real events can be denoted as... ,in, Let j represent the j-th real event. , This represents the start time of the j-th real event. Let represent the end time of the j-th real event. , This represents the description of driving behavior in the j-th real event. This represents the behavioral cause description of the j-th real event.

[0135] The process of obtaining the localization loss value described above may include:

[0136] The predicted start time and predicted end time of each predicted event are matched with the real start time and real end time of each real event in the current sample driving video to obtain a predicted event corresponding to each real event, thus obtaining each mapped event pair.

[0137] It should be noted that in the embodiments of this application, the predicted event set Q can be matched with each event in the real event set P. That is, predictive events that match each of the Q predicted events are found to match each of the P real events. For example, the Hungarian algorithm can be used to map each real event in P to a predicted event in Q with minimal cost.

[0138] The j-th real event in P The predicted events mapped in Q are ,but and A mapping event pair can be obtained as a single mapping event pair, and thus multiple mapping event pairs can be obtained.

[0139] For each mapped event pair, calculate the first loss value between the real event and the predicted event.

[0140] That is, for mapped event pairs and The first loss value of the mapped event pair can be calculated. Thus, the first loss value corresponding to each mapping event pair can be obtained.

[0141] After obtaining each first loss value, the time loss value can be obtained based on the first loss value corresponding to each mapped event pair; for example, the total time loss value L can be obtained by adding the first loss values ​​together. loc :

[0142] .

[0143] For each predicted event, a second loss value is calculated based on the predicted event category and the category label corresponding to the predicted event; wherein, a corresponding category label is pre-assigned to each predicted event.

[0144] It is understood that in the embodiments of this application, a set of multiple event categories can be predetermined, such as {left turn, right turn, acceleration, deceleration, parking, other}, and each predicted event can be pre-classified to determine the corresponding category label. For example, Q predicted events can be classified using a large language model GPT-4. The operation instruction can be: "Please classify the following driving behaviors according to the description. The selectable categories are {left turn, right turn, acceleration, deceleration, parking, other}. The driving behavior description is '[text description]'", thereby obtaining the category label for each predicted event. This category label is a pseudo-label y.

[0145] As described above, this application utilizes a classification unit to predict the category of each predicted event based on its encoded representation features, thereby obtaining the predicted event category for each event. Alternatively, a multilayer perceptron network can also be used to predict the predicted events. Encoding representation features Perform category prediction to obtain a label x for a predicted event category corresponding to the predicted event.

[0146] It should be noted that a second loss value corresponding to the cross-entropy loss can be further calculated based on the label x of the predicted event category corresponding to each predicted event and the pseudo-label y (which can be used as the real label) corresponding to that predicted event. The relationship of the cross-entropy loss (CrossEntropyLoss) is as follows:

[0147] ,in, This indicates that the cross-entropy loss function is calculated for x and y. Indicates the first The predicted event corresponds to the th... Labels for each category (i.e., labels for the predicted event categories). Indicates the first The predicted event is the... The true labels corresponding to each category, where C represents the number of categories and Q represents the total number of predicted events.

[0148] The category loss value is obtained based on the second loss value corresponding to each predicted event.

[0149] After obtaining the second loss value corresponding to each predicted event, the individual second loss values ​​can be summed to obtain the total category loss value L. cls .

[0150] Based on the time loss value L loc and category loss value L cls The positioning loss value is obtained.

[0151] In this embodiment of the application, L loc +L cls The overall loss for event localization detection, which is the sum of the two, is used as the localization loss value. This not only ensures the optimal correspondence between the real event and the predicted event, but also enhances the semantic similarity between the two, achieving the goal of minimizing the overall loss. Based on this, the model trained can improve the accuracy of event localization.

[0152] S450: The text generation module processes each encoded representation feature to obtain the predicted text description information corresponding to each predicted event, and combines it with the real text description information of each real event to obtain the text loss value.

[0153] It should be noted that in this embodiment, the current text module processes each encoded representation feature output by the encoder during the current iterative training process to obtain the predicted text description information corresponding to each predicted event. As described above, by matching each predicted event with each real event, each mapping event pair is obtained. Each mapping event pair includes a real event and a predicted event corresponding to that real event. For each predicted event, the corresponding real event can be determined based on the corresponding mapping event pair. Then, the real text description of the real event is obtained from the real text description information corresponding to each real event in the current sample driving video. Finally, by combining the real text description information of each real event, the text loss value can be obtained.

[0154] In one embodiment, the process implemented by S450 may include:

[0155] The second bidirectional long short-term memory network in the text generation module is used to process the encoded representation features corresponding to each predicted event to generate text description information corresponding to each predicted event.

[0156] It is understood that in this embodiment of the application, the text generation process is performed on the encoded representation features corresponding to each predicted event through the second bidirectional long short-term memory network BiLSTM based on the attention mechanism in the current text generation module. Furthermore, a maximum length limit for text generation (e.g., 50 words) can be preset. For example, if the j-th predicted event in set Q is mapped to the i-th real event in set P, the text generation process can be implemented through the following relationship:

[0157] ,in, This represents the encoded representation feature corresponding to the j-th predicted event. Let represent the sequence of predicted text description information corresponding to the j-th predicted event. Based on this relationship, a unique sequence of predicted text description information can be obtained for each predicted event.

[0158] Based on the real events corresponding to each predicted event, and based on the predicted text description information of the predicted event and the real text description information corresponding to the real event, a third loss value corresponding to the predicted event is determined.

[0159] Based on the mapping event pairs obtained above, we can obtain the real text description information of the actual event corresponding to each predicted event. This allows us to describe the information based on the generated predicted text. Corresponding real text description information The loss value is then calculated to obtain the corresponding third loss value. For example, the third loss value can be calculated using the following formula. :

[0160] ,in, Indicates to and Perform cross-entropy loss function calculation.

[0161] The text loss value is obtained based on the third loss value corresponding to each predicted event.

[0162] Understandably, after obtaining the third loss value corresponding to each predicted event, these third loss values ​​can be summed to obtain the overall text loss value L. text .

[0163] S460: Based on the state loss value, localization loss value, and text loss value, the parameters of each module in the motion feature extraction module, encoding and decoding module, vehicle state estimation module, event localization module, and text generation module are updated, and the trained driving data annotation model is obtained when the training termination condition is met.

[0164] It should be noted that, in this embodiment of the invention, after obtaining the state loss value, the localization loss value, and the text loss value, the individual loss values ​​can be summed to obtain the total loss value for this iteration of training. That is, the total loss value is: L mse +L loc +L cls +L text Then, it can be further determined whether the training termination condition is met. If not, gradient calculation and backpropagation of the neural network can be performed based on the total loss value to update the parameters of each network in each module of the driving data annotation model, thereby obtaining the updated modules and performing the next round of iterative training. If the training termination condition is met, the training ends, and the trained driving data annotation model is obtained. The training termination condition may include the total loss value being less than a preset loss value or the number of iterations reaching the maximum number of iterations. In this application, during the training process of the driving data annotation model, the loss functions of each stage are integrated to update the parameters of each module and corresponding network in the driving data annotation model, improving the accuracy of parameter updates and thus obtaining a more accurate driving data annotation model.

[0165] Therefore, this application pre-trains a driving data annotation model based on multiple sample driving videos and the corresponding real vehicle motion states for each sample driving video. During the annotation process, the current driving video is acquired and input into the driving data annotation model for analysis. This yields vehicle state information, event location information, and event text description information corresponding to each event in the current driving video. The current driving video is then annotated based on these vehicle state information, event location information, and event text description information. This application, when annotating and analyzing the current driving video, can simultaneously obtain event location information and event text description information for each event. The event text description information is not generated dependent on event location information, which improves annotation accuracy. Furthermore, vehicle state information can be obtained, thus acquiring more dimensional annotation information. This provides more dimensional data for training the autonomous driving model, thereby improving the training accuracy of the autonomous driving model.

[0166] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0167] Based on the above embodiments, this application also provides a driving data annotation device, please refer to... Figure 9 The device includes:

[0168] The first acquisition module 11 is used to acquire the current driving video of the vehicle;

[0169] Analysis module 12 is used to input the current driving video into a pre-trained driving data annotation model to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video.

[0170] The annotation module 13 is used to annotate the current driving video based on vehicle status information and event location information and event text description information corresponding to each event in the current driving video. The pre-trained driving data annotation model is obtained by the training module through model training based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.

[0171] In one embodiment, the driving data annotation model includes a motion feature extraction module, an encoding / decoding module, a vehicle state estimation module, an event localization module, and a text generation module.

[0172] Analysis module 12 includes:

[0173] The first extraction unit is used to input the current driving video into a pre-trained driving data annotation model, and extract the motion features of the current driving video through the motion feature extraction module to obtain various current motion features;

[0174] The encoding / decoding unit is used to encode and decode each current motion feature through the encoding / decoding module to obtain each current encoded motion feature and each current predicted event representation feature.

[0175] The state estimation unit is used to process each currently encoded motion feature through the vehicle state estimation module to obtain vehicle state information;

[0176] The localization unit is used to perform event localization processing on the representation features of each predicted event through the event localization module, so as to obtain the event localization information corresponding to each event in the current driving video.

[0177] The text generation unit is used to process the text of the representation features of each currently predicted event through the text generation module to obtain the event text description information corresponding to each event.

[0178] For a description of the features of the driving data annotation device in the embodiments of this application, please refer to the relevant description of the driving data annotation method in the embodiments, which will not be repeated here.

[0179] Based on the above embodiments, this application also provides a driving data annotation model training device, please refer to... Figure 10 The device includes:

[0180] The second acquisition module 21 is used to acquire multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video.

[0181] Training module 22 is used to train the model based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video, so as to obtain the trained driving data annotation model.

[0182] The driving data annotation model is used to analyze the current driving video of the vehicle to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and to annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video.

[0183] In one embodiment, the driving data annotation model includes a motion feature extraction module, an encoding / decoding module, a vehicle state estimation module, an event localization module, and a text generation module.

[0184] Training module 22 includes:

[0185] The second extraction unit is used to input N frames of images from the current sample driving video into the motion feature extraction module for each training iteration to obtain various motion features; the current sample driving video is labeled with real location information and real text description information of each real event; where N is an integer not less than 2;

[0186] The first processing unit is used to process each motion feature through the encoding and decoding module to obtain each encoded motion feature and the encoded representation feature corresponding to each predicted event.

[0187] The second processing unit is used to process each coded motion feature through the vehicle state estimation module to obtain the predicted vehicle motion state, and to determine the state loss value based on the predicted vehicle motion state and the real vehicle motion state corresponding to the current sample driving video.

[0188] The first analysis unit is used to perform localization analysis on each encoded representation feature through the event localization module to obtain the predicted localization information corresponding to each predicted event, and to obtain the localization loss value by combining the real localization information of each real event in the current sample driving video.

[0189] The third processing unit is used to process each encoded representation feature through the text generation module to obtain the predicted text description information corresponding to each predicted event, and to obtain the text loss value by combining the real text description information of each real event.

[0190] The update unit is used to update the parameters of the motion feature extraction module, encoding / decoding module, vehicle state estimation module, event localization module, and text generation module based on the state loss value, localization loss value, and text loss value, and obtain the trained driving data annotation model when the training termination condition is met.

[0191] In one embodiment, the second extraction unit includes:

[0192] The extraction subunit is used to extract motion features from N frames of the current sample driving video through the initial motion feature extraction unit in the motion feature extraction module, and obtain N initial motion features.

[0193] The expansion subunit is used to expand the N initial motion features by using multiple convolutional neural networks of different scales in the motion feature extraction module to obtain multiple motion features corresponding to each scale of the convolutional neural network; wherein, each scale of the convolutional neural network outputs multiple motion features corresponding to it.

[0194] In one embodiment, the encoding / decoding module includes an encoder and a decoder.

[0195] The first processing unit includes:

[0196] The first processing subunit is used to encode all motion features through the encoder to obtain various encoded motion features at different scales;

[0197] The second processing subunit is used to interactively process the initial representation features and coded motion features corresponding to each prediction event through the decoder using a cross-attention mechanism, so as to obtain the coded representation features corresponding to each prediction event.

[0198] In one embodiment, the second processing unit includes:

[0199] The state estimation subunit is used to estimate the coded motion features at different scales through the vehicle state estimation module, and obtain the predicted speed and predicted steering angle for each of the M frames in the current sample driving video; M is an integer greater than 0 and less than or equal to N.

[0200] The first loss value determination subunit is used to determine the state loss value based on the predicted speed and predicted steering angle corresponding to each frame in the M frames, and the real speed and real steering angle corresponding to each frame in the real vehicle motion state corresponding to the current sample driving video.

[0201] In one embodiment, the state estimation subunit includes:

[0202] The feature enhancement subunit is used to enhance each encoded motion feature at each scale through the first bidirectional long short-term memory network in the vehicle state estimation module, so as to obtain each enhanced encoded motion feature at each scale.

[0203] The downsampling processing subunit is used to downsample each enhanced coded motion feature at each scale using linear interpolation to obtain each downsampled coded motion feature at each scale; wherein, the number of downsampled coded motion features at any scale is M.

[0204] The max pooling subunit is used to perform max pooling on all downsampled encoded motion features to obtain M max pooled encoded motion features.

[0205] The first transformation subunit is used to convert the encoded motion features after max pooling into scalar values ​​using the first multilayer perceptron network, so as to obtain the predicted velocity and predicted turning angle corresponding to each frame in the M frames.

[0206] In one embodiment, the first analysis unit includes:

[0207] The location analysis subunit is used to perform location analysis on the encoded representation features corresponding to each predicted event through the event location module, so as to obtain the prediction start and end time and prediction event category corresponding to each predicted event.

[0208] The second loss value determination subunit is used to obtain the localization loss value based on the prediction start and end times and prediction event categories corresponding to each prediction event, combined with the real start and end times and real categories of each real event in the current sample driving video.

[0209] In one embodiment, the location analysis subunit includes:

[0210] The second conversion subunit is used to map the encoded representation features corresponding to each predicted event into a two-dimensional vector through the second multilayer perceptron network in the event localization module, so as to obtain the prediction start time and prediction end time corresponding to each predicted event.

[0211] The category prediction subunit is used to predict the category of each predicted event by using the classification unit in the event localization module to perform category prediction on the encoded representation features corresponding to each predicted event, thereby obtaining the predicted event category corresponding to each predicted event.

[0212] In one embodiment, the second loss value determining subunit includes:

[0213] The matching subunit is used to match the prediction start time and prediction end time corresponding to each predicted event with the real start time and real end time corresponding to each real event in the current sample driving video to obtain a predicted event corresponding to each real event, so as to obtain each mapping event pair.

[0214] The first computational subunit is used to calculate the first loss value of the real event and the predicted event for each mapping event pair;

[0215] The first determining subunit is used to obtain the time loss value based on the first loss value corresponding to each mapping event pair;

[0216] The second calculation subunit is used to calculate a second loss value for each predicted event based on the predicted event category and the category label corresponding to the predicted event; wherein, a corresponding category label is pre-assigned to each preset component.

[0217] The third determining subunit is used to obtain the category loss value based on the second loss value corresponding to each predicted event;

[0218] The fourth determination sub-unit is used to obtain the positioning loss value based on the time loss value and the category loss value.

[0219] In one embodiment, the third processing unit includes:

[0220] The first generation subunit is used to perform text generation processing on the encoded representation features corresponding to each prediction event through the second bidirectional long short-term memory network in the text generation module, so as to obtain the predicted text description information corresponding to each prediction event.

[0221] The fifth determining subunit is used to determine the third loss value corresponding to each predicted event based on the real event corresponding to each predicted event, the predicted text description information of the predicted event and the real text description information corresponding to the real event.

[0222] The sixth determination subunit is used to obtain the text loss value based on the third loss value corresponding to each predicted event.

[0223] In one embodiment, the device further includes:

[0224] The pre-acquisition module is used to pre-acquire multiple raw driving videos, which are labeled with the real location information and real text description information of each real event;

[0225] The expansion module is used to expand the original driving video samples to obtain the expanded initial sample driving videos.

[0226] The preprocessing module is used to preprocess each initial sample driving video to obtain each sample driving video; the sample driving video includes N equally spaced frames of images, where N is an integer not less than 2.

[0227] For a description of the features in the embodiment corresponding to the driving data annotation model training device in this application, please refer to the relevant description of the embodiment corresponding to the driving data annotation model training method, which will not be repeated here.

[0228] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the driving data annotation method, or to implement the steps of the driving data annotation model training method as described in any of the above-described embodiments.

[0229] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the driving data annotation method, or to implement the steps of the driving data annotation model training method as described in any of the above-described embodiments.

[0230] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0231] The embodiments of this application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-described driving data annotation method embodiments, or implements the steps of the driving data annotation model training method as described in any of the above-described embodiments.

[0232] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above-described driving data annotation method embodiments.

[0233] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0234] The foregoing has provided a detailed description of the driving data annotation method, program product, electronic device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for labeling driving data, characterized in that, include: Obtain the vehicle's current driving video; The current driving video is a driving video taken from the first-person perspective of the vehicle; The current driving video is input into a pre-trained driving data annotation model. The preprocessing module preprocesses the current driving video into a video with a preset number of frames per second. N frames are extracted from the preprocessed video at equal intervals. The motion feature extraction module extracts motion features from the N frames extracted from the preprocessed current driving video at equal intervals to obtain each current motion feature. The encoder in the encoding and decoding module encodes each current motion feature to obtain multiple current encoded motion features. The decoder in the encoding and decoding module uses a cross-attention mechanism to interact with each current encoded motion feature and the pre-stored initialization representation features of each preset time to obtain each current predicted event representation feature. The vehicle state estimation module processes each of the currently encoded motion features to obtain vehicle state information, which includes the speed and corresponding steering angle at multiple times. The event localization module performs event localization processing on the characterization features of each currently predicted event to obtain the event localization information corresponding to each event in the current driving video. The text generation module performs text generation processing on the representation features of each of the current predicted events to obtain event text description information corresponding to each event; the event text description information includes event behavior description information and / or event cause analysis information. The current driving video is labeled based on the vehicle state information, event location information, and event text description information corresponding to each event in the current driving video. The pre-trained driving data labeling model is trained using multiple sample driving videos and the real vehicle motion states corresponding to each sample driving video. During the training of the driving data labeling model, for each training iteration, the parameters in the motion feature extraction module, the encoding / decoding module, the vehicle state estimation module, the event location module, and the text generation module are updated based on the total loss value of this iteration. The trained driving data labeling model is obtained when the training termination condition is met. The total loss value is the sum of the state loss value corresponding to the vehicle state estimation module, the location loss value corresponding to the event location module, and the text loss value corresponding to the text generation module.

2. A method for training a driving data annotation model, characterized in that, include: Acquire multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video; The sample driving videos are driving videos taken from the first-person perspective of the vehicle. The model is trained based on multiple sample driving videos and the real vehicle motion state corresponding to each sample driving video to obtain the trained driving data annotation model. In the training process of the driving data annotation model, for each training round, the parameters of the motion feature extraction module, encoding / decoding module, vehicle state estimation module, event localization module, and text generation module are updated according to the total loss value of this iteration. The trained driving data annotation model is obtained when the training termination condition is met. The total loss value is the sum of the state loss value corresponding to the vehicle state estimation module, the localization loss value corresponding to the event localization module, and the text loss value corresponding to the text generation module. The driving data annotation model is used to analyze the current driving video of the vehicle to obtain vehicle status information, event location information and event text description information corresponding to each event in the current driving video, and to annotate the current driving video based on the vehicle status information and the event location information and event text description information corresponding to each event in the current driving video; wherein: The analysis of the vehicle's current driving video yields vehicle status information, event location information corresponding to each event in the current driving video, and event text description information, including: Obtain the current driving video of the vehicle; the current driving video is a driving video from the first-person perspective of the vehicle. The current driving video is input into a pre-trained driving data annotation model. The preprocessing module preprocesses the current driving video into a video with a preset number of frames per second. N frames are extracted from the preprocessed video at equal intervals. The motion feature extraction module extracts motion features from the N frames extracted from the preprocessed current driving video at equal intervals to obtain each current motion feature. The encoder in the encoding and decoding module encodes each current motion feature to obtain multiple current encoded motion features. The decoder in the encoding and decoding module uses a cross-attention mechanism to interact with each current encoded motion feature and the pre-stored initialization representation features of each preset time to obtain each current predicted event representation feature. The vehicle state estimation module processes each of the currently encoded motion features to obtain vehicle state information, which includes the speed and corresponding steering angle at multiple times. The event localization module performs event localization processing on the characterization features of each currently predicted event to obtain the event localization information corresponding to each event in the current driving video. The text generation module performs text generation processing on the representation features of each of the current predicted events to obtain event text description information corresponding to each event; the event text description information includes event behavior description information and / or event cause analysis information.

3. The driving data annotation model training method according to claim 2, characterized in that, The driving data annotation model includes a motion feature extraction module, an encoding / decoding module, a vehicle state estimation module, an event localization module, and a text generation module; The process of training a model based on multiple sample driving videos and the real vehicle motion states corresponding to each sample driving video to obtain a trained driving data annotation model includes: For each training iteration, N frames of images from the current sample driving video are input into the motion feature extraction module to obtain various motion features; the current sample driving video is labeled with the real location information and real text description information of each real event; where N is an integer not less than 2; The encoding and decoding module processes each motion feature to obtain each encoded motion feature and the encoded representation feature corresponding to each predicted event. The vehicle state estimation module processes each of the encoded motion features to obtain the predicted vehicle motion state, and determines the state loss value based on the predicted vehicle motion state and the real vehicle motion state corresponding to the current sample driving video. The event localization module performs localization analysis on each encoded representation feature to obtain predicted localization information corresponding to each predicted event, and combines it with the real localization information of each real event in the current sample driving video to obtain the localization loss value. The text generation module processes each encoded representation feature to obtain predicted text description information corresponding to each predicted event, and combines it with the real text description information of each real event to obtain the text loss value.

4. The driving data annotation model training method according to claim 3, characterized in that, The step involves inputting N frames of images from the current sample driving video into the motion feature extraction module to obtain various motion features, including: The initial motion feature extraction unit in the motion feature extraction module extracts motion features from N frames of the current sample driving video to obtain N initial motion features. The motion feature extraction module uses multiple convolutional neural networks of different scales to augment the N initial motion features, thereby obtaining multiple motion features corresponding to each scale of the convolutional neural network; wherein each scale of the convolutional neural network outputs multiple motion features corresponding to it.

5. The driving data annotation model training method according to claim 4, characterized in that, The encoding / decoding module includes an encoder and a decoder; The process of processing each motion feature through the encoding / decoding module to obtain each encoded motion feature and the encoded representation feature corresponding to each predicted event includes: The encoder encodes all the motion features to obtain coded motion features at different scales. The decoder uses a cross-attention mechanism to interactively process the initial representation features and the encoded motion features corresponding to each predicted event, thereby obtaining the encoded representation features corresponding to each predicted event.

6. The driving data annotation model training method according to claim 5, characterized in that, The process of processing each encoded motion feature through the vehicle state estimation module to obtain a predicted vehicle motion state, and determining a state loss value based on the predicted vehicle motion state and the real vehicle motion state corresponding to the current sample driving video, includes: The vehicle state estimation module estimates the encoded motion features at different scales to obtain the predicted speed and predicted steering angle for each of the M frames in the current sample driving video; M is an integer greater than 0 and less than or equal to N. The state loss value is determined based on the predicted speed and predicted steering angle corresponding to each frame in the M frames, and the real speed and real steering angle corresponding to each frame in the real vehicle motion state corresponding to the current sample driving video.

7. The driving data annotation model training method according to claim 6, characterized in that, The step of estimating the encoded motion features at different scales through the vehicle state estimation module to obtain the predicted speed and predicted steering angle for each of the M frames in the current sample driving video includes: The first bidirectional long short-term memory network in the vehicle state estimation module is used to enhance the coded motion features at each scale, so as to obtain the enhanced coded motion features at each scale. The enhanced coded motion features at each scale are downsampled using linear interpolation to obtain the downsampled coded motion features at each scale; where the number of downsampled coded motion features at any scale is M. Max pooling is performed on all downsampled encoded motion features to obtain M max-pooled encoded motion features; The first multilayer perceptron network is used to convert the max-pooled encoded motion features into scalar values ​​to obtain the predicted velocity and predicted turning angle for each frame in the M frames.

8. The driving data annotation model training method according to claim 4, characterized in that, The event localization module performs localization analysis on each encoded representation feature to obtain predicted localization information corresponding to each predicted event, and combines this with the real localization information of each real event in the current sample driving video to obtain a localization loss value, including: The event localization module performs localization analysis on the encoded representation features corresponding to each predicted event to obtain the prediction start and end time and prediction event category corresponding to each predicted event. Based on the predicted start and end times and predicted event category corresponding to each predicted event, and combined with the actual start and end times and actual categories of each real event in the current sample driving video, the localization loss value is obtained.

9. The driving data annotation model training method according to claim 8, characterized in that, The step of performing location analysis on the encoded representation features corresponding to each predicted event through the event location module to obtain the prediction start and end times and prediction event categories corresponding to each predicted event includes: The second multilayer perceptron network in the event localization module maps the encoded representation features corresponding to each predicted event into a two-dimensional vector to obtain the prediction start time and prediction end time corresponding to each predicted event. The classification unit in the event localization module performs category prediction on the encoded representation features corresponding to each predicted event to obtain the predicted event category corresponding to each predicted event.

10. The driving data annotation model training method according to claim 9, characterized in that, The localization loss value is obtained by combining the predicted start and end times and the predicted event category corresponding to each predicted event with the actual start and end times and actual categories of each real event in the current sample driving video. include: The prediction start time and prediction end time corresponding to each of the predicted events are matched with the real start time and real end time corresponding to each of the real events in the current sample driving video to obtain a predicted event corresponding to each of the real events, so as to obtain each mapping event pair; For each real event and predicted event in the mapped event pair, calculate a first loss value for the real event and the predicted event; Based on the first loss value corresponding to each of the mapped event pairs, the time loss value is obtained; For each predicted event, a second loss value is calculated based on the predicted event category and the category label corresponding to the predicted event; wherein, a corresponding category label is pre-assigned to each predicted event. Based on the second loss value corresponding to each of the predicted events, the category loss value is obtained; The positioning loss value is obtained based on the time loss value and the category loss value.

11. The driving data annotation model training method according to claim 10, characterized in that, The text generation module processes each encoded representation feature to obtain predicted text description information corresponding to each predicted event, and combines this with the real text description information of each real event to obtain a text loss value, including: The second bidirectional long short-term memory network in the text generation module is used to perform text generation processing on the encoded representation features corresponding to each predicted event, so as to obtain the predicted text description information corresponding to each predicted event. Based on the real events corresponding to each predicted event, and based on the predicted text description information of the predicted event and the real text description information corresponding to the real event, a third loss value corresponding to the predicted event is determined. The text loss value is obtained based on the third loss value corresponding to each predicted event.

12. The driving data annotation model training method according to any one of claims 2 to 9, characterized in that, Also includes: Multiple original driving videos are pre-acquired, and the original driving videos are labeled with the real location information and real text description information of each real event; Based on the original driving video, sample augmentation is performed to obtain each augmented initial sample driving video. Each of the initial sample driving videos is preprocessed to obtain a sample driving video; the sample driving video includes N equally spaced frames, where N is an integer not less than 2.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the driving data annotation method as described in claim 1, or the steps of the driving data annotation model training method as described in any one of claims 2 to 12.

14. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor is configured to implement the steps of the driving data annotation method as described in claim 1, or the steps of the driving data annotation model training method as described in any one of claims 2 to 12, when executing the computer program.

Citation Information

Patent Citations

  • Video labeling method and device, computer equipment and computer readable storage medium

    CN117710845A

  • End-to-end shipborne surveillance video dense description method and system based on dynamic feature memory

    CN117746328A

  • A method for generating information based on multimodal model and related equipment

    CN119763019A