Automatic driving video generation type prediction method guided by vehicle kinematics information

By integrating structured traffic information and kinematic parameter estimation in the generation of autonomous driving videos, the problem of failing to effectively integrate vehicle kinematic information in the prior art is solved, and autonomous driving videos that conform to physical laws are generated, which improves the accuracy and safety of predictions.

CN120378703AActive Publication Date: 2025-07-25NINGXIA UNIVERSITY

Patent Information

Application Number
CN202510516883.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing video generation method for predicting autonomous driving scenarios fails to effectively integrate vehicle kinematic information during the generation process, resulting in the generated video sequence being physically unreal or unreasonable, affecting the accuracy and practicality of the prediction.

Method used

A conditioned diffusion model is used to combine structured traffic information and kinematic parameter estimation neural network model to integrate visual characteristics and kinematic information to generate autonomous driving videos that conform to physical laws.

Benefits of technology

The generated video sequence conforms to the laws of vehicle kinematics, improves the accuracy and practicality of predictions, and ensures that the autonomous driving system has enough time to plan safe paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005372788890000054
    Figure BDA0005372788890000054
  • Figure BDA0005372788890000055
    Figure BDA0005372788890000055
  • Figure BDA0005372788890000061
    Figure BDA0005372788890000061
Patent Text Reader

Abstract

The invention provides a vehicle kinematics information guided automatic driving video generation type prediction method, belongs to the technical field of video generation of automatic driving scene prediction, and aims to solve the problem that a vehicle in a current end-to-end automatic driving model generation scene does not conform to a physical world vehicle kinematics rule. The method comprises the following steps: reconstructing an automatic driving scene video by using a conditional diffusion model and combining structured traffic information; establishing a neural network model for kinematics parameter estimation, and predicting future kinematics information of the vehicle; and utilizing a conditional diffusion model to guide generation of a future driving video by iteratively predicted kinematics information. According to the method, the future automatic driving video conforming to the kinematics law is predicted in a generated mode based on a real driving scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video generation for autonomous driving scene prediction, and particularly to a method for generating and predicting autonomous driving videos guided by vehicle kinematic information. Background Art

[0002] In recent years, autonomous driving technology is expected to become a disruptive technology in the transportation field, accelerating the intelligent transformation of the global transportation system and gradually becoming a transformative force in the transportation field. Scene prediction plays a crucial role in autonomous driving systems. It can anticipate traffic environment changes, enabling decision-makers to have sufficient time to plan reasonable routes, avoid colliding with roadside obstacles or straying into other lanes, providing key information support for decision-making and planning, and ensuring driving safety. Therefore, accurate prediction data is very important.

[0003] Currently, in the technical field of video generation for autonomous driving scene prediction, the related technologies are mainly divided into the following categories: First, based on Transformer, the core of which is the multi-head self-attention mechanism, which can process all elements in the video frame sequence in parallel, capturing global dependencies. Transformer avoids recursive calculations for sequence processing and significantly improves the efficiency of long-sequence modeling. However, such methods usually rely on pure visual data and lack explicit modeling of the kinematic laws of objects, which may lead to generated motions that do not conform to physical laws. Second, based on GAN, through adversarial training, the generative adversarial network (GAN) can create novel and realistic video content. Although GAN can produce high-quality video results, it does not impose direct constraints on physical rationality and may generate sudden motions that violate physical laws. Third, video generation based on diffusion models, which is an emerging generative AI technology that creates high-quality videos by simulating the gradual generation process of data from noise to real samples. The main role of the Unet part is to predict and remove noise. Although the Unet module can effectively predict and remove noise, the generation process is still dominated by visual features and does not explicitly integrate kinematic information, which may lead to physically unreasonable generated videos. In the current field of video generation for autonomous driving scene prediction, most methods mainly rely on visual data and deep learning models to generate video sequences of future scenes. However, these methods only consider the requirements of traffic regulations for driving speed during the generation process and do not consider the kinematic laws of vehicles in the physical world. This omission may result in physically unrealistic or unreasonable generated video sequences, thus affecting the accuracy and practicality of the prediction. Summary of the Invention

[0004] In view of this, the present invention provides a generative prediction method for autonomous driving videos guided by vehicle kinematic information. By considering vehicle kinematic information during the prediction process, the vehicle's motion trajectory conforms to the laws of motion, which can improve the accuracy of predicting the vehicle's motion trajectory, allowing the autonomous driving system sufficient time to plan a reasonable path and ensuring driving safety.

[0005] The technical solution adopted by the embodiments of the present invention to solve its technical problems is as follows:

[0006] A generative prediction method for autonomous driving videos guided by vehicle kinematic information, comprising:

[0007] Step S1, using a conditional diffusion model, combining structured traffic information to reconstruct an autonomous driving scenario video, where the structured traffic information includes a high-precision map and 3D boxes; concatenating the latent features generated by adding noise to the driving video through VAE with the structured traffic information, integrating the concatenation result and temporal weather data into the Unet of the conditional diffusion model, and sending the latent features generated after denoising by Unet into VAE to obtain the reconstructed driving video;

[0008] Step S2, establishing a neural network model for kinematic parameter estimation to predict the future kinematic information of the vehicle, including horizontal position, vertical position, inertial heading, longitudinal speed, lateral speed, yaw rate, throttle, and steering angle;

[0009] Step S3, using the conditional diffusion model to generate a future autonomous driving video guided by the future kinematic information of the vehicle;

[0010] Step S4, using a loss function to adjust the parameters of the conditional diffusion model and the neural network model.

[0011] Preferably, the step S1 includes:

[0012] Step S11, encoding the multi-frame image information of the vehicle video using a VAE variational autoencoder and continuously adding noise, and inputting it into the conditional diffusion model to obtain the latent feature Z after adding noise in the forward diffusion process, t∈[0,T]; t , t∈[0,T];

[0013] Step S12, projecting the traffic conditions onto the image plane to generate the structured traffic information, including a high-precision map and 3D boxes where N is the number of video frames, H is the height of the image, W is the width of the image, N B is the maximum number of boxes, and each 3D box is described by 16 numerical values to characterize features;

[0014] Step S13, the high-precision map is encoded through convolution, and the 3D box is encoded through convolution after Fourier embedding processing. The encoded high-precision map, the encoded 3D box, and Z t are connected to obtain a connection result P, which is input into the conditional diffusion model; the Fourier embedding processing result D is expressed as:

[0015] D = Fourier(B)

[0016] Step S14, clip generates an embedded box category feature G based on text conditions e , where the text condition is time-series weather data; G e and P are subjected to a fully connected operation to obtain a position encoding embedding E:

[0017] E = F fc ([P, G e )

[0018] In the formula, F fc (·) represents a fully connected layer operation;

[0019] Step S15, use the self-attention mechanism to integrate E with the visual feature v in the original Unet feature of the conditional diffusion model:

[0020] v = v + tanh(η)·TS(f s ([v, E]))

[0021] In the formula, η is a learnable parameter; f s ([v, E]) is the self-attention mechanism; TS(·) is a token selection operation that only considers visual tokens;

[0022] Step S16, reshape the integrated visual feature v from R N×C×H×W to R C×NHW , and add a temporal attention layer F to the conditional diffusion model to improve the dynamic consistency between frames:

[0023] F t (v) = Reshape NCHW (f s (Reshape CNHW (v + T pos )))

[0024] In the formula, T pos represents a temporal position embedding encoded by a sine function; Reshape NCHW represents the reshaping process, and f s is the self-attention mechanism;

[0025] Step S17, restore the visual feature v to the original dimension N×C×H×W, and send the multi-frame structured traffic information including the high-precision map and the 3D box, as well as the embedded box category feature G e into Unet. Unet continuously removes noise, and sends the latent features generated by Unet into VAE for decoding to obtain the reconstructed driving video.

[0026] Preferably, the step S2 includes:

[0027] Step S21, define the dynamic characteristic information for describing the vehicle system, including the state information S of the vehicle system t , the control input information U t , and the model coefficient set Φ:

[0028] where x t is the horizontal position, y t is the vertical position, θ t is the inertial heading, the longitudinal speed, the lateral speed, ω t is the yaw rate, T t is the throttle, δ t is the steering angle;

[0029] U t =[ΔT t ,Δδ t ∈R 2 , including the control amount ΔT of the vehicle throttle t and the change amount Δδ of the vehicle steering t ;

[0030] Φ={m,l f ,l r ,B f ,B r ,C f ,C r ,D f ,D r ,E f ,E r ,G f ,G r ,K f ,K r ,C m1 ,C m2 ,C r0 ,C d ,I z}∈R 20 , m is the mass of the vehicle, l f is the distance from the front axle to the vehicle's center of mass, l r is the distance from the rear axle to the vehicle's center of mass, Bf and B r are the front and rear tire stiffness coefficients, C f and C r are the front and rear tire damping coefficients, D f and D r are the front and rear tire rolling resistance coefficients, E f and E r are the front and rear tire nonlinearity coefficients, G f and G r are the front and rear tire cornering angle coefficients, K f and K r are the front and rear tire restoring moment coefficients, C m1 is the transmission model coefficient, C m2 is the transmission damping coefficient, C r0 is the rolling resistance coefficient, C d is the air resistance coefficient, I z is the moment of inertia of the vehicle; Let the known coefficient set Φ k ={m, l r , l f}; The unknown coefficient set Φ u ={B f , B r , C f , C r , D f , D r , E f , E r , G f , G r , K f , K r , C m1 , C m2 , C r0 , C d , I z};

[0031] Step S22, define the nominal range of each coefficient in Φ u , expressed as Extract features from the historical record [[X t-τ , U t-τ , …, [X t , U t of length τ through the hidden layer of the deep neural network, apply the Sigmoid activation function to calculate the output z of the obtained hidden layer, and scale the result Sigmoid(z) to range to obtain the estimated coefficient

[0032] Step S23, define the current vehicle motion state The neural network model f predicts the next vehicle motion state based on the given current state X t , U t , Φ k and the estimated coefficients as the kinematic information of the vehicle in the future:

[0033]

[0034] θ t+1 = θ t +(ω t )T s

[0035]

[0036] T t+1 = T t +Δt

[0037] δ t+1 = δ t +Δδ

[0038] where T s represents the time step; F fy represents the lateral tire force on the front wheels caused by the current vehicle motion and steering, F ry represents the lateral tire force on the rear wheels, F rx represents the longitudinal force of the current vehicle.

[0039] Preferably, the step S3 includes:

[0040] Step S31, input the initial frame image I0 and the structured traffic information {H0, B0} corresponding to I0;

[0041] Step S32, encode and flatten H0 and B0 into one-dimensional latent vectors, and each of the one-dimensional latent vectors is connected through a self-attention layer and aggregated through an MLP layer to generate a hidden state h0;

[0042] Step S33, use the LSTM-Transformer structure to encode the vehicle motion state The encoding result K t+1 includes the temporal relationship of the kinematic information of x t and x t+1 at different times captured by the LSTM network, as well as the relationship between different kinematic information at the same time captured by the Transformer;

[0043] Step S34, input K1 and h0 into the cross-attention layer Iterative updating is performed using gated recurrent units (GRUs) to obtain the structured traffic information h at future times t+1 :

[0044]

[0045] In the formula, represents cross-attention,[[]] represents the gated recurrent unit; K t+1 is the encoding result at time t + 1;

[0046] Step S35: Send h t+1 into the conditional diffusion model, and add h t+1 to the reconstructed driving video to generate a future autonomous driving video

[0047] Preferably, the loss function in step S4 is defined as:

[0048]

[0049] Among them, is the loss function of the conditional diffusion model, t is the time step, c is the conditional variable, and ∈ θ is the noise predicted by the model through the trainable parameter θ; is the predicted value and the loss function between the observed value

[0050] For the image information I ∈ R H×W×3 The feature encoding process is as follows: Use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 4×W / 4×4; Use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 16×W / 16×4. Finally, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 64×W / 64×8;

[0051] For the high-precision map H ∈ R H×W×3 The feature encoding process is as follows: Use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 4×W / 4×4; Use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 16×W / 16×4. Finally, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 64×W / 64×8;

[0052] For the 3D box in step S13 ​The feature encoding process is as follows: encoded using FourierEmbedder, with the output being N×NB×256; encoded using MLP, with the output being N×NB×512; encoded using MLP, with the output being N×NB×768;

[0053] The step S33 is for the kinematic information K∈R N×8 The feature encoding process is as follows: encoded using an LSTM network, with the output being N×64; mapped from 64 dimensions to 128 dimensions through a linear layer and adding positional encoding; encoded using a Transformer encoder, with the output being N×128.

[0054] As can be seen from the above technical solutions, for the problem that the vehicles in the scenarios generated by the current end-to-end autonomous driving models do not conform to the kinematic laws of vehicles in the physical world and it is impossible to rely on the generated autonomous driving videos to predict future driving scenarios, the vehicle kinematic information-guided autonomous driving video generative prediction method provided by the embodiments of the present invention proposes a conditional diffusion model guided by vehicle kinematic information to generate autonomous driving videos that conform to the kinematic laws of the physical world. First, input multiple frames of structured traffic information including high-precision maps and 3D boxes into the conditional diffusion model, and at the same time add a temporal attention layer to enable the conditional diffusion model to more deeply understand the complex motion conversions in the driving scenario and achieve the reconstruction of the current autonomous driving scenario video. Then, establish a neural network model for kinematic parameter estimation to predict future kinematic information. Finally, utilize the reverse diffusion process of the conditional diffusion model to guide the generation of future driving videos with the iteratively predicted kinematic information. The method of the present invention is based on real driving scenarios and predicts future autonomous driving videos that conform to kinematic laws in a generative manner. Brief Description of the Drawings

[0055] Figure 1 is a flowchart of the vehicle kinematic information-guided autonomous driving video generative prediction method of the present invention.

[0056] Figure 2 is the structured traffic condition, where the cuboid part in the figure is the 3D box information and the other part is the high-precision map information.

[0057] Figure 3 are six frames of images iteratively generated guided by kinematic information.

[0058] Figure 4 is a schematic block diagram of the conditional diffusion model. Detailed Description of the Embodiments

[0059] The following further elaborates in detail on the technical solutions and technical effects of the present invention in conjunction with the drawings of the present invention.

[0060] Referring to the problems in the background art, for the world model in current end-to-end autonomous driving, which lacks kinematic information of the real world and cannot accurately predict the future movement trajectory of vehicles, the present invention aims to provide a generative prediction method for autonomous driving videos guided by vehicle kinematic information, so as to provide a scientific basis for the accurate prediction of vehicle movement trajectories, enable the autonomous driving system to have sufficient time to plan a reasonable path, and avoid colliding with roadside obstacles or straying into other lanes.

[0061] The present invention provides a generative prediction method for autonomous driving videos guided by vehicle kinematic information. First, input multiple frames of structured traffic information including high-precision maps and 3D boxes into a conditional diffusion model to enable the conditional diffusion model to understand the structured traffic information. At the same time, add a temporal attention layer to enable the conditional diffusion model to more deeply understand the complex motion conversions in the driving scene, enhance the conditional diffusion model's ability to capture the inter-frame temporal relationship of driving videos, and thus achieve the reconstruction of the current autonomous driving scene video. Then, establish a kinematic neural network model to predict future kinematic information to assist in the generation of future driving videos. Finally, utilize the reverse diffusion process of the conditional diffusion model to guide the generation of future driving videos with the iteratively predicted kinematic information, so that the vehicle movement in the generated autonomous driving video conforms to the vehicle kinematic laws of the physical world.

[0062] The data comes from the real-world driving dataset nuScenes, including a total of 700 training videos and 150 validation videos. Each video contains approximately 20 seconds of footage captured by six surround-view cameras, and the frame rate of the video is 12 Hz. In step one, the reconstruction of the current driving video, the present invention uses the nuScenes development toolkit to obtain high-precision map annotations (lane boundaries, lane dividers) corresponding to the 12 Hz frames, and then projects them onto the image plane. The conditional diffusion model of the present invention is based on StableDiffusion v1.4. The following further describes the present invention in conjunction with the Figure 1 flowchart shown:

[0063] Step S1, use the conditional diffusion model to reconstruct the autonomous driving scene video in combination with structured traffic information, where the structured traffic information includes high-precision maps and 3D boxes; splice the latent features generated by adding noise to the driving video through VAE with the structured traffic information, integrate the splicing result, temporal weather data, and 3D box data in the conditional diffusion model to generate embedded box category features and enter the Unet of the conditional diffusion model, and send the latent features generated after denoising by Unet into VAE to obtain the reconstructed driving video;

[0064] Step S2, establish a neural network model for kinematic parameter estimation to predict the future kinematic information of the vehicle, including horizontal position, vertical position, inertial heading, longitudinal speed, lateral speed, yaw rate, throttle, and steering angle;

[0065] Step S3, use the conditional diffusion model to generate future autonomous driving videos guided by the future kinematic information of the vehicle;

[0066] Step S4, use the loss function to adjust the parameters of the conditional diffusion model and the neural network model.

[0067] Preferably, the conditional diffusion model refers to Figure 4 the structure shown. The specific implementation of Step S1 using the conditional diffusion model to reconstruct the autonomous driving scene video in combination with the structured traffic information includes:

[0068] Step S11, encode the multi-frame image information of the vehicle video using the VAE variational autoencoder and continuously add noise, and input it into the conditional diffusion model to obtain the latent feature Z after adding noise in the forward diffusion process, t ∈ [0, T]; t , t ∈ [0, T];

[0069] Step S12, project the traffic conditions onto the image plane to generate structured traffic information, including high-precision maps and 3D boxes where N is the number of video frames, H is the height of the image, W is the width of the image, N B is the maximum number of boxes, and each 3D box is described by 16 values to represent the features;

[0070] Step S13, encode the high-precision map through convolution, and after the 3D box is processed by Fourier embedding, encode it through convolution. Connect the encoded high-precision map, the encoded 3D box, and Z t to obtain the connection result P and input it into the conditional diffusion model; the Fourier embedding processing result D is expressed as:

[0071] D = Fourier(B) (1)

[0072] In formula (1), Fourier(·) is the Fourier embedding, and [·] is the connection operation.

[0073] For the image information I ∈ R of the present invention H×W×3 the feature encoding is as follows: First, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 4 × W / 4 × 4. Then, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 16 × W / 16 × 4. Finally, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 64 × W / 64 × 8.

[0074] For the high-precision map H ∈ R of the present invention H×W×3The feature encoding is as follows: First, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 4×W / 4×4. Then, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 16×W / 16×4. Finally, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 64×W / 64×8.

[0075] For the 3D box information in the present invention The feature encoding is as follows: First, use FourierEmbedder, and the output is N×NB×256. Then, use MLP, and the output is N×NB×512. Finally, use MLP, and the output is N×NB×768.

[0076] For the kinematic information K∈R in the present invention N×8 The feature encoding is as follows: First, use LSTM, and the output is N×64. Then, map the 64 dimensions to 128 dimensions through a linear layer and add positional encoding. Finally, use a Transformer encoder, and the output is N×128.

[0077] The high-precision map is represented as H, and the 3D box is represented as B. As shown in the appendix Figure 2 For the condition of spatial alignment, for example, the image information and the high-precision map information are downsampled using a series of 2D convolutional layers to ensure that the size of the final output is consistent with the size of the diffusion noise. For the 3D box, a multi-layer perceptron (MLP) layer is used for feature encoding.

[0078] Step S14, clip generates the embedded box category feature G based on the text condition e , where the text condition is the time-series weather data; perform a fully connected operation on G e and P to obtain the positional encoding embedding E:

[0079] E = F fc ([P, G e ) (2)

[0080] In formula (2), F fc (·) represents the fully connected layer operation; the CLIP feature can help generate different weather conditions for the driving video image;

[0081] Step S15, use the self-attention mechanism to integrate E with the visual feature υ in the original Unet feature without adding other conditions:

[0082] v = υ + tanh(η)·TS(f s ([v, E])) (3)

[0083] In formula (3), η is a learnable parameter; f s(υ, E) is the self-attention mechanism, which connects the visual feature υ and the feature after the positional embedding E. The self-attention mechanism captures the dependencies within the sequence and enhances the understanding of spatial information; TS(·) is the token selection operation that only considers visual tokens, which helps the conditional diffusion model focus on the visual features from the image;

[0084] Step S16, reshape the integrated visual feature υ from R N×C×H×W to R C×NHW , and add a temporal attention layer F t in the conditional diffusion model to improve the inter-frame dynamic consistency. The temporal attention layer enables the conditional diffusion model to focus on the temporal dynamics in the input data, thereby further enhancing the conditional diffusion model's ability to capture and interpret the inter-frame temporal relationship of the driving video:

[0085] F t (υ) = Reshape NCHW (f s (Reshape CNHW (υ + T pos ))) (4)

[0086] where T pos represents the temporal positional embedding encoded by the sine function; Reshape NCHW represents the reshaping process, and f s is the self-attention mechanism;

[0087] Step S17, restore the visual feature v to the original dimension N×C×H×W, and send the multi-frame structured traffic information including the high-precision map, 3D boxes, and the embedded box category feature G e into Unet. Unet continuously removes noise, and sends the latent features generated by Unet into VAE for decoding to obtain the reconstructed driving video. In this step, the present invention trains for 10 epochs with a batch size of 1, the video frame length N = 32, and the spatial size is 448×256.

[0088] Preferably, the specific implementation of step S2 to establish a neural network model for kinematic parameter estimation and predict the future kinematic information of the vehicle includes:

[0089] Step S21, define the dynamic characteristic information for describing the vehicle system, including the state information S t of the vehicle system, the control input information U t , and the model coefficient set Φ:

[0090] where x t is the horizontal position, y t is the vertical position, θ t is the inertial heading, Longitudinal speed Lateral speed, ω t Is the yaw rate, T t Is the throttle, δ t Is the steering angle;

[0091] U t =[ΔT t , Δδ t ∈ R 2 , including the control amount ΔT of the vehicle throttle t And the change amount Δδ of the vehicle steering t ;

[0092] Φ = {m, l f , l r , B f , B r , C f , C r , D f , D r , E f , E r , G f , G r , K f , K r , C m1 , C m2 , C r0 , C d , I z} ∈ R 20 , m is the mass of the vehicle, l f Is the distance from the front axle to the vehicle's center of mass, l r Is the distance from the rear axle to the vehicle's center of mass, B f And B r Are the front and rear tire stiffness coefficients, C f And C r Are the front and rear tire damping coefficients, D f And D r Are the front and rear tire rolling resistance coefficients, E f And E r Are the front and rear tire nonlinear coefficients, G f And G r Are the front and rear tire sideslip angle coefficients, K f And K r Are the front and rear tire restoring moment coefficients, C m1 Is the transmission model coefficient, C m2 Is the transmission damping coefficient, C r0 Is the rolling resistance coefficient, C d Is the air resistance coefficient, I z Is the moment of inertia of the vehicle; Let the known coefficient set Φ k ={m, lr , l f , set of unknown coefficients Φ u = {B f , B r , C f , C r , D f , D r , E f , E r , G f , G r , K f , K r , C m1 , C m2 , C r0 , C d , I z};

[0093] Step S22, define the nominal range of each coefficient in Φ, denoted as u (the maximum and minimum values can be the extreme values from a data set of multiple different types of vehicles), the evolution of the model state (This maximum and minimum values can be the extreme values from a data set of multiple different types of vehicles), the evolution of the model state is controlled by the state and the control input U t = [ΔT t , Δδ t ;

[0094] Feature extraction is performed on the historical records of length τ [[X t-τ , U t-τ , …, [X t , U t through the hidden layer of the deep neural network. The Sigmoid activation function is applied to calculate the output z of the obtained hidden layer, and the result Sigmoid(z) is scaled to range to obtain the estimated coefficient

[0095] Step S23, define the current vehicle motion state The neural network model f predicts the next vehicle motion state t , U t , Φ k and the estimated coefficient as the future kinematic information of the vehicle:

[0096]

[0097]

[0097] θ t+1 = θ t + (ω t)T s (8)

[0098]

[0099] T t+1 =T t +ΔT (12)

[0100] δ t+1 =δ t +Δδ (13)

[0101] Where, T s represents the time step; F fy represents the lateral tire force on the front wheel caused by the current vehicle motion and steering, F ry represents the lateral tire force on the rear wheel, F rx Indicates the current longitudinal force of the vehicle.

[0102] Preferably, step S3 includes:

[0103] Step S31, inputting the initial frame image I0 and the structured traffic information {H0, B0} corresponding to I0;

[0104] Step S32, H0 and B0 are encoded and flattened into one-dimensional latent vectors, each of which is connected through a self-attention layer and aggregated through an MLP layer to generate a hidden state h0;

[0105] Step S33: Use the LSTM-Transformer structure to analyze the vehicle motion state. Encode, the encoding result K t+1 Includes x captured by the LSTM network t With x t+1 The temporal relationship of kinematic information at different moments, and the relationship between different kinematic information at the same moment captured by Transformer; the encoding process of the LSTM-Transformer structure is as follows:

[0106] LSTM captures x t With x t+1 The temporal relationship of: The forget gate is based on the kinematic information x t+1 (K∈R N×8 )'s input dynamically retains how many historical states x t ; The input gate writes the kinematic changes at time t+1 into the cell state and updates the vehicle state; the output gate outputs the vehicle state at the current moment into a hidden state N×64;

[0107] The Transformer captures the relationships between different kinematic information at the same moment: the hidden state of the LSTM is linearly transformed to 128 dimensions, adapted to the standard input dimension of the Transformer, and then positional encoding is added to ensure that the Transformer can perceive the order of the sequence;

[0108] The multi-head attention mechanism captures the global dependencies at the same time step, and the final output is N×128 as the encoded result K t+1 。

[0109] In step S34, K1 and h0 are input and used in the cross-attention layer Iterative updates are performed using the gated recurrent unit GRUs to obtain the structured traffic information h at future moments t+1 :

[0110]

[0111] In Equation (14), represents cross-attention, represents the gated recurrent unit; these hidden states h t are concatenated with the kinematic information K t and decoded by the gated recurrent unit into future traffic structure information; K t+1 is the encoded result at time t+1; K1 is the initial value calculated based on the initial state x0 and the predicted state x1.

[0112] In step S35, h t+1 is fed into the conditional diffusion model, and h t+1 is added to the reconstructed driving video to generate the future autonomous driving video As shown in the appendix Figure 3 In this step, the present invention trains for 10 epochs with a batch size of 1, and the learning rate is set to 5×10 -5 。

[0113] Preferably, the loss function in step S4 is defined as:

[0114]

[0115] Equation (15) contains two parts, is the loss function of the conditional diffusion model, t is the time step, c is the conditional variable, and ε θ is the noise predicted by the model through the trainable parameters θ; is the predicted value and the observed value between the loss function.

[0116] The present invention gives the estimated coefficient The value is as follows:

[0117] B f / r The minimum value of the simulated data is 5.0, the maximum value of the simulated data is 30.0, the minimum value of the real data is 5.0, and the maximum value of the real data is 30.0. C f / r The minimum value of the simulated data is 0.5, the maximum value of the simulated data is 2.0, the minimum value of the real data is 0.5, and the maximum value of the real data is 2.0. D f / r The minimum value of the simulated data is 0.1, the maximum value of the simulated data is 1.9, the minimum value of the real data is 100.0, and the maximum value of the real data is 10000.0. E f / r The minimum value of the simulated data is -2.0, the maximum value of the simulated data is 0.0, the minimum value of the real data is -2.0, and the maximum value of the real data is 0.0. G f / r The minimum value of the simulated data is -0.02, the maximum value of the simulated data is 0.02, the minimum value of the real data is -0.02, and the maximum value of the real data is 0.02. K f / r The minimum value of the simulated data is -0.003, the maximum value of the simulated data is 0.003, the minimum value of the real data is -300.0, and the maximum value of the real data is 300.0. C m1 (N) The minimum value of the simulated data is 0.1435, the maximum value of the simulated data is 0.574, the minimum value of the real data is 500.0, and the maximum value of the real data is 2000.0. C m2 (kg s -1 ) The minimum value of the simulated data is 0.0273, the maximum value of the simulated data is 0.109, the minimum value of the real data is 0.0, and the maximum value of the real data is 1.0. C r0 (N) The minimum value of the simulated data is 0.0259, the maximum value of the simulated data is 0.1036, the minimum value of the real data is 0.1, and the maximum value of the real data is 1.4. C d (kg m -1 ) The minimum value of the simulated data is 1.75e -4 , the maximum value of the simulated data is 7.0e -4 , the minimum value of the real data is 0.1, and the maximum value of the real data is 1.0. I z (kgm 2 ) The minimum value of the simulated data is 1.39e -5 , the maximum value of the simulated data is 5.56e -5 , the minimum value of the real data is 500.0, and the maximum value of the real data is 2000.0.

[0118] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects.

[0119] (1) Incorporate the kinematic information to be modeled into the end-to-end autonomous driving video generation world model, and the generated video will be more accurate and meaningful.

[0120] (2) Kinematic parameters such as the position, driving speed, and steering angle of the vehicle follow strict physical laws, reflecting the authenticity of generation.

[0121] (3) This invention enables the provision of a scientific basis for the accurate prediction of the vehicle's motion trajectory, allowing the autonomous driving system to have sufficient time to plan a reasonable path and avoid colliding with roadside obstacles or straying into other lanes.

[0122] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A method for generating and predicting an autonomous driving video guided by vehicle kinematic information, characterized in that Including: Step S1: Using a conditional diffusion model, reconstruct an autonomous driving scenario video in combination with structured traffic information, where the structured traffic information includes a high-precision map and 3D bounding boxes; Add noise to the driving video and splice the latent features generated by VAE with the structured traffic information. Integrate the splicing result and the temporal weather data into the Unet of the conditional diffusion model, and send the latent features generated after denoising by Unet into VAE to obtain the reconstructed driving video; Step S2: Establish a neural network model for kinematic parameter estimation to predict the future kinematic information of the vehicle, including horizontal position, vertical position, inertial heading, longitudinal speed, lateral speed, yaw rate, throttle, and steering angle; Step S3: Use the conditional diffusion model to generate a future autonomous driving video guided by the future kinematic information of the vehicle; Step S4: Use a loss function to adjust the parameters of the conditional diffusion model and the neural network model.

2. The method for generating an autonomous driving video prediction guided by vehicle kinematic information according to claim 1, wherein The step S1 includes: Step S11, encode the multi-frame image information of the vehicle video using a VAE variational autoencoder and continuously add noise, input it into the conditional diffusion model, and obtain the latent feature Z after adding noise in the forward diffusion process t , t ∈ [0, T]; Step S12, project the traffic conditions onto the image plane to generate the structured traffic information, including a high-precision map and 3D bounding boxes where N is the number of video frames, H is the height of the image, W is the width of the image, N B is the maximum number of bounding boxes, and each 3D bounding box is described by 16 values to characterize features; Step S13, the high-precision map is encoded through convolution, and the 3D box is encoded through convolution after Fourier embedding processing. The encoded high-precision map, the encoded 3D box, and Z t are connected to obtain a connection result P and input it into the conditional diffusion model; the Fourier embedding processing result D is expressed as: D = Fourier(B) Step S14, clip generates the embedding box category feature G based on the text condition e , where the text condition is time-series weather data; take G e and P for a fully connected operation to obtain the position encoding embedding E: E = F fc ([P, G e ) where F fc (·) represents a fully connected layer operation; Step S15: Use the self-attention mechanism to integrate E with the visual feature v in the original Unet feature of the conditional diffusion model: v = v + tanh(η)·TS(f s ([v, E])) where η is a learnable parameter; f s ([v, E]) is the self-attention mechanism; TS(·) is the token selection operation that only considers visual tokens; Step S16, transform the integrated visual feature v from R N×C×H×W to R C×NHW , and add a temporal attention layer F t to the conditional diffusion model to improve the dynamic consistency between frames: F t (v) = Reshape NCHW (f s (Reshape CNHW (v + T pos ))) where T pos represents the time position embedding encoded by the sine function; Reshape NCHW represents the reshaping process, and f s is the self-attention mechanism; Step S17, restore the visual feature v to the original dimension N×C×H×W, and send the multi-frame structured traffic information including the high-precision map, the 3D box, and the embedded box category feature G e into Unet. Unet continuously removes noise, and sends the latent features generated by Unet into VAE for decoding to obtain the reconstructed driving video.

3. The method for generating an autonomous driving video prediction guided by vehicle kinematic information according to claim 2, wherein, The step S2 includes: Step S21, define the dynamic characteristic information for describing the vehicle system, including the state information S of the vehicle system t , the control input information U t , and the model coefficient set Φ: where, x t is the horizontal position, y t is the vertical position, θ t is the inertial heading, the longitudinal speed, the lateral speed, ω t is the yaw rate, T t is the throttle, δ t is the steering angle; U t = [ΔT t , Δδ t ∈ R 2 , including the control amount ΔT of the vehicle throttle t and the change amount Δδ of the vehicle steering t ; Φ = {m, l f , l r , B f , B r , C f , C r , D f , D r , E f , E r , G f , G r , K f , K r , C m1 , C m2 , C r0 , C d , I z} ∈ R 20 , m is the mass of the vehicle, l f is the distance from the front axle to the vehicle's center of mass, l r is the distance from the rear axle to the vehicle's center of mass, B f and B r are the front and rear tire stiffness coefficients, C f and C r are the front and rear tire damping coefficients, D f and D r are the front and rear tire rolling resistance coefficients, E f and E r are the front and rear tire nonlinear coefficients, G f and G r are the front and rear tire sideslip angle coefficients, K f and K r are the front and rear tire restoring moment coefficients, C m1 is the transmission model coefficient, C m2 is the damping coefficient of the gearbox, C r0 is the rolling resistance coefficient, C d is the air resistance coefficient, I z is the moment of inertia of the vehicle; Let the known coefficient set Φ k ={m, l r , l f}, and the unknown coefficient set Φ u ={B f , B r , C f , C r , D f , D r , E f , E r , G f , G r , K f , K r , C m1 , C m2 , C r0 , C d , I z}; Step S22, define Φ u the nominal range of each coefficient in through the hidden layer of the deep neural network for historical records of length τ [[X t-τ , U t-τ ,..., [X t , U t for feature extraction, apply the Sigmoid activation function to calculate the output z of the obtained hidden layer, and scale the result Sigmoid(z) to the range to obtain the estimated coefficient Step S23, define the current vehicle motion state Neural network model f predicts the next vehicle motion state based on the given current state X t , U t , Φ k and the estimated coefficients as the kinematic information of the vehicle in the future: ​ θ t+1 = θ t +(ω t )T s T t+1 = T t + ΔT δ t+1 = δ t + Δδ Where T s represents the time step; F fy represents the lateral tire force on the front wheels caused by the current vehicle motion and steering, F ry represents the lateral tire force on the rear wheels, F rx represents the longitudinal force of the current vehicle.

4. The method for generating an autonomous driving video prediction guided by vehicle kinematic information according to claim 3, wherein The step S3 includes: Step S31: Input the initial frame image I0 and the corresponding structured traffic information {H0, B0}; Step S32: Encode and flatten H0 and B0 into one-dimensional latent vectors, and connect and aggregate each one-dimensional latent vector through a self-attention layer and an MLP layer to generate a hidden state h0; Step S33, use the LSTM-Transformer structure to encode the vehicle motion state and the encoding result K t+1 includes the temporal relationship of kinematic information at different times captured by the LSTM network, and the relationship between different kinematic information at the same time captured by the Transformer; t between x t+1 and x at different times Step S34, K1 and h0 are input and processed using a cross-attention layer and iteratively updated using gated recurrent units (GRUs) to obtain the structured traffic information h at future times t+1 : In the formula, represents cross-attention, represents a gated recurrent unit; K t+1 is the encoding result at time t+1; Step S35, send h t+1 into the conditional diffusion model, and add h t+1 to the reconstructed driving video to generate a future autonomous driving video 5. The vehicle kinematic information-guided autonomous driving video generation prediction method according to claim 4, characterized in that, The loss function in the step S4 is defined as: Among them, is the loss function of the conditional diffusion model, t is the time step, c is the conditional variable, and ∈ θ is the noise predicted by the model through the trainable parameter θ; is the predicted value and the observed value is the loss function between them.

6. The method for generating and predicting an autonomous driving video guided by vehicle kinematic information according to claim 4, characterized in that: The step S11 is for the image information I ∈ R H×W×3 The feature encoding process is as follows: Use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 4×W / 4×4; Use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 16×W / 16×4. Finally, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 64×W / 64×8; The feature encoding process for the high-precision map H∈R H×W×3 is as follows: Use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 4×W / 4×4; use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 16×W / 16×4. Finally, use a 2D convolutional layer with a 4×4 convolutional kernel and a stride of 4, and the output is H / 64×W / 64×8; The above step S13 for the 3D box The feature encoding process is as follows: Use FourierEmbedder for encoding, and the output is N×NB×256; use MLP for encoding, and the output is N×NB×512; use MLP for encoding, and the output is N×NB×768; The feature encoding process for the kinematic information K ∈ R N×8 is as follows: Use an LSTM network for encoding, with the output being N × 64; map the 64 dimensions to 128 dimensions through a linear layer and add positional encoding; Use a Transformer encoder, and the output is N×128.

Citation Information

Patent Citations

  • Scene-level multi-agent track generation method and device based on consistent diffusion

    CN117473032A

  • Automatic driving decision-making method and system based on generative world large model and multi-step reinforcement learning

    CN118790287A

  • Driving condition data enhancement method and related equipment

    CN119646503A

  • Physics Embedded Neural Network Dynamics Model Structure of Autonomous Vehicle, Awareness of Hazardous Driving Situation and Stable Driving Methodology Using Latent Variables in the Model Hidden Layer

    KR102614816B1

  • Autonomous driving system

    WO2025072475A1

Cited By

  • Driver attention prediction method, system and device, storage medium and product

    CN120612675A

  • Decision model optimization method and device based on world model, medium and product

    CN120735801A

  • Automatic driving safety key multi-view video automatic generation method and device

    CN120954236A

  • An automatic driving safety-critical multi-view video automatic generation method and device

    CN120954236B

  • A driving video generation method based on natural language instructions

    CN122783708A