Automatic driving scene prediction method and related equipment

By introducing multimodal perception fusion and physical constraint diffusion models, the problems of data fusion error and robustness in autonomous driving scene prediction are solved, more accurate and stable trajectory generation and risk assessment are achieved, and the decision-making ability of the autonomous driving system is improved.

CN120633376APending Publication Date: 2025-09-12VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510542839.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing autonomous driving scene prediction methods have difficulties in time synchronization, spatial registration errors and information expression differences in multimodal sensor data fusion, resulting in insufficient prediction accuracy and stability. In addition, traditional data-driven methods lack robustness and generalization capabilities in complex traffic environments.

Method used

Through multimodal perception fusion, a physical constraint diffusion model is introduced for trajectory generation and risk assessment, including spatiotemporal alignment feature fusion, physical constraint loss optimized diffusion model and risk assessment, which improves the accuracy and robustness of prediction.

Benefits of technology

It effectively improves the accuracy, rationality and robustness of autonomous driving scenario predictions, can provide more reliable decision-making basis in complex traffic environments, avoid unrealistic trajectory generation, and maintain prediction performance in the absence of modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633376A_ABST
    Figure CN120633376A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving scene prediction method and related equipment, and relates to the technical field of automatic driving, and the method comprises the steps: obtaining the multi-mode sensor data of a target vehicle; carrying out feature fusion on the multi-modal sensor data to obtain a time-space aligned target fusion feature; performing diffusion generation on the target fusion features based on a physical constraint diffusion model to obtain a multi-modal future trajectory; and performing risk assessment on the multi-modal future trajectory to obtain an automatic driving prediction result of the target vehicle. According to the invention, through multi-modal perception fusion, the physical constraint diffusion model is introduced to carry out trajectory generation and risk assessment, and the accuracy, rationality and robustness of automatic driving scene prediction can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and more specifically, to an autonomous driving scene prediction method and related equipment. Background Art

[0002] With the rapid development of intelligent driving technology, autonomous driving systems are playing an increasingly important role in improving driving safety and reducing driving burdens. Scene prediction, a key component of autonomous driving, can predict the future behavior of the target vehicle and surrounding traffic participants, providing effective decision-making for the vehicle-mounted system. Currently, autonomous driving scene prediction primarily relies on multimodal sensor data, such as cameras, lidar, and millimeter-wave radar. However, due to difficulties in temporal synchronization, spatial registration errors, and differences in information representation between different modalities, the data fusion process is complex and prone to errors, which in turn affects the accuracy and stability of predictions.

[0003] Existing technologies often use traditional data-driven methods for trajectory prediction, such as recurrent neural networks or graph neural networks based on sequence modeling. While these methods can capture historical motion patterns, they are prone to generating trajectories that deviate from traffic rules or lack dynamic rationality, such as abrupt turns or unrealistic speed changes, thereby reducing the model's usability and safety in real-world driving environments. Furthermore, when there are modalities missing or quality degradation in perception data, the robustness and generalization capabilities of existing methods are difficult to guarantee, limiting the adaptability of autonomous driving systems in complex traffic environments. In other words, relevant technologies generally suffer from poor accuracy, rationality, and robustness in autonomous driving scenario prediction. Summary of the Invention

[0004] The Summary of the Invention section of this application introduces a series of simplified concepts that will be further described in detail in the Detailed Description of the Invention section. The Summary of the Invention section of this application is not intended to limit the key features and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.

[0005] The autonomous driving scene prediction method and related equipment provided in this application can generate trajectories and conduct risk assessment through multimodal perception fusion and the introduction of a physical constraint diffusion model, which can effectively improve the accuracy, rationality and robustness of autonomous driving scene prediction.

[0006] In a first aspect, the present application provides an autonomous driving scenario prediction method, comprising: acquiring multimodal sensor data of a target vehicle; performing feature fusion on the multimodal sensor data to obtain a target fusion feature that is aligned in time and space; performing diffusion generation on the target fusion feature based on a physical constraint diffusion model to obtain a multimodal future trajectory, wherein the physical constraint diffusion model is a model obtained by optimizing the generation process of the diffusion model through physical constraint loss; performing risk assessment on the multimodal future trajectory to obtain an autonomous driving prediction result of the target vehicle.

[0007] In some embodiments, the feature fusing of the multimodal sensor data to obtain a spatiotemporally aligned target fusion feature includes: performing interpolation alignment on the multimodal sensor data according to a preset timestamp set to obtain intermediate sensor data; performing feature extraction on the intermediate sensor data to obtain multimodal feature data; and performing feature fusing of the multimodal feature data through a cross-attention module to obtain the target fusion feature.

[0008] In some embodiments, the cross-attention module is used to perform feature fusion on the multimodal feature data to obtain the target fusion feature, including: adjusting the weight of the multimodal feature data according to the environmental condition information and sensor status of the target vehicle to obtain environmental right confirmation feature data; and performing feature fusion on the environmental right confirmation feature data through the cross-attention module to obtain the target fusion feature.

[0009] In some embodiments, the method of diffusing the target fusion features based on the physical constraint diffusion model to obtain the multimodal future trajectory includes: compressing the target fusion features through a latent space compression module to obtain low-dimensional latent space features; and inputting the low-dimensional latent space features into the physical constraint diffusion model to generate the multimodal future trajectory.

[0010] In some embodiments, inputting the low-dimensional latent space features into the physical constraint diffusion model to generate the multimodal future trajectory includes: inputting the low-dimensional latent space features into the physical constraint diffusion model so that the physical constraint diffusion model performs reverse denoising based on noise prediction loss and the physical constraint loss to generate the multimodal future trajectory, wherein the physical constraint loss includes at least one of acceleration constraint loss, curvature constraint loss, and vehicle dynamics model matching loss.

[0011] In some embodiments, the latent space compression module is a variational autoencoder, and the compression rate of the variational autoencoder is 1 / 16.

[0012] In some embodiments, a risk assessment is performed on the multimodal future trajectories to obtain an autonomous driving prediction result for the target vehicle, including: screening the multimodal future trajectories according to preset vehicle dynamics constraints to obtain a set of compliant future trajectories; determining a risk quantification score based on the collision probability and traffic rule violation probability of each compliant future trajectory in the set of compliant future trajectories; and sorting the set of compliant future trajectories according to the risk quantification score to obtain the autonomous driving prediction result.

[0013] In the second aspect, the present application also provides an autonomous driving scene prediction device, including: a data acquisition unit for acquiring multimodal sensor data of a target vehicle; a feature fusion unit for performing feature fusion on the multimodal sensor data to obtain a target fusion feature that is aligned in time and space; a trajectory prediction unit for performing diffusion generation on the target fusion feature based on a physical constraint diffusion model to obtain a multimodal future trajectory, wherein the physical constraint diffusion model is a model obtained by optimizing the generation process of the diffusion model through physical constraint loss; a risk assessment unit for performing risk assessment on the multimodal future trajectory to obtain an autonomous driving prediction result of the target vehicle.

[0014] In a third aspect, the present application also provides an electronic device comprising: a memory and a processor, wherein the processor is configured to implement the steps of the autonomous driving scene prediction method described in the first aspect when executing a computer program stored in the memory.

[0015] In a fourth aspect, the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the autonomous driving scene prediction method described in the first aspect.

[0016] In a fifth aspect, the present application also provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the autonomous driving scene prediction method provided in the embodiment of the present application is implemented.

[0017] In summary, this application achieves information complementarity and redundancy enhancement by acquiring multimodal sensor data from the target vehicle and fusing the features of different modal data. The time difference and spatial error between different sensors are processed through the spatiotemporal alignment mechanism, which improves the accuracy of the fused features and thus enhances the perception ability of complex traffic scenes. In the trajectory prediction stage, the physical constraint diffusion model is introduced and integrated into the diffusion model generation process as a loss function. This not only improves the rationality and feasibility of the generated trajectory, but also effectively avoids unrealistic trajectories that are prone to occur in traditional data-driven prediction methods, such as sudden lane changes and sharp turns. Since the diffusion model is used to model the joint distribution of multimodal features, even if there are modal missing in the input data, the missing information can still be reconstructed through the learned inter-modal relationship, thereby maintaining the prediction performance. This can greatly improve the robustness and generalization ability of the autonomous driving scene prediction method in actual complex scenarios. By performing risk assessment on the generated multimodal future trajectory, the risk situations that the target vehicle may face in the future can be more accurately judged, thereby providing a decision-making basis for the autonomous driving process. In summary, the autonomous driving scenario prediction method provided in this application can effectively improve the accuracy, rationality and robustness of autonomous driving scenario prediction by integrating multimodal perception and introducing a physical constraint diffusion model for trajectory generation and risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present description. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0019] Figure 1 A flowchart of an autonomous driving scenario prediction method provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of the structure of an autonomous driving scenario prediction device provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] Terms in the specification, claims, and drawings of this application, such as "first," "second," "third," "fourth," and the like (if any), are used to distinguish between similar objects, rather than to describe a particular order or precedence. Therefore, it is understood that these terms can be used interchangeably where appropriate, so that the embodiments described can be implemented in a different order, unless otherwise specified in the drawings or descriptions. In addition, the terms "is" and "has" and any variations thereof in this application are intended to cover all possible constituent elements on a non-exclusive basis. For example, a process, method, system, product, or apparatus that includes several steps or units is not necessarily limited to the steps or units that are explicitly listed, but may also include other steps or units that are not explicitly listed, or steps or units that are inherent to the process, method, product, or apparatus.

[0023] In this application, a "module" or "unit" refers to a computer program or part of a computer program that has a specific function and works in conjunction with other related parts to achieve a predetermined goal. These modules or units can be implemented by software, hardware (such as processing circuits or memories), or a combination of the two. One or more processors or memories can implement one or more modules or units. At the same time, each module or unit can also be part of a larger module or unit.

[0024] The technical solutions in this application will be described in detail below in conjunction with the accompanying drawings in the embodiments. It should be noted that the embodiments described are only part of this application, not all embodiments. In the following description, the "some embodiments" mentioned are only a subset of all possible embodiments, which may be the same or different subsets, and different embodiments can be combined with each other without conflict.

[0025] Figure 1 This is a flow chart of an autonomous driving scene prediction method provided by an embodiment of the present application. Figure 1 The autonomous driving scene prediction method provided in the embodiment of the present application may include the following steps 101 to 104:

[0026] Step 101, obtaining multimodal sensor data of a target vehicle;

[0027] In some examples, the target vehicle is the vehicle currently being predicted or monitored, and may be a vehicle equipped with an autonomous driving system. Autonomous driving refers to a driving method in which a computer system controls the vehicle's movement, eliminating or minimizing human driver intervention. Multimodal sensor data is perception information derived from multiple types of sensors on the target vehicle, covering vision (such as cameras), distance (such as radar and lidar), positioning (such as GPS / IMU), vehicle status (such as CAN bus data), environmental perception, and other data, and can be used to fully perceive the vehicle and its surrounding environment.

[0028] By implementing step 101, multimodal sensor data such as cameras, radars, and lidars are obtained, which can fully perceive the vehicle's environmental information and its own dynamic state from multiple dimensions. The different modalities have information complementarity, which can improve the integrity and reliability of the data and provide rich and diverse original inputs for subsequent feature extraction and scene prediction.

[0029] Step 102: perform feature fusion on the multimodal sensor data to obtain a target fusion feature aligned in time and space;

[0030] In some examples, the perception data from different types of sensors can be collaboratively processed and integrated in a unified feature space to obtain richer and more complementary expression information. Fusion methods may include timestamp interpolation, coordinate transformation, synchronization algorithm, splicing, weighted averaging, attention mechanism and other methods. The spatiotemporal aligned target fusion features refer to fusion features that are aligned in time and space, which can ensure that the environmental state expressed by each sensor data is consistent at the same time and in the same physical scene. First, a unified timestamp set can be used, such as the camera frame as the main time base, to interpolate or sample the data of other sensors. Then, the coordinate systems of different sensors, such as radar coordinates and image pixel coordinates, can be unified into the vehicle coordinate system or the world coordinate system to obtain the target fusion features.

[0031] For example, we can first interpolate and spatially transform multimodal data from cameras, radars, GPS, etc. according to a unified timestamp to achieve spatiotemporal alignment; then, we can fuse the aligned features through a deep feature extraction network and a cross-attention mechanism to output a unified representation of the target vehicle environment perception features, namely the target fusion features.

[0032] Through the implementation of step 102, the semantic, geometric and dynamic features in different modal data can be fused, and combined with the spatiotemporal alignment mechanism, the sampling time difference and spatial position error between different sensors can be resolved, which can effectively enhance the accuracy and consistency of the fused features and significantly improve the understanding and perception capabilities of the physical constraint diffusion model for complex traffic scenarios in subsequent steps.

[0033] Step 103: Based on a physical constraint diffusion model, the target fusion feature is diffused and generated to obtain a multimodal future trajectory, wherein the physical constraint diffusion model is a model obtained by optimizing the generation process of the diffusion model through physical constraint loss;

[0034] In some examples, the physical constraint diffusion model is a generative model that introduces physical constraint losses to regulate the generation process based on the traditional diffusion generation model. The physical constraint diffusion model can not only learn complex probability distributions, but also ensure that the generation results are physically reasonable and feasible. The target fusion features are input into the physical constraint diffusion model, and multimodal future trajectories can be generated through gradual denoising (backward diffusion). This process is guided by physical constraints to ensure the continuity and rationality of the trajectory in dimensions such as time, position, and speed. The multimodal future trajectory refers to the position information trajectory of the target vehicle in multiple future time steps, which is predicted after fusing multiple perception information. It contains displacement, speed, angle, etc. and can be used for autonomous driving decisions such as collision prediction and path planning.

[0035] For example, let the generated multimodal future trajectory τ={p0,p1,……,p T}, where p t =(x t ,y t ) identifies the position of the target vehicle at time step t. The physical constraints are defined as follows:

[0036] The acceleration is calculated by the second-order difference of the trajectory planned by the autonomous driving system, and the calculation formula is as follows:

[0037]

[0038] Where a t is the instantaneous acceleration vector of the target vehicle at time step t, p t+1 represents the position vector of the target vehicle at time step t+1, p t represents the position vector of the target vehicle at time step t, p t-1 represents the position vector of the target vehicle at time step t-1, Δt 2 Represents the square of the time interval between two adjacent time steps.

[0039] The acceleration constraint condition is that the acceleration of the vehicle does not exceed the maximum capacity of the vehicle a max , deriving the acceleration physical constraint Where, ||a t || represents the acceleration vector a t The Euclidean norm of ∑max() represents the accumulation of square penalties for the part that exceeds the acceleration upper limit at each time step, ensuring that losses only occur when physical constraints are exceeded.

[0040] Curvature inverse mapping of turning radius of planned trajectory of autonomous driving system in, represents the discrete approximation of the velocity in the x direction at time step t, represents the acceleration in the x direction, x t represents the lateral position coordinate of the vehicle at time step t, x t-1 Represents the lateral position coordinates of the vehicle at time step t-1. The curvature is constrained to not exceed the physical limit of the vehicle's turning and yaw. Derived to the curvature physical constraint in, Represents the trajectory curvature value calculated at time step t.

[0041] Using the vehicle dynamics model as a reference, the generated trajectory satisfies the kinematic equations

[0042]

[0043] Among them, v is the vehicle speed, θ is the heading angle, θ is the front wheel steering angle, and L is the wheelbase; the dynamic loss is defined as the difference between the generated trajectory and the model prediction, and the dynamic loss is Where, represents the trajectory position vector at time step t generated by the diffusion model or prediction model, represents the theoretical position vector at time step t obtained from the vehicle dynamics model simulation.

[0044] Through the implementation of step 103, the introduction of the physical constraint diffusion model can effectively compensate for the problem that traditional prediction models do not adequately consider vehicle dynamics constraints. By incorporating physical constraint losses into the trajectory generation process, the physical constraint diffusion model can ensure that the multimodal future trajectory is closer to actual traffic behavior in terms of temporal coherence and motion rationality, thereby avoiding the generation of unrealistic multimodal future trajectories and improving the feasibility and safety of trajectory prediction.

[0045] Step 104: perform risk assessment on the multimodal future trajectory to obtain a prediction result of autonomous driving of the target vehicle;

[0046] In some examples, by calculating the probability of collision with surrounding obstacles or other vehicles, and the probability of violation of road structures (such as lane markings and traffic lights) for each multimodal future trajectory, and combining it with the vehicle dynamics model, it is determined whether each trajectory has collision risks, rule violations, or operational infeasibility issues. The safety of multimodal future trajectories can be quantified through risk scoring, probability calculation, and other methods. Autonomous driving prediction results refer to the predicted information about the future behavior of the target vehicle output after evaluating multiple multimodal future trajectories, including the most likely driving path, steering actions, acceleration and deceleration trends, etc., which can be used by the autonomous driving system for path planning and risk avoidance decisions.

[0047] By implementing step 104, the generated future trajectory is evaluated for collision risk, abnormal behavior, etc., and potential dangerous scenarios can be identified in advance. This not only helps to improve the predictive ability of the autonomous driving system, but also provides a decision-making basis for active safety strategies such as early warning, emergency braking, and path adjustment, thereby significantly improving the safety and intelligent response capabilities of autonomous driving.

[0048] In summary, the embodiments of the present application achieve information complementarity and redundancy enhancement by acquiring multimodal sensor data of the target vehicle and fusing the features of different modal data. The time difference and spatial error between different sensors are processed through the spatiotemporal alignment mechanism, which improves the accuracy of the fused features and thus enhances the perception of complex traffic scenes. In the trajectory prediction stage, the physical constraint diffusion model is introduced and integrated into the diffusion model generation process as a loss function. This not only improves the rationality and feasibility of the generated trajectory, but also effectively avoids unrealistic trajectories that are prone to occur in traditional data-driven prediction methods, such as sudden lane changes and sharp turns. Since the diffusion model is used to model the joint distribution of multimodal features, even if there is a modality missing in the input data, the missing information can still be reconstructed through the learned inter-modal relationship, thereby maintaining the prediction performance. This can greatly improve the robustness and generalization ability of the autonomous driving scene prediction method in actual complex scenarios. By performing risk assessment on the generated multimodal future trajectory, the risk situations that the target vehicle may face in the future can be more accurately judged, thereby providing a decision-making basis for the autonomous driving process. In summary, the autonomous driving scenario prediction method provided in the embodiments of the present application can effectively improve the accuracy, rationality and robustness of autonomous driving scenario prediction by integrating multimodal perception and introducing a physical constraint diffusion model for trajectory generation and risk assessment.

[0049] In some embodiments, the aforementioned step 102 may include: interpolating and aligning the multimodal sensor data according to a preset timestamp set to obtain intermediate sensor data; extracting features from the intermediate sensor data to obtain multimodal feature data; and fusing the multimodal feature data through a cross-attention module to obtain target fusion features.

[0050] In some examples, a preset timestamp set is a predefined set of unified time points used to time-align data collected by different sensors, ensuring that all modal data are fused under the same time reference. For example, it can be set to a fixed time interval (such as every 100ms), or a timestamp set that is consistent with the time step of data acquisition by a certain sensor (such as a camera). To address the problem of inconsistent sampling frequencies and time synchronization of different sensors, a preset timestamp set can be used as a benchmark, and their data can be time-synchronized through time interpolation (such as linear interpolation or spline interpolation) to obtain intermediate sensor data. For example, if the radar samples every 200ms and the camera samples every 100ms, the radar data can be interpolated to align with the camera at every 100ms. Intermediate sensor data refers to various sensor data that have been time-aligned and have a unified time reference, laying the foundation for subsequent feature extraction. Multimodal feature data is a structured feature representation extracted from intermediate sensor data, such as 2D visual feature information from images, 3D geometric clustering features from radar, GPS location information, and grid-encoded semantic features extracted from map data. For example, image features can be extracted using a convolutional neural network, while radar geometric features can be extracted using a point cloud network. Multimodal feature data can be input into a cross-attention module, where they are interactively fused through an attention mechanism to generate a unified, high-quality fused feature vector with strong scene representation capabilities. The cross-attention module is a neural network structure used to establish associations between multiple modal features and dynamically adjust the contribution weight of each modality to the fusion result. The cross-attention mechanism in the Transformer can be used to perform weighted fusion on multimodal feature data input.

[0051] For example, the raw data from cameras, radars, and GPS can first be temporally interpolated according to a unified timestamp to generate spatiotemporally synchronized intermediate data, and the features of each modality can be extracted. Subsequently, the cross-attention module dynamically adjusts the fusion weights between modalities based on the environmental context, thereby generating target fusion features that contain multimodal information and are spatiotemporally aligned.

[0052] For example, when there are only two sensors, camera and lidar, the camera data sequence is Sampling frequency f c =30Hz, which means collecting 30 frames of images per second. Indicates that at time point t i The feature vector corresponding to the collected camera image data, N c Indicates the total number of data frames collected by the camera. The lidar data sequence is Sampling frequency f l =30Hz, which means collecting 30 frames of point cloud per second. Indicates that at time point t j The feature vector corresponding to the collected lidar point cloud data, N l Indicates the total number of data frames collected by the lidar. The preset timestamp set is K represents the total number of target alignment time points. For each target timestamp t k ∈T, T represents a set of unified time points for time synchronization of multimodal data such as cameras and radars, and finds the nearest adjacent sensor data frame t j <<t k <<t j+1 , t j <<t k <<t j+1 " represents the target time point t k Located at the actual sampling time point t of the sensor j and t j+1 Between, that is, interpolate between these two sampling points and calculate the difference weight Where t represents the value corresponding to the time variable; the aligned intermediate sensor data is and in, Indicates that at the target timestamp t k Alignment features obtained by temporal interpolation of camera image data, Indicates that at the target timestamp t k Alignment features obtained by temporal interpolation of lidar point cloud data, is the feature of the camera data of the jth frame, is the j-th frame laser radar data feature. The output aligned multimodal feature data is in, Represents a real number matrix with K rows and D columns, where K represents the number of frames after time alignment (i.e., the number of timestamps), and D represents the dimension of the fused feature vector at each moment.

[0053] Through the implementation of the above embodiment, the preset timestamp interpolation method is used to align data of different modalities, which solves the synchronization problem of multiple sensors in time and space and ensures the consistency of input data; and then the cross-attention mechanism is used for feature fusion to achieve deep semantic interaction between different modalities, effectively enhancing the expression ability of fused features for scene elements and improving the overall perception and understanding effect.

[0054] In some embodiments, the aforementioned feature fusion of multimodal feature data through the cross-attention module to obtain target fusion features may include: adjusting the weights of the multimodal feature data according to the environmental condition information and sensor status of the target vehicle to obtain environmental right confirmation feature data; and feature fusion of the environmental right confirmation feature data through the cross-attention module to obtain target fusion features.

[0055] In some examples, environmental condition information refers to the real-time state of the target vehicle's environment, including weather (sunny, rainy, foggy, etc.), lighting (daytime, nighttime, etc.), traffic density, and road type (urban roads, highways, tunnels, etc.). Weather and lighting conditions can be identified using camera images, road types can be identified using map data, and surrounding traffic density can be estimated using radar data. Sensor status represents the current operating status of each sensor, such as whether it is functioning properly, signal-to-noise ratio, field of view obstruction, and data acquisition frequency. System self-tests can be used to determine each sensor's operating status. Radar obstruction can be detected by changes in point cloud density, and image blur can be used to assess camera effectiveness. Based on environmental conditions and sensor status, sensor data features from each modality can be weighted differently, prioritizing higher-quality or more reliable data modalities. For example, in foggy weather, the weight of camera features can be reduced, while the relative weight of radar or GPS can be increased. Environmentally authenticated feature data refers to multimodal feature data after modal weighting has been adjusted. Its fusion process incorporates environmental and sensor confidence information to better reflect the current situation. On the confidence-weighted environmental confirmation feature data, the cross-attention mechanism is used to further explore the correlation between modalities, which can be fused into a unified target fusion feature to enhance scene understanding capabilities. The Transformer architecture can be used to send the weighted modal features into the cross-attention layer and output the fused unified feature vector, which is the target fusion feature.

[0056] Through the implementation of the above embodiments, a weight adjustment strategy for environmental conditions and sensor states is introduced in the fusion stage, so that the physical constraint diffusion model can adaptively adjust the modal weight according to the current scenario, thereby improving the perception sensitivity to key modal information. In particular, when some modal information is incomplete or the signal-to-noise ratio is low, the robustness and reliability of the fusion results can still be guaranteed, effectively improving the scene understanding effect.

[0057] In some embodiments, the aforementioned step 103 may include: compressing the target fusion features through a latent space compression module to obtain low-dimensional latent space features; and inputting the low-dimensional latent space features into a physical constrained diffusion model to generate a multimodal future trajectory.

[0058] In some examples, a latent space compression module is used to encode high-dimensional target fusion features and extract the compressed representation, which is a low-dimensional latent space feature. This can reduce the computational complexity of subsequent models and improve generalization. The latent space compression module is a neural network module that compresses high-dimensional features into a low-dimensional latent representation. It can be a variational autoencoder, a self-attention compression network, or a dimensionality reduction encoder. The latent space compression module aims to retain key features while reducing redundant information to facilitate subsequent model processing. The low-dimensional latent space feature is a more compact and information-rich representation compressed from the high-dimensional target fusion feature. It is located in the "latent space" learned by the model and expresses the deeper nature of the features. The compressed low-dimensional latent space feature is used as the input to the physical constraint diffusion model. In the physical constraint diffusion model, through stepwise denoising sampling and the introduction of physical constraints, it can guide the generation of future trajectory data that conforms to the physical characteristics of the vehicle.

[0059] Through the implementation of the above embodiment, compressing high-dimensional fusion features into a low-dimensional latent space helps to reduce the computational complexity of the diffusion model and improve the generation stability. Through the structured latent space, the diffusion process can more effectively learn the distribution characteristics of trajectory generation, thereby generating more reasonable and coherent future trajectory results, and improving the efficiency and accuracy of trajectory modeling.

[0060] In some embodiments, the aforementioned inputting of low-dimensional latent space features into a physical constraint diffusion model to generate a multimodal future trajectory may include: inputting the aforementioned low-dimensional latent space features into the physical constraint diffusion model, so that the physical constraint diffusion model performs reverse denoising based on noise prediction loss and physical constraint loss to obtain a multimodal future trajectory, wherein the physical constraint loss may include at least one of acceleration constraint loss, curvature constraint loss, and vehicle dynamics model matching loss.

[0061] In some examples, during trajectory generation using a physically constrained diffusion model, not only is a standard noise prediction loss used to restore the true trajectory, but an additional physical constraint loss is also introduced to regularize the generation process, guiding the model to generate trajectories that conform to physical laws. For example, during diffusion model training, the loss function can be set as: L = L_noise + λ*L_physics, where L is the total loss, L_noise is the error between the predicted noise and the true noise, L_physics is the physical constraint loss, including acceleration or curvature errors, and λ is a weighting factor. The noise prediction loss is the difference between the predicted noise and the true noise calculated during the training phase of the physically constrained diffusion model to learn how to recover the original data from the noisy data. This is used to optimize the model's denoising capabilities. The mean squared error loss can be used to measure the difference between the model's predicted noise and the actual added noise. The physical constraint loss is a loss function that measures whether the generated trajectory conforms to physical laws, such as smoothness, feasibility, and vehicle dynamic constraints, during the model's trajectory generation process, to guide the generated results to be more realistic. Acceleration constraint loss is a loss that limits the acceleration change in the generated trajectory. It can prevent unrealistic sudden acceleration or deceleration and improve the smoothness and safety of the trajectory. It can ensure that the acceleration is within the specified threshold by calculating the velocity change rate of each frame of the trajectory. The excess is punished as a loss. Curvature constraint loss is a loss used to limit the spatial curvature of the trajectory. It can avoid abnormal sharp turns or jitters in the trajectory and make the trajectory more natural and smooth. The curvature change can be calculated based on adjacent trajectory points. If it exceeds the normal steering range, it is accumulated as a loss value. Vehicle dynamics model matching loss is a loss that the trajectory should be consistent with the vehicle's dynamics model to ensure that the trajectory is physically executable. The state transfer equation derived based on vehicle dynamics can be used to compare with the generated trajectory, and the error is the loss.

[0062] By implementing the above-mentioned embodiments, the introduction of physical constraint losses including acceleration, curvature, and dynamic model matching can ensure that the generated future trajectory is physically reasonable and complies with behavioral restrictions such as staying in the same lane and normal driving, avoiding unrealistic situations such as sudden acceleration and sharp turns. Combined with noise prediction loss, it can also maintain the stability and accuracy of the generated model, further improving the prediction quality.

[0063] In some embodiments, the latent space compression module is a variational autoencoder, and the compression rate of the variational autoencoder is 1 / 16.

[0064] In some examples, the module used to compress high-dimensional input data into a low-dimensional latent feature representation uses a variational autoencoder (VAE), a neural network structure with probabilistic modeling capabilities that can learn the underlying distribution of data. Compared to conventional autoencoders, VAEs introduce a regularization mechanism to generate a more continuous and stable latent space, which is beneficial for subsequent trajectory generation tasks. After compression by the VAE, the input data is reduced to 1 / 16 of its original dimension. This compression ratio significantly reduces the data dimension and the computational burden while preserving key semantic information. For example, if the original fused feature dimension is 1024, the compressed latent feature dimension is 64, resulting in a compression ratio of 64 / 1024 = 1 / 16.

[0065] Through the implementation of the above embodiment, not only the data dimension is significantly reduced and the computational burden is lowered, but also the potential representation ability of the data is retained, which promotes the diffusion model to learn the essential characteristics of trajectory distribution; the variational autoencoder structure can also improve the generative expression ability of compressed features, providing a more compact and efficient input representation for subsequent trajectory generation.

[0066] In some embodiments, step 104 may include: screening multimodal future trajectories according to preset vehicle dynamics constraints to obtain a set of compliant future trajectories; determining a risk quantification score based on the collision probability and traffic rule violation probability of each compliant future trajectory in the set of compliant future trajectories; and sorting the set of compliant future trajectories according to the risk quantification score to obtain an autonomous driving prediction result.

[0067] In some examples, preset vehicle dynamics constraints refer to a set of constraints pre-set based on the basic laws of vehicle physical performance and driving behavior, such as maximum acceleration / deceleration, maximum steering angle, tire grip limit, etc., which can be used to determine whether the future trajectory is physically feasible; preset vehicle dynamics constraints can be determined based on vehicle manufacturing parameters, road traffic regulations or vehicle dynamics models; for example, the steering angle is set not to exceed 35°, the longitudinal acceleration is set not to exceed 3m / s 2The compliant future trajectory set is the set of trajectories in the multimodal future trajectory that satisfy the aforementioned dynamic constraints. This means that physically infeasible "abnormal trajectories" are eliminated, leaving only executable future paths. Each multimodal future trajectory can be analyzed point by point. If its speed, acceleration, or steering angle exceeds a set threshold, it is eliminated; otherwise, it is included in the "compliant set." The risk quantification score comprehensively evaluates the risk factors faced by each compliant future trajectory in the future and converts them into a unified numerical score, which is used to rank the trajectories by risk level. A weighted scoring model can be used, such as: risk quantification score = α × collision probability + β × violation probability, where α and β are empirically determined weight coefficients. For example, if a trajectory has a collision probability of 0.2 and a violation probability of 0.1, and α = 0.7 and β = 0.3, the risk score is 0.2 × 0.7 + 0.1 × 0.3 = 0.17. The collision probability is the predicted probability that a future trajectory will collide with other traffic participants (such as vehicles or pedestrians) within a certain period of time. It can be calculated based on Bayesian methods, Markov models, or deep neural networks, combined with dynamic targets in the current environment (such as adjacent vehicle trajectories). For example, the longer a trajectory overlaps with the multimodal future trajectory of the preceding vehicle within 3 seconds, the higher the collision probability. The traffic rule violation probability is the probability that a multimodal future trajectory will violate traffic rules (such as crossing a lane, running a red light, or driving against traffic) within a certain period of time. Violations can be detected by comparing the trajectory with lane lines, signal light status, and speed limit information in a high-precision map. For example, if a trajectory exceeds the speed limit or runs a red light, its violation probability will be significantly increased. Each compliant future trajectory is scored based on its potential risk (such as collision probability and probability of violating traffic rules), sorted, and the optimal trajectory is selected as the prediction output of the autonomous driving system.

[0068] Through the implementation of the above embodiment, it is possible to ensure that only executable trajectory plans are retained; combining indicators such as collision probability and traffic violation probability for risk scoring and ranking can not only enhance the practicality of autonomous driving prediction results, but also provide a priority basis for proactive decision-making, thereby significantly improving the safety and intelligence of autonomous driving.

[0069] The calculation process of the physical constrained diffusion model of the embodiment of the present application will be described below:

[0070] 1. Model input

[0071] Historical trace of input data

[0072] Among them, P t =(x t ,y t ,θ t ,v t) contains the position, heading angle, and velocity in the ego-vehicle coordinate system. The multimodal sensor data S = (I, L, M) contains image, point cloud, and map information.

[0073] Vehicle dynamics model:

[0074] Among them, δ is the vehicle’s front road steering angle, l f and l r Respectively represent the front and rear axle wheelbase and L = l f +l r .

[0075] Generate multimodal trajectories for the next T steps Satisfy vehicle dynamics constraints.

[0076] The system calculation process is divided into 5 steps:

[0077] Step 1: Multimodal feature extraction and fusion

[0078] Sensor Data Encoding: Image Features Point cloud features Map Features

[0079] Cross-modal attention fusion Where d is the fusion feature dimension.

[0080] Step 2: Latent Space Diffusion Modeling

[0081] Encoding compression: where d z is the latent space dimension, usually d z <<<D, same below.

[0082] The forward diffusion process gradually adds noise in the latent space:

[0083] Inverse denoising network, train the denoising network with vehicle dynamics constraints: ∈ θ (z t ,t,C)=U-Net(z t ,t,C).

[0084] Step 3: Mathematical Embedding of Dynamic Constraints

[0085] Generate trajectory decoding, decode the original space trajectory from the latent variable z0: Y gen =Decoder ψ (z0)={p1,……,p T}

[0086] Dynamic loss calculation. Based on the vehicle dynamics model, from the previous state Pt-1 Estimate the next state to generate a reference trajectory where δ t and v t Vehicle steering angles and velocities generated for the diffusion model.

[0087] Trajectory matching loss

[0088] Joint optimization objective, using a total loss function to fuse noise prediction and vehicle dynamics model constraints. This includes using a segmented training strategy:

[0089] 1. Pre-training phase: The goal is to initially learn the data distribution. To capture the diversity of autonomous driving scenarios, set the hyperparameters λ1 = λ2 = λ3 = 0. The termination condition is the diffusion loss. Convergence, as in loss value fluctuations < 1%.

[0090] 2. Joint optimization phase: The goal is to gradually inject dynamic constraints to balance the quality of the generated data with the physical compliance. The initial value of the hyperparameter adjustment is set to λ1 = 0.1, λ2 = 0.05, and λ3 = 0.05. Every 10 epochs are trained. Epoch is the process in which the neural network traverses the entire training data set once. i ←λ i +0.1 increments until λ1=1,λ2=0.5,λ3=0.5。The termination condition is the total loss It has stabilized and the violation rate is less than 0.5%.

[0091] Contains the use of gradient backpropagation and optimization:

[0092] 1. Gradient fusion, diffusion loss gradient Gradient with physical constraints Weighted sum, total gradient

[0093] 2. Gradient clipping. To prevent gradient explosion, limit the gradient norm

[0094] 3. Optimizer configuration. Use AdamW optimizer, initial learning rate η = 10 -4 , weight decay 10-3.

[0095] After the total loss function integrates noise prediction and dynamic constraints, high-quality generation and high safety of autonomous driving scenario predictions can be achieved through phased training strategies, gradient optimization design and dynamic hyperparameter adjustment.

[0096] Step 4: Denoising Diffusion Implicit Models (DDIM) Accelerated Sampling and Trajectory Correction

[0097] Deterministic sampling, using DDIM to shorten the sampling steps to S<<T, so that Established.

[0098] Feasibility correction of the vehicle dynamics model, the generated trajectory Y gen Post-processing and projection into feasible space Use QP quadratic programming or PID control to adjust the trajectory in real time.

[0099] Step 5: Multimodal output and risk assessment

[0100] Multimodal sampling, sampling K times from the diffusion model to generate a diverse set of trajectories

[0101] Vehicle dynamics compliance screening, eliminating trajectories that do not conform to the vehicle model vaild ={Y k |KinematicViolatio(Y k )<τ}.

[0102] Risk sensitivity ranking, sorted by collision probability and traffic regulations probability

[0103] The calculation process of the joint optimization objective is as follows:

[0104] The joint optimization objective ensures that the generated future scenarios are both consistent with the data distribution and physically feasible by weighted summing the noise prediction loss and physical constraint loss of the diffusion model. The total loss function is formally expressed as:

[0105]

[0106] in, represents the noise prediction loss of the diffusion model, represents the vehicle dynamics model matching loss, represents the vehicle acceleration constraint loss, represents the curvature constraint loss, λ1, λ2, and λ3 represent hyperparameters used to balance the contribution of each loss term.

[0107] The diffusion model is trained by predicting noise ∈, and the loss function is the diffusion loss in, is the latent variable after noise addition, ∈ θ It is a denoising network, and the input conditions are time t and multimodal features C.

[0108] Reference trajectory Y generated by the vehicle dynamics model ref and generate trajectory Y gen Mean square error, dynamics model matching loss

[0109] Among them, p t =(x t, y t ,θ t ,v t ) contains the vehicle position, heading angle and speed, represents the trajectory position vector at time step t generated by the diffusion model or prediction model, represents the theoretical position vector at time step t obtained from the vehicle dynamics model simulation.

[0110] Penalize acceleration that exceeds the physical limits of the vehicle in the generated trajectory, acceleration constraint loss Among them, a t is the acceleration, which is calculated by the second-order difference of the trajectory:

[0111] Limit the curvature of the generated trajectory to not exceed the maximum value corresponding to the minimum turning radius of the vehicle, and the curvature constraint loss

[0112] Hyperparameter selection and balancing strategies are divided into initial training, mid- and late-stage training, and adaptive adjustment strategies.

[0113] Initial training is dominated by diffusion loss, with λ1, λ2, and λ3 set to small values ​​(e.g., 0.1) to prioritize learning the data distribution pattern. Mid- to late-stage training is a phase of physical constraint enforcement, with gradually increasing hyperparameters (e.g., λ1 = 1.0, λ2 = 0.5, λ3 = 0.5) to enforce that the generated trajectories conform to physical laws. The adaptive adjustment strategy is dynamically adjusted based on the violation rate on the validation set.

[0114] Gradient calculation and back propagation, all loss terms are designed as differentiable functions to ensure the feasibility of end-to-end training. The diffusion loss gradient is

[0115]

[0116] The loss gradient of the vehicle dynamics model is

[0117]

[0118] Calculate the acceleration and curvature loss gradients frame by frame using the chain rule

[0119]

[0120] The joint optimization objective achieves the following core advantages by balancing the noise prediction loss and the physical constraint loss:

[0121] 1. The trade-off between generation quality and safety: ensuring that trajectories conform to vehicle dynamics while maintaining scene diversity;

[0122] 2. Scalability: supports flexible addition of other constraints (such as traffic rules and energy consumption optimization).

[0123] 3. Training stability: Adaptive hyperparameter adjustment strategy avoids divergence in the optimization process.

[0124] 4. This design provides an innovative framework that combines data-driven and physical guidance for autonomous driving scenario prediction.

[0125] Furthermore, as an implementation of the aforementioned method embodiment, the present application also provides an autonomous driving scene prediction device for implementing the aforementioned method embodiment. This device embodiment corresponds to the aforementioned method embodiment. For ease of reading, this autonomous driving scene prediction device embodiment will no longer describe the details of the aforementioned method embodiment one by one, but it should be clear that the device in the embodiment of the present application can correspond to and implement all the contents of the aforementioned method embodiment. Figure 2 As shown, the autonomous driving scene prediction device 20 includes: a data acquisition unit 201, a feature fusion unit 202, a trajectory prediction unit 203 and a risk assessment unit 204, wherein the data acquisition unit 201 is used to acquire multimodal sensor data of the target vehicle; the feature fusion unit 202 is used to perform feature fusion on the multimodal sensor data to obtain a target fusion feature aligned in time and space; the trajectory prediction unit 203 is used to diffuse and generate the target fusion feature based on a physical constraint diffusion model to obtain a multimodal future trajectory, wherein the physical constraint diffusion model is a model obtained by optimizing the generation process of the diffusion model through physical constraint loss; the risk assessment unit 204 is used to perform risk assessment on the multimodal future trajectory to obtain an autonomous driving prediction result of the target vehicle.

[0126] In some embodiments, the feature fusion unit 202 is also used to interpolate and align the multimodal sensor data according to a preset timestamp set to obtain intermediate sensor data; perform feature extraction on the intermediate sensor data to obtain multimodal feature data; and perform feature fusion on the multimodal feature data through a cross-attention module to obtain target fusion features.

[0127] In some embodiments, the feature fusion unit 202 is also used to adjust the weight of the multimodal feature data according to the environmental condition information and sensor status of the target vehicle to obtain environmental confirmation feature data; and perform feature fusion on the environmental confirmation feature data through the cross-attention module to obtain the target fusion feature.

[0128] In some embodiments, the trajectory prediction unit 203 is further configured to compress the target fusion features through a latent space compression module to obtain low-dimensional latent space features; and input the low-dimensional latent space features into the physical constraint diffusion model to generate a multimodal future trajectory.

[0129] In some embodiments, the trajectory prediction unit 203 is further used to input the low-dimensional latent space features into the physical constraint diffusion model, so that the physical constraint diffusion model performs reverse denoising based on the noise prediction loss and the physical constraint loss to obtain a multimodal future trajectory, wherein the physical constraint loss includes at least one of the acceleration constraint loss, the curvature constraint loss, and the vehicle dynamics model matching loss.

[0130] In some embodiments, the latent space compression module is a variational autoencoder, and the compression rate of the variational autoencoder is 1 / 16.

[0131] In some embodiments, the risk assessment unit 204 is further configured to screen multimodal future trajectories according to preset vehicle dynamics constraints to obtain a set of compliant future trajectories; determine a risk quantification score based on the collision probability and traffic rule violation probability of each compliant future trajectory in the set of compliant future trajectories; and sort the set of compliant future trajectories according to the risk quantification score to obtain an autonomous driving prediction result.

[0132] The present application also provides a computer-readable storage medium, which stores computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute any step of the autonomous driving scene prediction method provided in the present application.

[0133] In some embodiments, the computer-readable storage medium may be a random access memory (RAM), a read-only memory (ROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); or it may be various devices including one or any combination of the above memories.

[0134] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0135] In some embodiments, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (for example, files storing one or more modules, subroutines, or code portions).

[0136] In some embodiments, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0137] like Figure 3 As shown, the present application also provides an electronic device 30, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, any step of the above-mentioned autonomous driving scene prediction method is implemented.

[0138] The present application also provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform any step of the autonomous driving scenario prediction method described above.

[0139] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for predicting an autonomous driving scene, characterized in that: include: Acquire multimodal sensor data of the target vehicle; Performing feature fusion on the multimodal sensor data to obtain a target fusion feature aligned in time and space; Based on a physical constraint diffusion model, the target fusion feature is diffusely generated to obtain a multimodal future trajectory, wherein the physical constraint diffusion model is a model obtained by optimizing the generation process of the diffusion model through a physical constraint loss; A risk assessment is performed on the multimodal future trajectory to obtain an autonomous driving prediction result of the target vehicle.

2. The autonomous driving scene prediction method according to claim 1, characterized in that: The step of performing feature fusion on the multimodal sensor data to obtain a spatiotemporally aligned target fusion feature includes: interpolating and aligning the multimodal sensor data according to a preset timestamp set to obtain intermediate sensor data; Performing feature extraction on the intermediate sensor data to obtain multimodal feature data; The multimodal feature data is subjected to feature fusion through a cross attention module to obtain the target fusion feature.

3. The autonomous driving scene prediction method according to claim 2, characterized in that: The step of performing feature fusion on the multimodal feature data through a cross attention module to obtain the target fusion feature includes: According to the environmental condition information and sensor status of the target vehicle, the multimodal feature data is weighted to obtain environmental right confirmation feature data; The cross-attention module is used to perform feature fusion on the environmental right confirmation feature data to obtain the target fusion feature.

4. The autonomous driving scene prediction method according to claim 1, characterized in that: The physical constraint diffusion model is used to diffuse and generate the target fusion features to obtain a multimodal future trajectory, including: Compressing the target fusion features through a latent space compression module to obtain low-dimensional latent space features; The low-dimensional latent space features are input into the physical constrained diffusion model to generate the multimodal future trajectory.

5. The automatic driving scene prediction method according to claim 4, characterized in that: Inputting the low-dimensional latent space features into the physical constrained diffusion model to generate the multimodal future trajectory includes: The low-dimensional latent space features are input into a physical constraint diffusion model so that the physical constraint diffusion model performs reverse denoising based on a noise prediction loss and a physical constraint loss to obtain the multimodal future trajectory, wherein the physical constraint loss includes at least one of an acceleration constraint loss, a curvature constraint loss, and a vehicle dynamics model matching loss.

6. The method for predicting an autonomous driving scenario according to claim 4, wherein: The latent space compression module is a variational autoencoder, and the compression rate of the variational autoencoder is 1 / 16.

7. The method for predicting an autonomous driving scene according to any one of claims 1 to 6, wherein: Performing a risk assessment on the multimodal future trajectory to obtain an autonomous driving prediction result for the target vehicle includes: According to preset vehicle dynamics constraints, the multimodal future trajectories are screened to obtain a set of compliant future trajectories; determining a risk quantification score based on a collision probability and a traffic rule violation probability of each compliant future trajectory in the set of compliant future trajectories; The set of compliant future trajectories is sorted according to the risk quantification score to obtain the autonomous driving prediction result.

8. An autonomous driving scene prediction device, characterized in that: include: A data acquisition unit, used to acquire multimodal sensor data of a target vehicle; A feature fusion unit, configured to perform feature fusion on the multimodal sensor data to obtain a target fusion feature aligned in time and space; a trajectory prediction unit, configured to perform diffusion generation on the target fusion features based on a physical constraint diffusion model to obtain a multimodal future trajectory, wherein the physical constraint diffusion model is a model obtained by optimizing the generation process of the diffusion model through a physical constraint loss; A risk assessment unit is used to perform risk assessment on the multimodal future trajectory to obtain an autonomous driving prediction result of the target vehicle.

9. An electronic device comprising: A memory and a processor, characterized in that the processor is used to implement the steps of the autonomous driving scene prediction method according to any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the autonomous driving scene prediction method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Track generation method and system for autonomous vehicle and storage medium

    CN120840665A