An automatic driving safety-critical multi-view video automatic generation method and device

By acquiring autonomous driving scenario information and using a diffusion model and the Nuscenes dataset to generate multi-view safety-critical videos, the problem of low generation efficiency and poor quality in existing technologies is solved, and efficient and accurate multi-view video generation suitable for vision-based end-to-end autonomous driving systems is achieved.

CN120954236BActive Publication Date: 2026-01-20ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511478440.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-20
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing technologies lack efficient and accurate automated methods for generating multi-view videos, making it impossible to generate safety-critical scenarios suitable for vision-based end-to-end autonomous driving systems. Furthermore, the videos generated by existing methods are of low quality and have limited controllability.

Method used

By acquiring information about autonomous vehicles, non-autonomous vehicles, and maps in the initial autonomous driving scenario, a diffusion model is used to select safety-critical vehicles and simulate trajectories. Combined with the Nuscenes dataset for annotation and multi-view video generation, the automated generation of safety-critical videos is achieved.

Benefits of technology

It achieves efficient generation of vehicle realism and diversity, produces high-quality multi-view safety-critical videos, significantly improves data utilization and system development efficiency, and reduces manual intervention and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954236B_ABST
    Figure CN120954236B_ABST
Patent Text Reader

Abstract

The application discloses an automatic driving safety key multi-view video automatic generation method and device, and relates to the technical field of automatic driving. The method comprises the following steps: based on map information, safety key vehicle information is obtained by selecting safety key vehicles according to self-vehicle information and non-self-vehicle information; based on a reasoning time guide function, an overall guide value is obtained by performing calculation according to the map information, the self-vehicle information, the non-self-vehicle information and the safety key vehicle information; based on the overall guide value, simulation of a safety trajectory is performed by using a pre-trained diffusion model according to the map information, the self-vehicle information, the non-self-vehicle information and the safety key vehicle information, and simulation trajectory information is obtained; the simulation trajectory information is labeled, a video is generated by using a preset multi-view video generation model, and multi-view safety key videos are obtained. The application is a kind of efficient and accurate multi-view video automatic generation method for automatic driving safety key.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic driving, in particular to an automatic driving safety-critical multi-view video automatic generation method and device. BACKGROUND

[0002] Vision-based end-to-end autonomous driving systems are increasingly gaining attention and gradually being applied in real-world scenarios due to their ability to directly map visual inputs to driving decisions. However, ensuring the safety of these systems in various complex scenarios remains a significant challenge. To address this issue, large-scale, diverse driving datasets, particularly those involving safety-critical situations, are required for testing. Collecting such data in the real world is both expensive and inherently dangerous. As a promising alternative, the synthetic generation of safety-critical driving scenarios offers a scalable, low-risk, and cost-effective solution.

[0003] Common safety-critical autonomous driving scenario generation primarily involves safety traffic flow generation methods, which generate the trajectories of all vehicles within a scene, predict dangerous vehicles approaching with dangerous trajectories, and generate the trajectory data of the ego vehicle. Safety traffic flow generation mainly utilizes diffusion models to simulate traffic trajectories, typically modeling vehicle trajectories using diffusion models and training diffusion models on large driving datasets to generate results close to real trajectories. In the denoising process, adversarial loss is used to guide the denoising process to achieve controllable trajectory generation and obtain dangerous collision trajectories.

[0004] Since the input of vision-based end-to-end autonomous driving systems is image data captured by cameras, safety traffic flow generation cannot be applied in the current popular vision-based end-to-end autonomous driving system testing. Video generation models based on diffusion models provide the possibility of generating safety-critical scenarios at the image level. Video generation models based on diffusion models are pre-trained on massive videos, modeling the ability of video generation. To further understand autonomous driving scenarios, further fine-tuning on autonomous driving datasets is required, and current diffusion models contain controllable signals such as text and image signals from the pre-trained model itself. All vehicle motion information and map information in the autonomous driving scenario are used as control signals to guide video generation. However, the common multi-view video generation framework relies on the trajectories in existing conventional autonomous driving datasets for video generation, and the generated multi-view videos lack safety-criticality. Manually operating vehicle motion information for video generation control cannot effectively simulate vehicle motion, and generating safety-critical scenarios for different scenarios requires independent design, resulting in low generation efficiency.

[0005] Currently, most methods for safety-critical autonomous driving data generation focus on using diffusion models to generate realistic and controllable adversarial trajectories. These trajectories are mainly used to evaluate and improve the performance of the planning module of autonomous driving systems. The trajectories generated based on the above methods are non-visual trajectories, which are incompatible with end-to-end autonomous driving systems that require visual input. Although some methods use simulators to generate safety-critical driving videos, their effectiveness is limited by the gap between simulation and reality. In addition, although some deep video generation models have also been explored for generating real-world driving accident videos, the generated videos are generally of low quality and limited to single-view output, and the controllability is limited. With the extensive exploration of multi-view generation models for autonomous driving, it provides a possibility for generating multi-view safety-critical videos, but current multi-view generation models usually rely on fixed open-source datasets such as vehicle data recorded in Nuscenes for automated generation, and the generated scenes are usually non-safety-critical data, and manual modification of vehicle trajectories often brings greater vehicle trajectory inauthenticity and is extremely inefficient.

[0006] In the prior art, there is a lack of an efficient and accurate multi-view video automated generation method for safety-critical autonomous driving. SUMMARY

[0007] To solve the technical problems of safety-critical trajectory generation methods and single-view safety-critical video generation that do not adapt to the input of visual end-to-end systems in the prior art, the embodiments of the present application provide an autonomous driving safety-critical multi-view video automated generation method and device. The technical solution is as follows:

[0008] On the one hand, an autonomous driving safety-critical multi-view video automated generation method is provided, which is implemented by a multi-view video automated generation device, and the method comprises:

[0009] Obtain self-vehicle information, non-self-vehicle information, and map information of an initial autonomous driving scene;

[0010] Based on the map information, safety-critical vehicle selection is performed according to the self-vehicle information and the non-self-vehicle information to obtain safety-critical vehicle information;

[0011] Based on the reasoning time guide function, the map information, the self-vehicle information, the non-self-vehicle information, and the safety-critical vehicle information are calculated to obtain an overall guide value;

[0012] Based on the overall guide value, the map information, the self-vehicle information, the non-self-vehicle information, and the safety-critical vehicle information are used to perform safety trajectory simulation using a pre-trained diffusion model to obtain simulation trajectory information;

[0013] According to the annotated trajectory information, a preset multi-view video generation model is used for video generation to obtain multi-view safety-critical video.

[0014] According to the annotated trajectory information, a preset multi-view video generation model is used for video generation to obtain multi-view safety-critical video.

[0015] In another aspect, an automatic driving safety-critical multi-view video automatic generation device is provided, which is applied to the automatic driving safety-critical multi-view video automatic generation method, and the device comprises:

[0016] An information acquisition module is configured to acquire ego vehicle information, non-ego vehicle information and map information of an initial scene of automatic driving.

[0017] A safety-critical vehicle selection module is configured to select safety-critical vehicles based on the map information, the ego vehicle information and the non-ego vehicle information to obtain safety-critical vehicle information.

[0018] A total guidance value calculation module is configured to calculate a total guidance value based on a guidance function at an inference time and the map information, the ego vehicle information, the non-ego vehicle information and the safety-critical vehicle information.

[0019] A trajectory simulation module is configured to simulate safety trajectories based on the total guidance value and the map information, the ego vehicle information, the non-ego vehicle information and the safety-critical vehicle information using a pre-trained diffusion model to obtain simulated trajectory information.

[0020] An information annotation module is configured to annotate the simulated trajectory information based on a data form of a Nuscenes dataset to obtain annotated trajectory information.

[0021] A video generation module is configured to generate video based on the annotated trajectory information using a preset multi-view video generation model to obtain multi-view safety-critical video.

[0022] In another aspect, a multi-view video automatic generation device is provided, which comprises a processor and a memory having computer readable instructions stored thereon, wherein the computer readable instructions are executed by the processor to implement any one of the above automatic driving safety-critical multi-view video automatic generation methods.

[0023] In another aspect, a computer readable storage medium is provided, which stores at least one instruction, wherein the at least one instruction is loaded and executed by a processor to implement any one of the above automatic driving safety-critical multi-view video automatic generation methods.

[0024] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:

[0025] The application provides an automatic driving safety-critical multi-view video automatic generation method, calculates a danger score of each non-self vehicle according to vehicle information and map information, selects a vehicle with the highest danger score as a safety-critical vehicle, realizes vehicle authenticity and diversity, simulates a safety trajectory by using a controllable trajectory diffusion model, solves the problem of low efficiency of manual adjustment, can automatically generate a large amount of high-quality videos, greatly reduces manual intervention and cost, and automatically generates multi-view safety-critical videos, perfectly meets the input requirements of a visual end-to-end automatic driving system, and significantly improves data utilization and system development efficiency. The application is an efficient and accurate multi-view video automatic generation method for automatic driving safety-critical. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0027] Figure 1 It is a flow chart of an automatic driving safety-critical multi-view video automatic generation method provided by the embodiment of the present application.

[0028] Figure 2 It is a block diagram of an automatic driving safety-critical multi-view video automatic generation device provided by the embodiment of the present application.

[0029] Figure 3 It is a structural schematic diagram of a multi-view video automatic generation equipment provided by the embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the present application will be described below with reference to the drawings.

[0031] In the embodiments of the present application, the words such as "example", "for example" are used to represent as an example, illustration or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be either one of the two.

[0032] In the embodiments of the present application, "image" and "picture" can be used interchangeably, and it should be pointed out that their meanings are consistent when their differences are not emphasized.

[0033] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1, and their meanings are consistent when their differences are not emphasized.

[0034] To make the technical problems, technical solutions and advantages to be solved by the present application clearer, the following will be described in detail in conjunction with the drawings and specific embodiments.

[0035] The embodiments of the present application provide an automatic driving safety key multi-view video automatic generation method, which can be realized by a multi-view video automatic generation device, which can be a terminal or a server. Figure 1 As shown in the automatic driving safety key multi-view video automatic generation method flow chart, the processing flow of the method can include the following steps:

[0036] S1, obtaining self-vehicle information, non-self-vehicle information and map information of an initial scene of automatic driving.

[0037] In a feasible implementation manner, the current state information of the self-vehicle (target vehicle) obtained by the present application includes but is not limited to position, speed and heading angle. The current state information of all non-self-vehicles in the scene obtained includes but is not limited to position, speed and heading angle. The high-precision map information obtained includes but is not limited to lane lines, intersection topologies, etc.

[0038] Some application example data in this part can be obtained from the open source automatic driving data set Nuscenes, or vehicle map features can be extracted using the map on OpenStreetMap (OSM).

[0039] S2, based on the map information, performing safety key vehicle selection according to the self-vehicle information and the non-self-vehicle information to obtain safety key vehicle information.

[0040] Optionally, based on the map information, performing safety key vehicle selection according to the self-vehicle information and the non-self-vehicle information to obtain safety key vehicle information, including:

[0041] According to the self-vehicle information and the non-self-vehicle information, the Euclidean distance set is calculated;

[0042] According to the ego vehicle information and the non-ego vehicle information, a relative speed absolute value set is obtained by calculation;

[0043] Based on a preset conflict value, a map conflict factor set is obtained by conflict evaluation according to the map information, the ego vehicle information and the non-ego vehicle information;

[0044] According to the Euclidean distance set, the relative speed absolute value set and the map conflict factor set, a danger degree score set is obtained by calculation;

[0045] The non-ego vehicle information corresponding to the maximum value in the danger degree score set is selected as the safety critical vehicle information.

[0046] In an available implementation, three vehicle map features are used in the present application for screening safety critical vehicle information, including Euclidean distance, relative speed absolute value and map conflict factor.

[0047] The Euclidean distance (D) is calculated between the ego vehicle and the non-ego vehicle. The relative speed absolute value (V) is calculated as the difference between the ego vehicle speed vector and the non-ego vehicle speed vector. The map conflict factor (C) is evaluated based on the current position, heading of the non-ego vehicle and high-definition map information.

[0048] The relative speed absolute value (V) is calculated as the difference between the ego vehicle speed vector and the non-ego vehicle speed vector. The map conflict factor (C) is evaluated based on the current position, heading of the non-ego vehicle and high-definition map information.

[0049] The map conflict factor (C) is evaluated based on the current position, heading of the non-ego vehicle and high-definition map information. For each non-ego vehicle in the scene, a danger score (SDS) is calculated according to the above-mentioned feature parameters, and the calculation formula is as follows:

[0050] For each non-ego vehicle in the scene, a danger score (SDS) is calculated according to the above-mentioned feature parameters, and the calculation formula is as follows:

[0051] (1);

[0052] The danger degree score (SDS) describes the possibility of collision with the ego vehicle in the current scene, and is used as the object of safety trajectory simulation in the following collision simulation, which can greatly simulate the probability of successful collision and improve the generation success and efficiency.

[0053] Optionally, based on a preset conflict value, a map conflict factor set is obtained by conflict evaluation according to the map information, the ego vehicle information and the non-ego vehicle information, including:

[0054] ​​​​​Based on map information and vehicle information, perform same-direction following conflict verification on non-vehicle information to obtain the first set of verification results.

[0055] Based on map information and vehicle information, cross-road trajectory conflict verification is performed on non-vehicle information to obtain a second set of verification results.

[0056] Based on map information and vehicle information, adjacent lane cut-in conflict verification is performed on non-vehicle information to obtain a third set of verification results.

[0057] Based on preset conflict values, conflict factors are assigned to non-autonomous vehicles according to the first verification result set, the second verification result set, and the third verification result set to obtain a map conflict factor set.

[0058] The map conflict factor set includes high conflict factor, medium conflict factor, and low conflict factor types.

[0059] In one feasible implementation, the map conflict factor includes a high conflict factor, a medium conflict factor, and a low conflict factor; wherein, the high conflict factor is assigned a value of 3; the medium conflict factor is assigned a value of 2; and the low conflict factor is assigned a value of 1.

[0060] For non-autonomous vehicle information, perform same-direction following conflict verification. When the first verification result satisfies that the non-autonomous vehicle is located in the current lane of the autonomous vehicle, the distance between them is less than 30 meters, and the difference in heading angle is less than 15°, the conflict factor of the non-autonomous vehicle is high conflict factor.

[0061] When the first verification result satisfies that the non-autonomous vehicle is located in an adjacent lane, the difference in heading angle is less than 30°, and the longitudinal distance is less than 40 meters, the conflict factor of the non-autonomous vehicle is a medium conflict factor.

[0062] When the first verification result satisfies that the non-autonomous vehicle is located in the oncoming lane and there are physical separation structures such as a central median strip and guardrail, the conflict factor of the non-autonomous vehicle is low.

[0063] Cross-trajectory conflict verification is performed on non-autonomous vehicle information. When the second verification result satisfies that the non-autonomous vehicle is in the cross-trajectory area and its predicted path intersects with the autonomous vehicle's path within the next 3 seconds, the conflict factor of the non-autonomous vehicle is a high conflict factor.

[0064] When the second verification result satisfies that the non-autonomous vehicle is in the lane merging / merging area and the closest distance between the path trajectory and the autonomous vehicle's path within the next 5 seconds is less than 10 meters, the conflict factor of the non-autonomous vehicle is a medium conflict factor.

[0065] When the second verification result satisfies that the Euclidean distance between the non-autonomous vehicle and the autonomous vehicle is greater than 50 meters, and the minimum distance between the paths of the two vehicles in the next 5 seconds is greater than 20 meters, the conflict factor of the non-autonomous vehicle is low.

[0066] When the third check result meets that the non-self vehicle is located in the adjacent lane and has a lane-changing trend with a lateral speed greater than 0.5 m / s, and the distance from the non-self vehicle to the ego vehicle is less than 20 meters, the conflict factor of the non-self vehicle is a high conflict factor;

[0067] When the third check result meets that the non-self vehicle is not currently on the path of the ego vehicle, but has a trend of approaching the path of the ego vehicle, and the distance is within the range of 30-50 meters, the conflict factor of the non-self vehicle is a medium conflict factor;

[0068] When the third check result meets that the non-self vehicle is stationary and located in a non-driving area (such as the edge of the road, a parking area, the outside of a sidewalk, etc.) or a safety area marked in the map, the conflict factor of the non-self vehicle is a low conflict factor.

[0069] In the above three checks, a determination result and a corresponding conflict factor assignment can be obtained as long as any one of the check determination conditions is met, and since the physical scenarios of the above check determination conditions are non-overlapping and the determination logic is mutually exclusive, the check result of the non-self vehicle information meeting the determination condition is unique.

[0070] S3, based on the reasoning-time guidance function, calculating the overall guidance value based on the map information, the ego vehicle information, the non-self vehicle information, and the safety-critical vehicle information.

[0071] Optionally, based on the reasoning-time guidance function, the overall guidance value is calculated based on the map information, the ego vehicle information, the non-self vehicle information, and the safety-critical vehicle information, including:

[0072] Based on the vehicle influence guidance function, the first guidance value is calculated based on the map information, the ego vehicle information, and the safety-critical vehicle information;

[0073] Based on the obstacle avoidance safety guidance function, the second guidance value is calculated based on the map information, the ego vehicle information, and the non-self vehicle information;

[0074] Based on the road adhesion guidance function, the third guidance value is calculated based on the map information and the ego vehicle information;

[0075] The first guidance value, the second guidance value, and the third guidance value are weighted and calculated to obtain the overall guidance value.

[0076] In a feasible implementation, the present application designs the following reasoning-time guidance function for controllable trajectory diffusion, which can perform safe trajectory simulation, and uses the following reasoning-time guidance function combination to guide the denoising process to achieve collision trajectory generation.

[0077] wherein the calculation of the vehicle influence guidance function is as follows:

[0078] (2);

[0079] where, is the distance between ego and safety critical vehicle in future K time steps; is the minimum safety distance of ego vehicle; is an indicator function, which returns 1 when the condition in the bracket is true, and 0 otherwise;

[0080] The obstacle avoidance safety guiding function is calculated as follows in equation (3):

[0081] (3);

[0082] where, is the distance between ego vehicle x with speed and non-ego vehicle y with speed in future K time steps; is the minimum safety distance between ego vehicle x and non-ego vehicle y; is the safety mechanism mask; is the stationary threshold of ego vehicle x relative to non-ego vehicle y;

[0083] The road adhesion guiding function is calculated as follows in equation (4):

[0084] (4);

[0085] where, is the sampled point position; is the off-road region; is the on-road point position; is the on-road region; is the Euclidean distance calculation; is the diagonal length of the vehicle bounding box of ego vehicle.

[0086] In one possible implementation, the vehicle influence guiding function is proposed to encourage the opponent vehicle to collide with the ego vehicle. When the distance between the two vehicles is less than the minimum safety distance , an influence penalty is generated. To ensure that only the opponent vehicle exhibits "aggressive" behavior, the trajectory of the ego vehicle is fixed when the gradient is calculated. Once the two vehicles actually collide, the loss stops calculating immediately to avoid unnatural sticking or abnormal behavior after the collision.

[0087] The obstacle avoidance safety guiding function is used to prevent collisions between non-ego vehicles in the scene other than the ego vehicle and the opponent vehicle. If the distance between the vehicles and in future K time steps less than their minimum safe distance , and moving (whose speed exceeded the static threshold ) will incur a safety penalty. In calculating this penalty, the motion of the vehicle is treated as fixed, meaning that when a moving vehicle approaches a stationary vehicle, the moving vehicle will proactively adjust its path to avoid collision, which is more consistent with real-world driving behavior. The mask is used to specify which vehicles need to enable this safety mechanism. In the “collision simulation” phase, this mechanism applies to all vehicles except the host and opponent vehicles; in the “avoidance simulation” phase, it applies to all vehicles.

[0088] The road attachment loss aims to ensure that the vehicle trajectory is always within the drivable area. When the vehicle is moving (whose speed exceeded the static threshold ), and there are sampling points located outside the road ( ) within its bounding box, an attachment penalty will be incurred. The size of this penalty is inversely proportional to the Euclidean distance from the point to the nearest road point ( ), and is normalized by the diagonal length of the vehicle bounding box . The gradient of this loss will guide the vehicle's trajectory so that the part outside the road is “pulled back” to the nearest drivable area, ensuring the rationality of the vehicle's motion.

[0089] Finally, the guiding function is constructed when reasoning, as follows (5):

[0090] (5);

[0091] where , , are the corresponding weighting coefficients, respectively.

[0092] Under the overall guiding value optimization, the safety-critical gradient is introduced in the denoising process and the authenticity of the scene is ensured, generating fixed time-sequential vehicle trajectory signals.

[0093] S4, based on the overall guiding value, according to the map information, self-vehicle information, non-self-vehicle information and safety-critical vehicle information, using a pre-trained diffusion model to simulate a safe trajectory, obtaining simulated trajectory information.

[0094] In a feasible implementation, the application modifies the Diffusion model in the controllable text generation (CTG) work to implement safe trajectory simulation, and the CTG is a trajectory generation Diffusion model pre-trained on trajectory data. The input is the initial vehicle state and map information. The subsequent motion state of the vehicle is obtained by denoising from Gaussian noise, and the output is the motion information of all vehicles in the subsequent period of time. By calculating the overall guide value of the output trajectory of each step of the Diffusion model in the Diffusion model denoising process, the gradient of the overall guide value is used to optimize the vehicle trajectory of each step, so as to realize the generation of safe key trajectories.

[0095] S5, based on the data form of the Nuscenes dataset, the simulated trajectory information is labeled to obtain labeled trajectory information.

[0096] Optionally, based on the data form of the Nuscenes dataset, the simulated trajectory information is labeled to obtain labeled trajectory information, comprising:

[0097] Based on the data form of the Nuscenes dataset, the frequency alignment is performed according to the translation vector of the simulated trajectory information to obtain first trajectory information;

[0098] Based on the data form of the Nuscenes dataset, the three-dimensional pose update is performed according to the two-dimensional trajectory of the first trajectory information to obtain second trajectory information;

[0099] Based on the data form of the Nuscenes dataset, the collision logic filtering is performed according to the labeling of the second trajectory information to obtain labeled trajectory information.

[0100] In a feasible implementation, the existing multi-view video generation model is usually based on the Nuscenes dataset, so in the process of generating safe key trajectories to safe key videos by using the multi-view video generation model, it is necessary to convert the trajectory signal into a labeled file in the format of the Nuscenes dataset, and there is no ready-made conversion tool for this part of the content.

[0101] The method for converting vehicle trajectory signals into Nuscenes dataset format labels proposed by the application is mainly implemented from the following key aspects:

[0102] Frequency alignment, considering that the original vehicle trajectory data and the sensor data (for example, the camera image is usually 12Hz, and the key frame is 2Hz) of the Nuscenes dataset may have different sampling frequencies, the patent application tool can intelligently adapt the frequency of the trajectory data.

[0103] Specifically, by linearly interpolating the translation vectors of the trajectory data, the smooth transition of the trajectory pose on the time axis is ensured. At the same time, the timestamps of all newly generated Nuscenes sample data (sample_data) are accurately updated according to the target frequency, thereby ensuring the time sequence consistency of the trajectory information and the sensor data.

[0104] Three-dimensional pose update, the input vehicle trajectory signal is usually represented in the form of two-dimensional trajectory information, and the two-dimensional trajectory is converted into a 3D translation vector (translation) and a 3D rotation quaternion (rotation) required by the Nuscenes annotation, so as to construct the pose of the vehicle in the three-dimensional space.

[0105] For the ego vehicle, its pose information will be directly used to update the ego_pose table in the Nuscenes dataset. For other related vehicles in the scene, their pose information is used to update the corresponding fields in the sample_annotation table. Although the input trajectory data is unified, the present application will accurately assign it to the corresponding different annotations under the Nuscenes dataset according to the token of the vehicle, ensuring the accuracy and compatibility of the pose information.

[0106] Selection of vehicle instances in the scene, the input processed in this step is the trajectory information generated based on the initial scene simulation of Nuscenes. In order to maintain the logical consistency of the generated scene, the tool reuses the original Nuscens annotation when processing vehicle annotation for information other than vehicle trajectory.

[0107] The present application only performs trajectory alignment and pose update on the vehicle instances that already exist in the initial scene. For vehicles that appear in subsequent frames in the Nuscenes original dataset, the tool will selectively remove them from the newly generated Nuscenes annotation. Due to changes in the ego vehicle trajectory, it may cause the problem of sudden appearance of vehicles. At the same time, for the vehicles in the initial scene, their original Bounding Box, category and other information will be reused, so as to ensure that the generated scene only contains vehicles that exist from the starting time, and maintains their original attributes, avoiding unnecessary complexity or inconsistency.

[0108] S6、According to the labeled trajectory information, a preset multi-view video generation model is used to generate a video to obtain a multi-view safety-critical video.

[0109] In a feasible implementation, the application utilizes the open-source multi-view autonomous driving generation framework OpenDWM to combine the obtained labeled trajectory information, and realizes multi-view safety-critical collision scene generation, the input of the framework being Nuscenes format data, from which a map, vehicle trajectory, ego trajectory and camera internal and external parameters are extracted as control signals, which are converted into time-series multi-view video data output.

[0110] The application provides an automatic driving safety-critical multi-view video automatic generation method, calculates a danger score of each non-ego vehicle according to vehicle information and map information, selects a vehicle with the highest danger score as a safety-critical vehicle, realizes vehicle authenticity and diversity, uses a controllable trajectory diffusion model for safety trajectory simulation, solves the problem of low efficiency of manual adjustment, can autonomously generate a large amount of high-quality videos, greatly reduces manual intervention and cost, and perfectly meets the input requirements of a visual end-to-end autonomous driving system, significantly improves data utilization and system development efficiency.

[0111] Figure 2 is an automatic driving safety-critical multi-view video automatic generation device block diagram provided by an embodiment of the application, and the device is used for the automatic driving safety-critical multi-view video automatic generation method. Figure 2 The device includes an information acquisition module 210, a safety-critical vehicle selection module 220, a total guidance value calculation module 230, a trajectory simulation module 240, an information labeling module 250 and a video generation module 260.

[0112] The information acquisition module 210 is used for acquiring ego vehicle information, non-ego vehicle information and map information of an initial scene of autonomous driving.

[0113] The safety-critical vehicle selection module 220 is used for selecting a safety-critical vehicle based on the map information according to the ego vehicle information and the non-ego vehicle information, and obtaining safety-critical vehicle information.

[0114] The total guidance value calculation module 230 is used for calculating a total guidance value based on an inference time guidance function according to the map information, the ego vehicle information, the non-ego vehicle information and the safety-critical vehicle information.

[0115] The trajectory simulation module 240 is used for simulating a safety trajectory based on the total guidance value according to the map information, the ego vehicle information, the non-ego vehicle information and the safety-critical vehicle information using a pre-trained diffusion model, and obtaining simulation trajectory information.

[0116] The information labeling module 250 is configured to label the simulation trajectory information based on a data form of the Nuscenes dataset to obtain labeled trajectory information.

[0117] The video generation module 260 is configured to generate a video based on the labeled trajectory information and using a preset multi-view video generation model to obtain a multi-view safety-critical video.

[0118] Optionally, the safety-critical vehicle selection module 220 is further configured to:

[0119] The Euclidean distance set is calculated based on the ego vehicle information and the non-ego vehicle information.

[0120] The relative speed absolute value set is calculated based on the ego vehicle information and the non-ego vehicle information.

[0121] The map conflict factor set is obtained by performing conflict evaluation based on the map information, the ego vehicle information and the non-ego vehicle information, based on a preset conflict value.

[0122] The danger degree score set is calculated based on the Euclidean distance set, the relative speed absolute value set and the map conflict factor set.

[0123] The non-ego vehicle information corresponding to the maximum value in the danger degree score set is determined as the safety-critical vehicle information.

[0124] Optionally, the safety-critical vehicle selection module 220 is further configured to:

[0125] The first verification result set is obtained by performing same-direction following conflict verification on the non-ego vehicle information based on the map information and the ego vehicle information.

[0126] The second verification result set is obtained by performing intersection trajectory conflict verification on the non-ego vehicle information based on the map information and the ego vehicle information.

[0127] The third verification result set is obtained by performing adjacent lane cut-in conflict verification on the non-ego vehicle information based on the map information and the ego vehicle information.

[0128] The map conflict factor set is obtained by assigning conflict factors to the non-ego vehicle based on the first verification result set, the second verification result set and the third verification result set, based on a preset conflict value.

[0129] The map conflict factor types in the map conflict factor set include a high conflict factor, a medium conflict factor and a low conflict factor.

[0130] Optionally, the total guidance value calculation module 230 is further configured to:

[0131] The first guide value is obtained according to map information, ego vehicle information and safety key vehicle information based on a vehicle influence guide function;

[0132] The second guide value is obtained according to map information, ego vehicle information and non-ego vehicle information based on an obstacle avoidance safety guide function;

[0133] The third guide value is obtained according to map information and ego vehicle information based on a road surface adhesion guide function;

[0134] The overall guide value is obtained by weighted calculation according to the first guide value, the second guide value and the third guide value.

[0135] The vehicle influence guide function is calculated as formula (1) as follows:

[0136] (1);

[0137] Wherein, is the distance between the ego vehicle and the safety key vehicle in the future K time steps; is the minimum safety distance of the ego vehicle; is an indicator function, which returns 1 when the condition in the parentheses is true, and returns 0 when the condition in the parentheses is not true;

[0138] The obstacle avoidance safety guide function is calculated as formula (2) as follows:

[0139] (2);

[0140] Wherein, is the distance between the ego vehicle x with a speed of and the non-ego vehicle y with a speed of in the future K time steps; is the minimum safety distance between the ego vehicle x and the non-ego vehicle y; is a safety mechanism mask; is the static threshold of the ego vehicle x relative to the non-ego vehicle y;

[0141] The road surface adhesion guide function is calculated as formula (3) as follows:

[0142] (3);

[0143] Wherein, is the sampling point position; is the off-road area; is the on-road point position; is the on-road area; is the Euclidean distance calculation; is the diagonal length of the vehicle bounding box of the ego vehicle.

[0144] Optionally, the information labeling module 250 is further used for:

[0145] Based on the data form of the Nuscenes dataset, frequency alignment is performed according to the translation vector of the simulation trajectory information, and first trajectory information is obtained;

[0146] Based on the data form of the Nuscenes dataset, three-dimensional pose updating is performed according to the two-dimensional trajectory of the first trajectory information, and second trajectory information is obtained;

[0147] Based on the data form of the Nuscenes dataset, collision logic filtering is performed according to the labeling of the second trajectory information, and labeled trajectory information is obtained.

[0148] The present application provides an automatic driving safety key multi-view video automatic generation method, calculates the danger score of each non-self vehicle according to vehicle information and map information, selects the vehicle with the highest danger score as the safety key vehicle, realizes the vehicle authenticity and diversity, uses the controllable trajectory diffusion model to simulate the safety trajectory, solves the low efficiency problem of manual adjustment, can automatically generate a large amount of high-quality video, greatly reduces the artificial intervention and cost, the present application automatically generates multi-view safety key video, perfectly fits the input requirements of the visual end-to-end automatic driving system, significantly improves the data utilization and system development efficiency. The present application is an efficient and accurate multi-view video automatic generation method for automatic driving safety key.

[0149] Figure 3 is a structural schematic diagram of a multi-view video automatic generation device provided by an embodiment of the present application, as Figure 3 shown, the multi-view video automatic generation device can include the automatic driving safety key multi-view video automatic generation apparatus shown in Figure 2 above. Optionally, the multi-view video automatic generation device 310 can include a first processor 2001.

[0150] Optionally, the multi-view video automatic generation device 310 can further include a memory 2002 and a transceiver 2003.

[0151] Among them, the first processor 2001 and the memory 2002 and the transceiver 2003, such as can be connected through the communication bus.

[0152] The following will be Figure 3 introduced in detail:

[0153] The first processor 2001 is a control center of the multi-view video automated generation device 310, and can be one processor or a collective term of multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), and can also be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present application, such as one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0154] Optionally, the first processor 2001 can execute various functions of the multi-view video automated generation device 310 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0155] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as the CPU0 and the CPU1 shown in FIG. 2. Figure 3

[0156] In a specific implementation, as an embodiment, the multi-view video automated generation device 310 can also include multiple processors, such as the first processor 2001 and the second processor 2004 shown in FIG. 2. Each of the processors can be a single-CPU or a multi-CPU. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Figure 3

[0157] The memory 2002 is configured to store software programs for implementing the schemes of the present application, and the first processor 2001 is configured to control execution. For specific implementation, refer to the above method embodiments, which will not be repeated here.

[0158] ​​Optionally, the memory 2002 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 2002 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through an interface circuit (not shown in the figure) of the multi-view video automated generation device 310, and the embodiments of the present application do not make specific limitations here. Figure 3

[0159] The transceiver 2003 is configured to communicate with a network device or a terminal device.

[0160] Optionally, the transceiver 2003 can include a receiver and a transmitter (not shown separately in the figure). The receiver is configured to implement a receiving function, and the transmitter is configured to implement a transmitting function. Figure 3

[0161] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through an interface circuit (not shown in the figure) of the multi-view video automated generation device 310, and the embodiments of the present application do not make specific limitations here. Figure 3

[0162] It should be noted that the structure of the multi-view video automated generation device 310 shown in the figure does not constitute a limitation on the router, and the actual multi-view video automated generation device can include more or fewer components than those shown in the figure, or combine certain components, or different component arrangements. Figure 3 In addition, the technical effects of the multi-view video automated generation device 310 can refer to the technical effects of the autonomous driving safety-critical multi-view video automated generation method described in the above method embodiments, which will not be repeated here.

[0163]

[0164] ​​​​It is to be understood that the first processor 2001 in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or can also be any conventional processor.

[0165] It is also to be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM) and direct rambus RAM (DR RAM).

[0166] The above-described embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.

[0167] It should be understood that the term "and / or" herein merely describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents that the associated objects before and after it are in an "or" relationship, but it can also represent an "and / or" relationship, which can be understood according to the context before and after it.

[0168] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0169] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0170] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0171] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0172] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0173] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0174] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.

[0175] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0176] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An automated safety-critical multi-view video automated generation method, characterized in that, The method comprises: obtaining self-vehicle information, non-self-vehicle information and map information of an initial autonomous driving scene; based on the map information, performing safety key vehicle selection according to the self-vehicle information and the non-self-vehicle information to obtain safety key vehicle information; based on a reasoning guidance function, performing calculation according to the map information, the self-vehicle information, the non-self-vehicle information and the safety key vehicle information to obtain an overall guidance value, comprising: based on a vehicle influence guidance function, performing calculation according to the map information, the self-vehicle information and the safety key vehicle information to obtain a first guidance value; the vehicle influence guidance function is calculated as follows (1): (1); wherein, is the distance of the ego car - safety critical car in the future K time steps; is the minimum safety distance of the ego car; is an indicator function that returns 1 when the condition in the parentheses is true, and 0 when the condition in the parentheses is not true; based on an obstacle avoidance safety guidance function, performing calculation according to the map information, the self-vehicle information and the non-self-vehicle information to obtain a second guidance value; the obstacle avoidance safety guidance function is calculated as follows (2): (2); wherein, is the distance between the ego vehicle x with speed and the non-ego vehicle y with speed in the future K time steps; is the minimum safe distance between the ego vehicle x and the non-ego vehicle y; is the safety mechanism mask; is the stationary threshold of the ego vehicle x relative to the non-ego vehicle y; based on a road surface adhesion guidance function, performing calculation according to the map information and the self-vehicle information to obtain a third guidance value; the road surface adhesion guidance function is calculated as follows (3): (3); wherein, is a sampling point position; is an off-road area; is an on-road point position; is an on-road area; is a Euclidean distance calculation; is a diagonal length of a vehicle bounding box of the ego vehicle; performing weighted calculation according to the first guidance value, the second guidance value and the third guidance value to obtain the overall guidance value; based on the overall guidance value, performing safety trajectory simulation using a pre-trained diffusion model according to the map information, the self-vehicle information, the non-self-vehicle information and the safety key vehicle information to obtain simulation trajectory information; based on the data form of the Nuscenes dataset, labeling the simulation trajectory information to obtain labeled trajectory information; using a preset multi-view video generation model to generate a video according to the labeled trajectory information to obtain a multi-view safety key video. 2.The automated generation of safety-critical multi-view video for autonomous driving method of claim 1, wherein, based on the map information, performing safety key vehicle selection according to the self-vehicle information and the non-self-vehicle information to obtain safety key vehicle information, comprising: performing calculation according to the self-vehicle information and the non-self-vehicle information to obtain a Euclidean distance set; performing calculation according to the self-vehicle information and the non-self-vehicle information to obtain a relative speed absolute value set; based on a preset conflict value, performing conflict evaluation according to the map information, the self-vehicle information and the non-self-vehicle information to obtain a map conflict factor set; performing calculation according to the Euclidean distance set, the relative speed absolute value set and the map conflict factor set to obtain a danger degree score set; selecting non-self-vehicle information corresponding to the maximum value in the danger degree score set to determine as the safety key vehicle information. 3.The automated generation of safety-critical multi-view video for autonomous driving method of claim 2, wherein, based on the preset conflict value, performing conflict evaluation according to the map information, the self-vehicle information and the non-self-vehicle information to obtain a map conflict factor set, comprising: performing same-direction following conflict verification on the non-self-vehicle information according to the map information and the self-vehicle information to obtain a first verification result set; performing intersection trajectory conflict verification on the non-self-vehicle information according to the map information and the self-vehicle information to obtain a second verification result set; performing adjacent lane cut-in conflict verification on the non-self-vehicle information according to the map information and the self-vehicle information to obtain a third verification result set; based on the preset conflict value, assigning conflict factors to the non-self-vehicle according to the first verification result set, the second verification result set and the third verification result set to obtain the map conflict factor set. 4.The method of claim 3, wherein, the map conflict factor types in the map conflict factor set comprise high conflict factors, medium conflict factors and low conflict factors.

5. The automated generation of safety-critical multi-view video for autonomous driving method of claim 1, wherein, The data form based on the Nuscenes dataset is used to label the simulation trajectory information to obtain labeled trajectory information, including: The data form based on the Nuscenes dataset is used to perform frequency alignment according to the translation vector of the simulation trajectory information to obtain first trajectory information; The data form based on the Nuscenes dataset is used to perform three-dimensional pose updating according to the two-dimensional trajectory of the first trajectory information to obtain second trajectory information; The data form based on the Nuscenes dataset is used to perform collision logic filtering according to the labeling of the second trajectory information to obtain labeled trajectory information.

6. An apparatus for automated generation of safety-critical multi-view videos for autonomous driving, the apparatus being configured to implement the method for automated generation of safety-critical multi-view videos for autonomous driving according to any one of claims 1 to 5, characterized in that, The device comprises: An information acquisition module for acquiring ego vehicle information, non-ego vehicle information and map information of an initial autonomous driving scene; A safety-critical vehicle selection module for performing safety-critical vehicle selection based on the map information according to the ego vehicle information and the non-ego vehicle information to obtain safety-critical vehicle information; A total guidance value calculation module for calculating a total guidance value based on a guidance function at reasoning time according to the map information, the ego vehicle information, the non-ego vehicle information and the safety-critical vehicle information; A vehicle influence guidance function is used to calculate a first guidance value based on the map information, the ego vehicle information and the safety-critical vehicle information; The calculation of the vehicle influence guidance function is as follows (1): (1); wherein, is the distance of the ego car - safety critical car in the future K time steps; is the minimum safety distance of the ego car; is an indicator function that returns 1 when the condition in the parentheses is true, and 0 when the condition in the parentheses is not true; An obstacle avoidance safety guidance function is used to calculate a second guidance value based on the map information, the ego vehicle information and the non-ego vehicle information; The calculation of the obstacle avoidance safety guidance function is as follows (2): (2); wherein, is the distance between the ego vehicle x with speed and the non-ego vehicle y with speed in the future K time steps; is the minimum safe distance between the ego vehicle x and the non-ego vehicle y; is the safety mechanism mask; is the stationary threshold of the ego vehicle x relative to the non-ego vehicle y; A road surface adhesion guidance function is used to calculate a third guidance value based on the map information and the ego vehicle information; The calculation of the road surface adhesion guidance function is as follows (3): (3); wherein, is a sampling point position; is an off-road area; is an on-road point position; is an on-road area; is a Euclidean distance calculation; is a diagonal length of a vehicle bounding box of the ego vehicle; The first guidance value, the second guidance value and the third guidance value are weighted to obtain the total guidance value; A trajectory simulation module for performing safety trajectory simulation using a pre-trained diffusion model based on the total guidance value according to the map information, the ego vehicle information, the non-ego vehicle information and the safety-critical vehicle information to obtain simulation trajectory information; An information labeling module for labeling the simulation trajectory information based on the data form of the Nuscenes dataset to obtain labeled trajectory information; A video generation module for generating a video using a preset multi-view video generation model based on the labeled trajectory information to obtain a multi-view safety-critical video.

7. A multi-view video automatic generation apparatus characterized by comprising: The multi-view video automatic generation device comprises: A processor; A memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the method of any one of claims 1 to 5.

8. A computer readable storage medium, characterized in that, The computer readable storage medium stores program code that can be called and executed by the processor to implement the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Automatic driving video generation type prediction method guided by vehicle kinematics information

    CN120378703A

  • Automatic driving image data generation method and device, equipment and medium

    CN120655847A