A terminal state guided automatic driving high-risk scene data generation method and device
Patent Information
- Application Number
- CN202611047653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]本发明的目的是为了解决传统自动驾驶测试场景生成技术存在视觉瑕疵与物理谬误的问题,提出了一种终态引导的自动驾驶高风险场景数据生成方法及装置
Smart Images

Figure CN122551312A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of autonomous driving technology, specifically relating to a method and apparatus for generating data for high-risk autonomous driving scenarios with final state guidance. Background Technology
[0002] With the rapid development of autonomous driving technology, its safety verification in real traffic environments has become a key challenge restricting large-scale deployment and commercial application. Compared to high-cost and high-risk field testing, simulation-based verification methods, due to their controllable environment and repeatable scenarios, have become an important means of evaluating the performance of autonomous driving systems and exploring their safety boundaries. However, as autonomous driving architectures gradually evolve towards end-to-end cognition and decision-making based on large vision-language models, the underlying limitations of existing technologies in generating test scenarios with high realism, high physical threat levels, and high semantic diversity are becoming increasingly apparent.
[0003] On the one hand, traditional adversarial testing of autonomous driving heavily relies on rule-based or heuristic search-based 3D simulator pipelines. While these methods offer highly parametric control over the trajectories of vehicles or pedestrians, the generated images often possess strong synthetic texture features, simplified ambient lighting models, and are limited by pre-bound skeletal animation templates. This lack of representation leads to a significant gap between virtual and real-world transfer. More limitingly, the adversarial target behavior in traditional simulation environments is constrained by a predefined maneuver state space, failing to realistically replicate the semantic diversity of complex real-world traffic environments. For example, in simulated line-of-sight occlusion conflict scenarios, pedestrian actions often consist only of movements like moving and standing still, while in the real world, pedestrians often exhibit subtle movements such as leaning forward. This dual lack of visual detail and physical realism makes simulation scenarios easily obscured by the generalization capabilities of large autonomous driving models, failing to effectively expose the system's potential vulnerabilities in long-tailed complex interactions.
[0004] On the other hand, while deep learning-based video generation models have significantly improved pixel-level visual realism in recent years, they struggle to achieve deterministic synthesis of high-risk scenes. First, most existing video generation technologies employ open-loop forward prediction mechanisms based on initial frames. Since their training sets primarily consist of massive amounts of conventional, safe real-world driving data, the models exhibit strong safety biases. When performing video inference in adversarial scenarios, the models often converge to high-frequency safety events based on maximum likelihood, thus spontaneously correcting their adversarial intent and creating logical illusions that violate real-world physics, such as pedestrians instantly disappearing or conflict trajectories being forcibly stretched. Second, conventional 2D image fusion or simple object insertion techniques used to address these issues lack effective perception of the physical depth and geometric topology of the 3D world, easily leading to pixel-level artifacts and spatial geometric distortions in the synthesized adversarial targets, such as missing contact shadows.
[0005] These underlying visual flaws and physical errors can be easily identified and intercepted in advance by the anomaly detection module or the underlying perception noise filtering mechanism of the autonomous driving front end. As a result, the synthetic adversarial target cannot penetrate and reach the high-order cognitive reasoning and planning decision-making module of the autonomous driving system. Consequently, it is difficult to accurately, comprehensively and objectively assess the decision safety of the autonomous driving system based on the multimodal large model under extreme pressure. Summary of the Invention
[0006] The purpose of this invention is to solve the problems of visual defects and physical errors in traditional autonomous driving test scenario generation technology, and to propose a final state-guided method and device for generating data for high-risk autonomous driving scenarios.
[0007] The technical solution of this invention is as follows: This invention provides a method for generating data for high-risk autonomous driving scenarios with final-state guidance, comprising: Step 1: Acquire basic traffic video and extract the first frame of the basic traffic video, which is a sequence of continuous pixel frames; Step 2: Perform visual depth estimation and geometric feature decomposition on each frame of the basic traffic video to obtain a dense depth map and a tag set image corresponding to each frame. The tag set image contains discrete digital labels used to identify each key semantic target in the scene. Step 3: Input the dense depth map and the labeled set image corresponding to each frame in the basic traffic video as joint visual features into the multimodal large-scale language model, identify the video frame where the interaction vulnerability point in the basic traffic video is located as the selected frame, and infer and output the catastrophic final state specification. Step 4: Based on the catastrophic final state specification, generate an adversarial proxy image using an image generation model, and fuse the adversarial proxy image into the selected frame by combining physical constraints and image inpainting techniques to obtain an adversarial catastrophic final state image; Step 5: Using the first frame of the extracted basic traffic video as the starting anchor point constraint and the adversarial disaster final state image as the ending anchor point constraint, use the video generation model to perform time-series dynamic simulation and synthesize a video of key safety scenarios for autonomous driving.
[0008] Preferably, step 2 includes: Step 2.1: Use a depth estimation model to perform visual depth estimation on each frame of the basic traffic video to obtain the dense depth map that reflects the three-dimensional spatial distribution characteristics of the traffic scene; Step 2.2: Use the panoramic segmentation algorithm to perform pixel-level semantic analysis on each frame of the basic traffic video, identify key semantic targets with potential physical interaction value, and generate an independent local semantic mask for each key semantic target; Step 2.3: For each frame in the basic traffic video, remove redundant masks that do not meet the requirements of adversarial experiments from the local semantic mask according to the preset multi-level filtering rules, assign a unique digital identifier to the retained local semantic mask, and visually overlay the digital identifier and the corresponding local semantic mask on the current frame image to obtain the tag set image.
[0009] Preferably, step 3 includes: Step 3.1: Stack and stitch the labeled image corresponding to each frame in the basic traffic video with the normalized dense depth map to obtain the joint visual features, and input the joint visual features corresponding to each frame in the basic traffic video into the multimodal large-scale language model; Step 3.2: The multimodal large-scale language model evaluates the collision risk of each frame in the basic traffic video based on the joint visual features, determines the video frame where the interaction vulnerability is located as the selected frame based on the collision risk assessment results, performs logical filtering on the digital identifiers in the selected frame, selects the blind spot with the most security threat, and specifies the physical intrusion vector parameters of the adversarial agent in the disaster final state based on the selected blind spot. Step 3.3: The multimodal large-scale language model, based on a two-stage prompting strategy, infers and outputs the catastrophic final state specification, which includes: the appearance attribute specification of the adversarial agent, the local spatial constraints of the target insertion region, and the dynamic action logic specification of the catastrophic final state.
[0010] Preferably, in step 3.3, the two-stage prompting strategy includes: The first phase describes the constant physical identity attributes of adversarial agents using natural language; The second stage describes the dynamic actions and postures of the adversarial agent in the catastrophe final state using natural language.
[0011] Preferably, step 4 includes: Step 4.1: Based on the catastrophic final state specification, generate a high-resolution image using an image generation model, and perform foreground separation processing on the high-resolution image to extract an adversarial proxy asset without background. Step 4.2: Determine the target metric depth according to the catastrophic final state specification, calculate the pixel scaling ratio of the adversarial proxy asset at the target metric depth in the selected frame based on the pinhole camera perspective model, and perform geometric scaling on the adversarial proxy asset according to the pixel scaling ratio; Step 4.3: Determine the local exploration space according to the catastrophic final state specification, detect effective ground with flat features in the local exploration space of the selected frame, solve for the minimum value of the constructed local cost minimization objective function according to the candidate insertion coordinate set corresponding to the effective ground, and determine the optimal insertion coordinate of the adversarial proxy asset; Step 4.4: Superimpose the adversarial proxy assets after geometric scaling onto the optimal insertion coordinates of the selected frame to obtain a preliminary superimposed image. Perform boundary dilation on the mask of the adversarial proxy assets in the preliminary superimposed image to generate a repair mask. Perform multiple rounds of denoising iteration with linear dynamic attenuation intensity on the preliminary superimposed image within the repair mask area to obtain the adversarial catastrophe final state image.
[0012] Preferably, in step 4.2, the pixel scaling ratio is expressed as: ; in, Indicates the pixel scaling ratio. This represents the reference height benchmark value of the adversarial proxy asset in the real physical world. This indicates the physical focal length parameter of the camera lens corresponding to the scene. Indicates the depth of the target measurement. It is a lower bound truncation function. This is the lower limit of the safety clamping protection distance.
[0013] Preferably, in step 4.3, the objective function for minimizing local costs is expressed as: ; in, Indicates the candidate insertion coordinates The comprehensive value of the product and These represent the center coordinates of the most threatening blind spot in the selected frame. This represents the weighting factor for lateral spatial displacement. This represents the longitudinal spatial displacement weighting factor. This indicates that the center is approaching the penalty item.
[0014] Preferably, in step 4.4, the noise interference intensity in each round of the multi-round denoising iteration process with linear dynamic attenuation intensity is... Degradation is performed according to the following constraint equations: ; in, Indicates the index of the current iteration round. This represents the total number of iterations, and ; This represents the initial denoising intensity threshold. This indicates the threshold for terminating the denoising intensity.
[0015] Preferably, step 5 includes: Step 5.1: Use a logic parsing program to extract the location information of the adversarial agent in the adversarial disaster final state image; Step 5.2: Based on the location information of the adversarial agent, perform cross-consistency logic intervention processing to forcibly set the starting orientation condition of the adversarial agent in the first frame; Step 5.3: The starting orientation condition is injected into the video generation model as a negative cue rule. Based on the starting anchor point constraint and the ending anchor point constraint, the video generation model is able to deduce a coherent video with a substantial spatial crossing trajectory, which serves as the video for the key safety scenario of autonomous driving.
[0016] This invention provides a final-state guided high-risk scenario data generation system for autonomous driving, applicable to the final-state guided high-risk scenario data generation method for autonomous driving described in any of the preceding embodiments, comprising: The video acquisition module is used to acquire basic traffic video and extract the first frame of the basic traffic video, wherein the basic traffic video is a sequence of continuous pixel frames. The video preprocessing module is used to perform visual depth estimation and geometric feature decomposition on each frame of the basic traffic video to obtain a dense depth map and a tag set image corresponding to each frame. The tag set image contains discrete digital tags used to identify each key semantic target in the scene. The final state inference module is used to input the dense depth map and the labeled set image corresponding to each frame in the basic traffic video as joint visual features into the multimodal large-scale language model, identify the video frame where the interaction vulnerability point in the basic traffic video is located as the selected frame, and infer and output the catastrophic final state specification. The final state fusion module is used to generate an adversarial proxy image based on the catastrophic final state specification using an image generation model, and to fuse the adversarial proxy image into the selected frame by combining physical constraints and image inpainting techniques to obtain an adversarial catastrophic final state image. The video synthesis module is used to synthesize videos of key safety scenarios for autonomous driving by using the first frame of the extracted basic traffic video as the starting anchor point constraint and the adversarial disaster final state image as the ending anchor point constraint, and by performing time-series dynamic simulation using a video generation model.
[0017] The beneficial effects of this invention are: The present invention provides a final-state-guided method for generating high-risk scenario data for autonomous driving. By constructing bidirectional temporal constraints at both ends, it forces the model to evolve towards a catastrophic final state, effectively overcoming the tendency of traditional open-loop forward prediction mechanisms to always converge to low-risk events. This provides high-value, deterministic adversarial test assets for autonomous driving systems. Through visual depth estimation combined with purely logical discrete digital indexing, along with physical constraints and image inpainting techniques, it fundamentally eliminates geometric fallacies of the adversarial target, improves the physical realism of the generated scene, and utilizes forced code rules to constrain the behavioral boundaries of the generation model, effectively avoiding spatiotemporal logical discontinuities that easily occur during video generation that violate physical principles. Attached Figure Description
[0018] Figure 1 The diagram shows a flowchart of a method for generating data for high-risk autonomous driving scenarios guided by the final state. Figure 2 The diagram shows the logical architecture of the data generation method for high-risk autonomous driving scenarios guided by the final state. Figure 3 The diagram shows the structural block diagram of a high-risk scenario data generation system for autonomous driving guided by the final state. Detailed Implementation
[0019] Exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary and are intended to illustrate the principles and spirit of the invention, and are not intended to limit the scope of the invention.
[0020] In the first aspect, embodiments of the present invention provide a method for generating high-risk scenario data for autonomous driving guided by the final state. The core concept is as follows: first, a catastrophic final state with deterministic collision risk is determined; then, starting from a benign first frame and ending at the catastrophic final state, the intermediate process is generated in reverse using a video diffusion model constrained by the first and last frames; at the same time, physical geometric constraints and logical interventions are introduced to generate a high-fidelity, high-threat test scenario video.
[0021] Please see Figure 1 and Figure 2 , Figure 1 The diagram shows a flowchart of a method for generating data for high-risk autonomous driving scenarios guided by the final state. Figure 2 The diagram shows the logical architecture of the final state-guided method for generating data in high-risk scenarios of autonomous driving.
[0022] like Figure 1 As shown in this embodiment, the method for generating high-risk scenario data for autonomous driving with final state guidance includes the following steps: Step 1: Acquire basic traffic video and extract the first frame of the basic traffic video. The basic traffic video is a sequence of continuous pixel frames.
[0023] In this embodiment, a sequence of continuous pixel frames containing natural street scenes can be extracted from an existing real autonomous driving dataset as a basic traffic video, and the first frame on the timeline of the basic traffic video can be extracted as the first frame to provide an initial anchor point constraint for subsequent video spatiotemporal extrapolation, which is used to constrain the initial state of video generation.
[0024] Understandably, the basic traffic video does not contain any adversarial events; that is, the basic traffic video is a continuous sequence of pixel frames representing a benign, normal scene.
[0025] Step 2: Perform visual depth estimation and geometric feature decomposition on each frame of the basic traffic video to obtain the dense depth map and the label set image corresponding to each frame. The label set image contains discrete digital labels used to identify each key semantic target in the scene.
[0026] In this embodiment, step 2 includes: Step 2.1: Use a depth estimation model to perform visual depth estimation on each frame of the basic traffic video to obtain a dense depth map that reflects the three-dimensional spatial distribution characteristics of the traffic scene.
[0027] Alternatively, the DepthAnythingV3 (General 3D Visual Geometry Reconstruction) model can be used to perform visual depth estimation on each frame of the underlying traffic video, outputting a dense depth map with the same resolution as the original image.
[0028] In this embodiment, the dense depth map maps each pixel to a relative or absolute physical spatial distance, reflecting the three-dimensional spatial distribution characteristics of the traffic scene.
[0029] Step 2.2: Use the panoramic segmentation algorithm to perform pixel-level semantic analysis on each frame of the basic traffic video, identify key semantic targets with potential physical interaction value, and generate an independent local semantic mask for each key semantic target.
[0030] Alternatively, the MobileSAM (Mobile Segmentation Model) algorithm can be used to perform pixel-level semantic parsing on each frame of the basic traffic video.
[0031] For example, key semantic targets with potential physical interaction value can be objects that may create blind spots, such as parked vehicles, roadside green belts, and large traffic signs.
[0032] Step 2.3: For each frame in the basic traffic video, remove redundant masks that do not meet the requirements of the adversarial experiment from the local semantic mask according to the preset multi-level filtering rules, assign a unique digital identifier to the retained local semantic mask, and visually overlay the digital identifier and the corresponding local semantic mask on the current frame image to obtain the labeled set image.
[0033] In this embodiment, the multi-level filtering rules include: (1) Remove masks with depth values greater than a preset distance threshold corresponding to dense depth maps, that is, remove distant targets with lower threat. For example, the preset distance threshold is 50 meters.
[0034] (2) Remove masks whose pixel area accounts for less than a preset area ratio threshold of the total pixels in the original image, i.e., remove noise or irrelevant targets. For example, the preset area ratio threshold is 0.5%.
[0035] Optionally, a unique numerical identifier, such as 1, 2, 3, etc., can be assigned to each valid mask that remains after filtering.
[0036] Step 3: Input the dense depth map and labeled image corresponding to each frame in the basic traffic video as joint visual features into the multimodal large-scale language model, identify the video frames where the interactive vulnerability points in the basic traffic video are located as selected frames, and infer and output the catastrophic final state specification.
[0037] In this embodiment, step 3 includes: Step 3.1: Stack and stitch the labeled image corresponding to each frame in the basic traffic video with the normalized dense depth map to obtain joint visual features. Input the joint visual features corresponding to each frame in the basic traffic video into the multimodal large-scale language model.
[0038] In this embodiment, the dense depth map corresponding to each frame in the basic traffic video is normalized and used as an independent visual channel. It is then stacked and stitched with a tag set image (RGB image) that is superimposed with digital identifiers and semantic masks to form a joint visual feature, which is then input into a multimodal large-scale language model.
[0039] Alternatively, the multimodal large-scale language model can be Qwen3-VL-32B.
[0040] Step 3.2: The multimodal large-scale language model evaluates the collision risk of each frame in the basic traffic video based on joint visual features. Based on the collision risk assessment results, the video frame where the interaction vulnerability is located is determined as the selected frame. The digital identifiers in the selected frame are logically filtered to select the blind spot with the greatest security threat. Based on the selected blind spot, the physical intrusion vector parameters of the adversarial agent at the end state of the disaster are specified.
[0041] In this embodiment, the multimodal large-scale language model evaluates the occlusion relationship, relative velocity, and collision time (TTC) of the target in each frame based on the input joint visual feature sequence, and selects the frame with the highest collision risk as the selected frame, which serves as the base frame for the subsequent final image.
[0042] In this embodiment, security threats can be quantified using Time-of-Collision (TTC). The multimodal large-scale language model anchors the blind spot with the highest security threat by solving for the coordinates that minimize TTC within the effective area of a selected frame, and accordingly specifies the physical intrusion direction of the adversarial agent in the final state, for example, rushing out from left to right. The collision time... , Depth is measured for the target, provided by a dense depth map. The vehicle speed can be estimated from basic traffic video.
[0043] Step 3.3: The multimodal large-scale language model is based on a two-stage prompting strategy to infer catastrophic final state specifications.
[0044] In this embodiment, the two-stage prompting strategy includes: The first phase describes the constant physical identity attributes of adversarial agents using natural language.
[0045] For example, the prompt for generating a constant physical identity attribute could be: "Adult male, wearing a dark gray jacket...".
[0046] The second stage describes the dynamic actions and postures of the adversarial agent in the catastrophe final state using natural language.
[0047] For example, the prompt words for generating dynamic actions and posture features at the end of a disaster could be: "A pedestrian suddenly rushes out, runs quickly to the left, leans forward, and focuses on crossing the lane...".
[0048] In this embodiment, the catastrophic final state specification includes: the appearance attribute specification of the adversarial agent, the local spatial constraints of the target insertion region, and the dynamic action logic specification of the catastrophic final state. To achieve precise semantic control and physical alignment, the process of outputting the catastrophic final state specification by the multimodal large-scale language model follows the following multimodal mapping equation at the underlying logic level: ; in, This represents the catastrophic final state specification of the output. The appearance attribute specification for an adversarial agent is represented as "adult male, wearing a dark gray jacket and blue jeans"; The local spatial constraints representing the target insertion region include: the coordinates of the center of the blind spot. Information such as target depth and intrusion direction; A dynamic action logic specification representing the final state of a disaster. Represents the inference mapping function of a multimodal large-scale language model. , , These represent the color pixel features, metric depth features, and digital identifier mask features of the labeled image set, respectively. This represents the input aggressive task instructions. The joint mapping equation ensures that the generated specification not only possesses semantic rationality but is also strictly constrained by three-dimensional physical depth.
[0049] Step 4: Based on the catastrophic final state specification, generate adversarial proxy images using an image generation model. Combine physical constraints and image inpainting techniques to fuse the adversarial proxy images into selected frames to obtain adversarial catastrophic final state images.
[0050] In this embodiment, step 4 includes: Step 4.1: Based on the catastrophic final state specification, generate high-resolution images using an image generation model, and extract background-free adversarial proxy assets by performing foreground separation processing on the high-resolution images.
[0051] In this embodiment, an image generation model is used to generate a high-resolution image based on the textual prompts describing the appearance attributes of the adversarial agent in the catastrophic final state specification. This high-resolution image is then used to extract a background-free RGBA image from the foreground separation algorithm (such as background removal or matting), resulting in an adversarial agent asset with an alpha (transparency channel).
[0052] Optionally, the image generation model can employ Stable Diffusion 3.5.
[0053] Step 4.2: Determine the target metric depth according to the catastrophic final state specification, calculate the pixel scaling ratio of the adversarial proxy asset at the target metric depth in the selected frame based on the pinhole camera perspective model, and perform geometric scaling on the adversarial proxy asset according to the pixel scaling ratio.
[0054] In this embodiment, the pixel scaling ratio is expressed as: ; in, Indicates the pixel scaling ratio. This represents the reference height of the adversarial proxy asset in the real physical world; for example, it could be 1.75 meters. This indicates the physical focal length parameter of the camera lens corresponding to the scene. Indicates the depth of the target measurement. This is a lower bound truncation function. By using the lower bound truncation function, numerical overflow distortion caused by extremely close spatial distance calculations can be effectively eliminated. For example, the lower limit of the safety clamping protection distance can be 2 meters.
[0055] It is understandable that when the lower limit parameter of the safety clamp protection distance is... When triggered, it forces a truncation to avoid screen crashes caused by extremely close proximity.
[0056] Step 4.3: Determine the local exploration space according to the catastrophic final state specification, detect effective ground with flat features in the local exploration space of the selected frame, solve for the minimum value of the constructed local cost minimization objective function based on the candidate insertion coordinate set corresponding to the effective ground, and determine the optimal insertion coordinates of the adversarial proxy asset.
[0057] In this embodiment, the local exploration space refers to a rectangular or circular neighborhood centered on the center of the blind spot, with a certain width and height. For example, the local exploration space can be a window centered on the center of the blind spot, with a side length of 50-100 pixels. This local exploration space defines the search range for physically ground-based optimization.
[0058] The specific process for detecting effective ground with flat features is as follows: Set a local search window, for example, 50×50 pixels, in the local exploration space of the selected frame. Calculate the mean square error of the depth values of all pixels within the window. Regions with a mean square error less than a preset threshold (e.g., 0.05 or 0.1) are determined to be effective flat ground. Otherwise, slide the window to an adjacent position to continue the detection until all window regions that meet the flatness condition are found in the local exploration space, forming a set of candidate insertion regions for adversarial proxy asset insertion. The pixel coordinates of all pixels in the candidate insertion region set constitute the candidate insertion coordinate set.
[0059] Furthermore, the optimal insertion coordinates of the adversarial proxy asset are determined by iterating through the set of candidate insertion coordinates to find the minimum value of the local cost minimization objective function.
[0060] In this embodiment, the objective function for minimizing local cost is expressed as: ; in, Indicates the candidate insertion coordinates The comprehensive value of the product and These represent the center coordinates of the most threatening blind spot in the selected frame. This represents the weighting factor for lateral spatial displacement. This represents the longitudinal spatial displacement weighting factor. and This is used to constrain the generated foot coordinates to not deviate from reasonable blind zone boundaries. The center approximation penalty term is used to apply a geometric gradient tendency, forcing candidate insertion coordinates to converge toward the geometric centerline of the vehicle's expected driving trajectory while satisfying the physical flatness property, in order to construct the maximum terminal interaction threat risk.
[0061] In this embodiment, the specific functional form of the center approximation penalty term is as follows: ,in This represents the lateral coordinates of the geometric centerline of the vehicle's predicted trajectory on the pixel plane. This is the penalty coefficient.
[0062] Step 4.4: Superimpose the adversarial proxy assets after geometric scaling onto the optimal insertion coordinates of the selected frame to obtain a preliminary superimposed image. Perform boundary dilation on the mask of the adversarial proxy assets in the preliminary superimposed image to generate a repair mask. Perform multiple rounds of denoising iteration with linear dynamic attenuation intensity on the preliminary superimposed image within the repair mask area to obtain the adversarial catastrophe final state image.
[0063] Specifically, based on the optimal insertion coordinates, the geometrically scaled adversarial proxy asset is initially superimposed onto the corresponding position in the selected frame using alpha blending. At this point, there may be issues such as harsh edges, inconsistent lighting, and missing shadows between the adversarial proxy asset and the background, requiring further processing.
[0064] Further, obtain the binary mask of the adversarial proxy asset. Perform a morphological dilation operation on this mask, expanding the mask area outward by several pixels, for example, dilating it by 5-10 pixels, to generate a repair mask slightly larger than the proxy area. The dilated repair mask is used to indicate the areas that need to be repaired in subsequent denoising iterations, ensuring that the transition zone between the edge of the adversarial proxy and the background is adequately processed, avoiding a harsh cutting effect.
[0065] Within the area defined by the repair mask, multiple rounds of denoising iterations with linear dynamic attenuation are performed on the selected frame, with the denoising interference intensity in each round being... Degradation is performed according to the following constraint equations: ; in, Indicates the index of the current iteration round. This represents the total number of iterations, and ; This represents the initial denoising intensity threshold. This indicates the threshold for terminating the denoising intensity. Less than , It can be 0.8. It can be 0.2.
[0066] In this embodiment, the initial denoising intensity threshold strongly blends the bottom of the adversarial proxy with the background environment, generating foot contact shadows that conform to the original street scene lighting mapping and eliminating boundary artifacts. The termination denoising intensity threshold preserves the high-frequency texture details of the adversarial proxy (such as clothing wrinkles and facial features) to avoid over-smoothing. Each round of denoising operation performs noise prediction and residual noise removal on the pixels within the repair mask while maintaining the structural information of the proxy subject. In the multi-round denoising process, each step is typically combined with an alpha composite operation to ensure a smooth transition of adversarial proxy pixels with background pixels according to transparency, ultimately outputting a fully blended adversarial catastrophe final state image.
[0067] Step 5: Using the first frame of the extracted basic traffic video as the starting anchor point constraint and the adversarial disaster final state image as the ending anchor point constraint, the video generation model is used to perform time-series dynamic simulation to synthesize videos of key safety scenarios for autonomous driving.
[0068] In this embodiment, step 5 includes: Step 5.1: Use a logic parsing program to extract the location information of the adversarial agent in the adversarial disaster final state image.
[0069] Specifically, using logical parsing procedures, such as object detection or pixel coordinate reading, the foot coordinates and lane sides of the adversarial agent are obtained from the adversarial disaster final state image, including: left, right and center.
[0070] Step 5.2: Based on the location information of the adversarial agent, perform cross-consistency logic intervention processing to forcibly set the starting orientation condition of the adversarial agent in the first frame.
[0071] In this embodiment, the cross-consistency logic intervention includes: Lateral opposition rule: If the final state position of the adversarial agent is determined to be on the preset first side of the vehicle lane, such as the center of the left side, then the adversarial agent is forced to set its initial motion entry point in the first frame to be on the opposite second side, such as the right blind spot. This simulates a complete trajectory that suddenly crosses from one side to the other.
[0072] Vertical opposition rule: If the final state position of the adversarial agent is determined to be in the vertical near blind zone, such as less than 5 meters away from the vehicle, then it is forced to be set to the vertical far blind zone in the first frame, such as more than 30 meters away, to avoid "instantaneous movement" and require the video generation model to generate a reasonable spatial trajectories.
[0073] Step 5.3: The starting orientation condition is injected into the video generation model as a negative cue rule. Based on the starting anchor point constraint and the ending anchor point constraint, the video generation model can deduce a coherent video with a substantial spatial crossing trajectory, which serves as a video for key safety scenarios of autonomous driving.
[0074] In this embodiment, the video generation model uses a video diffusion model with joint control capabilities for the first and last frames, such as Wan2.1-FLF2V (Wan2.1-FLF2V, the Wan-Phase First and Last Frame Model). This video generation model uses the first frame of the basic traffic video and the disaster final state image as structured rigid boundary conditions for the latent space noise reduction process. Negative cue rule injection specifically involves injecting words that violate physical laws, such as "pedestrians appear out of thin air," "instantaneous movement," and "location jump," as negative cue words into the sampling process of the video diffusion model through a classifier-free guidance mechanism, thereby forcibly suppressing spatiotemporal illusions during the generation process.
[0075] In this embodiment, the state evolution guided by the target conditions is performed using a video diffusion model. The underlying spatiotemporal deduction logic can be formally represented as a conditional probability generation equation with boundary constraints: ; in, This represents the final synthesized video of key safety scenarios for autonomous driving. This refers to a video diffusion network that possesses the capability for joint control of the first and last frames. The first frame serves as the starting anchor point; This is the final state image of an adversarial catastrophe, serving as the terminal anchor point. The dynamic action logic specification for the final state of a disaster; This represents the parameters of the cross-consistency enforcement rule that triggers the intervention. This is achieved through the initial and final bistate (…). ) as rigid boundary conditions, and code-level hard rules ( The injection into the latent space denoising trajectory of the diffusion model fundamentally eliminates the probability of the model spontaneously generating hallucinatory trajectories.
[0076] The present invention provides a final-state-guided method for generating high-risk scenario data for autonomous driving. By constructing bidirectional temporal constraints at both ends, it forces the model to evolve towards a catastrophic final state, effectively overcoming the tendency of traditional open-loop forward prediction mechanisms to always converge to low-risk events. This provides high-value, deterministic adversarial test assets for autonomous driving systems. Through visual depth estimation combined with purely logical discrete digital indexing, along with physical ground-hugging constraints based on center approximation and adaptive perspective scaling, it fundamentally eliminates geometric fallacies of adversarial targets, improves the physical realism of the generated scene, and introduces a cross-consistency logic intervention mechanism. By using forced code rules to constrain the behavioral boundaries of the generation model, it effectively avoids spatiotemporal logical discontinuities that easily occur during video generation, which violate the common sense of physics.
[0077] Secondly, embodiments of the present invention provide a final-state guided autonomous driving high-risk scenario data generation system, which can be used to implement the final-state guided autonomous driving high-risk scenario data generation method provided in the first aspect.
[0078] Please see Figure 3 , Figure 3 The diagram shows the structural block of a high-risk scenario data generation system for autonomous driving guided by the final state. Figure 3 As shown, the final state-guided autonomous driving high-risk scenario data generation system of this embodiment includes: a video acquisition module, a video preprocessing module, a final state inference module, a final state fusion module, and a video synthesis module.
[0079] The video acquisition module acquires basic traffic video and extracts its first frame, which is a sequence of continuous pixel frames. The video preprocessing module performs visual depth estimation and geometric feature decomposition on each frame of the basic traffic video, obtaining a dense depth map and a tag set image for each frame. The tag set image contains discrete digital labels used to identify key semantic targets in the scene. The final state inference module takes the dense depth map and tag set image corresponding to each frame of the basic traffic video as joint visual features and inputs them into a multimodal large-scale language model. It identifies the video frames containing interactive vulnerability points in the basic traffic video as selected frames and infers and outputs catastrophic final state specifications. The final state fusion module generates adversarial proxy images based on the catastrophic final state specifications using an image generation model. It then fuses these adversarial proxy images into the selected frames using physical constraints and image inpainting techniques to obtain an adversarial catastrophic final state image. The video synthesis module uses the first frame of the extracted basic traffic video as the starting anchor constraint and the adversarial catastrophic final state image as the ending anchor constraint. It then uses a video generation model to perform temporal dynamic extrapolation to synthesize videos of key safety scenarios for autonomous driving.
[0080] For details regarding the final state-guided autonomous driving high-risk scenario data generation system and its corresponding beneficial effects, please refer to the relevant content of the final state-guided autonomous driving high-risk scenario data generation method provided in the first aspect, which will not be elaborated here.
[0081] To verify the effectiveness of the safety-critical scenario videos generated by this invention against autonomous driving systems, a multi-dimensional Joint Security Assessment Index (JOS score) was constructed, and quantitative flaw detection was performed on various mainstream autonomous driving systems based on vision-language large models. The JOS score integrates the degradation of three dimensions of the system when facing threats: Perception Alignment (PA), Expectation Error (AE), and Planning Compliance (PC). By running under the same hardware environment, the experimental results compared the basic driving scenario with the high-risk scenario synthesized by this invention. Compared with traditional simulators and ordinary forward video generation models, the tested system showed a sharp drop in its Overall Defense Score (JOS) when facing the test scenario synthesized by this invention. This proves that the method of this invention not only effectively overcomes the safety bias in the generation stage and completely eliminates easily detectable visual and physical artifacts, but also successfully breaks through the perception and prediction defenses of autonomous driving systems through cross-lane game actions generated by head-to-tail constraints, exposing the vulnerability of its planning decisions.
[0082] It should be noted that the above embodiments are merely preferred examples for explaining the technical concept of the present invention. Those skilled in the art should realize that the synthesis method described in this invention is not limited to the nuScenes dataset or real street view videos, but can also be applied to test video sequences rendered and output by high-fidelity simulators (such as CARLA, Unreal Engine, etc.); at the same time, the specific model architecture mentioned in the above embodiments is only one option for implementing the function, and any alternative architecture with equivalent spatial awareness or condition generation capabilities is within the protection scope of this invention without departing from the core logic of this invention.
[0083] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that an article or apparatus comprising a list of elements includes not only those elements but also other elements not expressly listed. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or apparatus that includes said element. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0084] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0085] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A terminal state guided automatic driving high-risk scene data generation method, characterized in that, include: Step 1: Acquire basic traffic video and extract the first frame of the basic traffic video, which is a sequence of continuous pixel frames; Step 2: Perform visual depth estimation and geometric feature decomposition on each frame of the basic traffic video to obtain a dense depth map and a tag set image corresponding to each frame. The tag set image contains discrete digital labels used to identify each key semantic target in the scene. Step 3: Input the dense depth map and the labeled set image corresponding to each frame in the basic traffic video as joint visual features into the multimodal large-scale language model, identify the video frame where the interaction vulnerability point in the basic traffic video is located as the selected frame, and infer and output the catastrophic final state specification. Step 4: Based on the catastrophic final state specification, generate an adversarial proxy image using an image generation model, and fuse the adversarial proxy image into the selected frame by combining physical constraints and image inpainting techniques to obtain an adversarial catastrophic final state image; Step 5: Using the first frame of the extracted basic traffic video as the starting anchor point constraint and the adversarial disaster final state image as the ending anchor point constraint, use the video generation model to perform time-series dynamic simulation and synthesize a video of key safety scenarios for autonomous driving.
2. The end-to-end guided autonomous driving high-risk scenario data generation method according to claim 1, characterized in that, Step 2 includes: Step 2.1: Use a depth estimation model to perform visual depth estimation on each frame of the basic traffic video to obtain the dense depth map that reflects the three-dimensional spatial distribution characteristics of the traffic scene; Step 2.2: Use the panoramic segmentation algorithm to perform pixel-level semantic analysis on each frame of the basic traffic video, identify key semantic targets with potential physical interaction value, and generate an independent local semantic mask for each key semantic target; Step 2.3: For each frame in the basic traffic video, remove redundant masks that do not meet the requirements of adversarial experiments from the local semantic mask according to the preset multi-level filtering rules, assign a unique digital identifier to the retained local semantic mask, and visually overlay the digital identifier and the corresponding local semantic mask on the current frame image to obtain the tag set image.
3. The method for generating high-risk scenario data for autonomous driving with final state guidance according to claim 1, characterized in that, Step 3 includes: Step 3.1: Stack and stitch the labeled image corresponding to each frame in the basic traffic video with the normalized dense depth map to obtain the joint visual features, and input the joint visual features corresponding to each frame in the basic traffic video into the multimodal large-scale language model; Step 3.2: The multimodal large-scale language model evaluates the collision risk of each frame in the basic traffic video based on the joint visual features, determines the video frame where the interaction vulnerability is located as the selected frame based on the collision risk assessment results, performs logical filtering on the digital identifiers in the selected frame, selects the blind spot with the most security threat, and specifies the physical intrusion vector parameters of the adversarial agent in the disaster final state based on the selected blind spot. Step 3.3: The multimodal large-scale language model, based on a two-stage prompting strategy, infers and outputs the catastrophic final state specification, which includes: the appearance attribute specification of the adversarial agent, the local spatial constraints of the target insertion region, and the dynamic action logic specification of the catastrophic final state.
4. The method for generating high-risk scenario data for autonomous driving with final state guidance according to claim 3, characterized in that, In step 3.3, the two-stage prompting strategy includes: The first phase describes the constant physical identity attributes of adversarial agents using natural language; The second stage describes the dynamic actions and postures of the adversarial agent in the catastrophe final state using natural language.
5. The method for generating high-risk scenario data for autonomous driving with final state guidance according to claim 1, characterized in that, Step 4 includes: Step 4.1: Based on the catastrophic final state specification, generate a high-resolution image using an image generation model, and perform foreground separation processing on the high-resolution image to extract an adversarial proxy asset without background. Step 4.2: Determine the target metric depth according to the catastrophic final state specification, calculate the pixel scaling ratio of the adversarial proxy asset at the target metric depth in the selected frame based on the pinhole camera perspective model, and perform geometric scaling on the adversarial proxy asset according to the pixel scaling ratio; Step 4.3: Determine the local exploration space according to the catastrophic final state specification, detect effective ground with flat features in the local exploration space of the selected frame, solve for the minimum value of the constructed local cost minimization objective function according to the candidate insertion coordinate set corresponding to the effective ground, and determine the optimal insertion coordinate of the adversarial proxy asset; Step 4.4: Superimpose the adversarial proxy assets after geometric scaling onto the optimal insertion coordinates of the selected frame to obtain a preliminary superimposed image. Perform boundary dilation on the mask of the adversarial proxy assets in the preliminary superimposed image to generate a repair mask. Perform multiple rounds of denoising iteration with linear dynamic attenuation intensity on the preliminary superimposed image within the repair mask area to obtain the adversarial catastrophe final state image.
6. The method for generating high-risk scenario data for autonomous driving with final state guidance according to claim 5, characterized in that, In step 4.2, the pixel scaling ratio is expressed as: ; in, Indicates the pixel scaling ratio. This represents the reference height benchmark value of the adversarial proxy asset in the real physical world. This indicates the physical focal length parameter of the camera lens corresponding to the scene. Indicates the depth of the target measurement. It is a lower bound truncation function. This is the lower limit of the safety clamping protection distance.
7. The method for generating high-risk scenario data for autonomous driving with final state guidance according to claim 5, characterized in that, In step 4.3, the objective function for minimizing local costs is expressed as: ; in, Indicates the candidate insertion coordinates The comprehensive value of the product and These represent the center coordinates of the most threatening blind spot in the selected frame. This represents the weighting factor for lateral spatial displacement. This represents the longitudinal spatial displacement weighting factor. This indicates that the center is approaching the penalty item.
8. The method for generating high-risk scenario data for autonomous driving with final state guidance according to claim 5, characterized in that, In step 4.4, the noise interference intensity in each round of the multi-round denoising iteration process with linear dynamic attenuation intensity is... Degradation is performed according to the following constraint equations: ; in, Indicates the index of the current iteration round. This represents the total number of iterations, and ; This represents the initial denoising intensity threshold. This indicates the threshold for terminating the denoising intensity.
9. The method for generating high-risk scenario data for autonomous driving with final state guidance according to claim 1, characterized in that, Step 5 includes: Step 5.1: Use a logic parsing program to extract the location information of the adversarial agent in the adversarial disaster final state image; Step 5.2: Based on the location information of the adversarial agent, perform cross-consistency logic intervention processing to forcibly set the starting orientation condition of the adversarial agent in the first frame; Step 5.3: The starting orientation condition is injected into the video generation model as a negative cue rule. Based on the starting anchor point constraint and the ending anchor point constraint, the video generation model is able to deduce a coherent video with a substantial spatial crossing trajectory, which serves as the video for the key safety scenario of autonomous driving.
10. A data generation system for high-risk scenarios of autonomous driving with final state guidance, characterized in that, The method for generating high-risk autonomous driving scenario data with final state guidance as described in any one of claims 1-9 includes: The video acquisition module is used to acquire basic traffic video and extract the first frame of the basic traffic video, wherein the basic traffic video is a sequence of continuous pixel frames. The video preprocessing module is used to perform visual depth estimation and geometric feature decomposition on each frame of the basic traffic video to obtain a dense depth map and a tag set image corresponding to each frame. The tag set image contains discrete digital tags used to identify each key semantic target in the scene. The final state inference module is used to input the dense depth map and the labeled set image corresponding to each frame in the basic traffic video as joint visual features into the multimodal large-scale language model, identify the video frame where the interaction vulnerability point in the basic traffic video is located as the selected frame, and infer and output the catastrophic final state specification. The final state fusion module is used to generate an adversarial proxy image based on the catastrophic final state specification using an image generation model, and to fuse the adversarial proxy image into the selected frame by combining physical constraints and image inpainting techniques to obtain an adversarial catastrophic final state image. The video synthesis module is used to synthesize videos of key safety scenarios for autonomous driving by using the first frame of the extracted basic traffic video as the starting anchor point constraint and the adversarial disaster final state image as the ending anchor point constraint, and by performing time-series dynamic simulation using a video generation model.