Automatic driving image generation method for controllable injection of road traffic conditions
By injecting road traffic condition information into the Unet model and using the diffusion model to generate autonomous driving images, the problem of difficulty in collecting high-quality extreme driving scenario data in the existing technology is solved, high-quality and highly controllable data set generation is achieved, and the generalization ability of the autonomous driving system is improved.
Patent Information
- Application Number
- CN202510614950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing technology is difficult to collect high-quality and highly controllable extreme driving scenario data on actual roads, and faces challenges such as data privacy, security regulations, and label consistency.
A method of autonomous driving image generation based on diffusion model is proposed, and a trained Unet model is used to generate high-quality autonomous driving images under the guidance of road traffic conditions from a preset perspective. This method controls the image generation process by injecting traffic road conditions text information and image information, and realizes controllability of road traffic conditions.
It has realized the construction of high-quality and highly controllable synthetic data sets, enriched the real autonomous driving data sets, and improved the generalization capabilities of the autonomous driving system.
Smart Images

Figure CN120147995A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the processing or generation of autonomous driving image data, and particularly to a method for generating autonomous driving images with controllable injection of road traffic conditions. Background Art
[0002] The development of autonomous driving systems relies on large amounts of real and accurately labeled multi-modal perception data for model training, validation, and testing. These data not only include images but also cover multi-source perception channels such as semantic segmentation, depth information, optical flow, and sensor trajectories, forming the core basis in the perception-decision-control closed loop. High-quality data can significantly improve the model's understanding ability of complex traffic environments, thereby enhancing the robustness and generalization of the system, and is a key resource for promoting the continuous evolution of autonomous driving technology.
[0003] However, collecting extreme driving scenarios (such as aggressive lane changes by surrounding vehicles, rapid approach of rear vehicles, etc.) in real roads is not only costly but also difficult to achieve controllable combinations of extreme conditions, making it difficult to systematically cover more possible situations. In addition, challenges such as data privacy, safety regulations, and annotation consistency are also faced during the data collection process. Therefore, constructing a high-quality and highly controllable synthetic dataset has become an important direction for improving the generalization ability of autonomous driving systems. Summary of the Invention
[0004] To solve the above problems existing in the prior art, the present disclosure proposes a method for generating autonomous driving data in real scenarios by restricting road traffic conditions, using a diffusion model to achieve the conversion from conditional control to high-quality autonomous driving scenario images, thereby enriching the real autonomous driving dataset. The specific technical solutions are as follows.
[0005] In a first aspect, the present disclosure proposes a method for generating autonomous driving images with controllable injection of road traffic conditions. Using a trained Unet model, under the guidance of road traffic condition information at a preset perspective, an autonomous driving image is generated based on a noise image, where the road traffic condition information includes traffic condition text information and traffic condition image information; wherein: the training steps of the Unet model include: obtaining temporally continuous traffic scenes, for each traffic scene, taking real driving images obtained from multiple perspectives as a group of samples; adding noise to each image in each group of samples, and using the Unet model to denoise under the guidance of the road traffic conditions in each real driving image to restore the real driving images of each perspective.
[0006] In an implementation of the above technical solution, the guidance is achieved by injecting road traffic conditions into each downsampling feature and upsampling feature of the Unet model. The steps include: injecting the corresponding traffic condition text information into the downsampling features, and injecting both the corresponding traffic condition text information and traffic condition image information into the upsampling features.
[0007] In an implementation of the above technical solution, the steps for injecting traffic condition text information include: using the feature to be injected as the query feature of cross-attention, using the traffic condition text information as the key feature and value feature, performing attention calculation using the cross-attention mechanism, and taking the calculation result as the new downsampling feature.
[0008] In an implementation of the above technical solution, the injection of traffic condition image information is achieved by adding and fusing the road image information feature with the feature to be injected.
[0009] In an implementation of the above technical solution, the traffic condition text information includes descriptions of weather and task information.
[0010] In an implementation of the above technical solution, the traffic condition image information includes a set of road condition reference images, the camera projection of the traffic instance mask, and the camera projection of the lane line topology information.
[0011] In an implementation of the above technical solution, for each group of samples, the same traffic instance in different perspectives in the traffic condition image information is identified by the mask id.
[0012] In an implementation of the above technical solution, the size of the set of road condition reference images is Nr, where Nr is a set value, and the Nr road condition reference images are the Nr historical images that are closest in time to the current denoised image P k in time; the historical images are real driving images during training and are the autonomous driving images generated at ; the historical images are real driving images during training and are the autonomous driving images generated at , , …, at the time of inference.
[0013] In an implementation of the above technical solution, the total loss used in training is calculated as follows:
[0014] In the formula: and are weights; the reconstruction loss , is the noise added at the th step, is the original image, is the image according to the current control condition At time step Using the step noise map predicted noise; E is the expectation; conditional controllability loss , is the image generated after denoising under the control condition , is to extract the condition from the generated image , is the structural similarity metric; consistency loss , is a pre-trained network with image feature extraction capabilities is the historical road condition reference image
[0015] In a second aspect, the present disclosure provides a computer-readable storage medium storing a computer program that can be loaded and executed by a processor to perform any of the above methods
[0016] Advantageous technical effects of the present disclosure: A high-quality and highly controllable synthetic data set can be constructed to enrich the real autonomous driving data set for the learning and training of the autonomous driving system, and improve the generalization ability of the autonomous driving system BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts
[0018] Figure 1 Schematic diagram of a driving scenario generation model with high road traffic controllability, high time series, and cross-view consistency
[0019] Figure 2 Schematic diagram of the topological camera projection of the surrounding vehicle instance
[0020] Figure 3 Schematic diagram of the topological camera projection of the lane line
[0021] Figure 4 Schematic diagram of the historical road condition reference image update mechanism DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] As can be seen from the background art, collecting extreme driving scenarios in the actual road in the prior art is not only costly, but also it is very difficult to achieve a controllable combination of extreme conditions, it is difficult to systematically cover more possible situations, and there are also practical challenges such as data privacy, safety regulations, and annotation consistency during the data collection process
[0023] With the continuous development of generative artificial intelligence technology in recent years, it has become possible to generate real autonomous driving scenario data. This case proposes a method for generating an autonomous driving dataset for real scenarios based on a diffusion model, combining the ControlNet structure with a cross-attention conditional injection mechanism, and generating a rich autonomous driving dataset through conditionally controllable images and videos, thereby significantly improving the performance of the generated images in terms of road traffic controllability.
[0024] The following clearly and completely describes how the technical solution of this case is implemented. Obviously, the described implementation manners are only part of the implementation manners of this case, rather than all the implementation manners. Based on the implementation manners in this case, all other implementation manners obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this application.
[0025] (I) Image generation model (1) Diffusion model The diffusion model consists of a diffusion process and a denoising process. In the diffusion process, Gaussian noise is gradually added to the input data, and finally it is destroyed into approximately pure Gaussian noise. In the denoising process, through learnable inverse diffusion operations, it is gradually restored from the noise to the original input data.
[0026] Specifically, the diffusion model generates intermediate states by adding noise to the original image, and learns the reverse process to restore the original image during the training process. Its core modeling is as follows:
[0027] where the original image is , and the forward diffusion process perturbs it into a noisy image , is the number of steps for adding noise, is the cumulative product of the noise attenuation coefficients at each step, is the noise attenuation coefficient at the -th step, which controls the intensity of noise injection. In the reverse generation process, the model restores the original image by learning , where is the control condition, is the number of denoising steps. In terms of the specific implementation method, train the noise prediction function , and subtracting the predicted noise from the noisy image can obtain a clearer image. Therefore, the optimization objective is:
[0028] Here is random noise, is the forward noise addition process, and in this method, the control condition Refers to road traffic condition information, including traffic condition image information and traffic condition text information.
[0029] (2)Noise prediction model based on Unet model This solution uses a trained Unet model to generate an autonomous driving image from a noise image under the guidance of road traffic condition information at a preset perspective. The road condition information includes traffic condition text information and traffic condition image information.
[0030] Among them, the training steps of the Unet model include: obtaining consecutive traffic scenes in time, and for each traffic scene, using the real driving images obtained from multiple perspectives as a set of samples; adding noise to each image in each set of samples, and using the Unet model to denoise under the guidance of the road traffic conditions in each real driving image to restore the real driving images of each perspective.
[0031] The Unet model consists of a downsampling part and an upsampling part. The traffic condition text information is a low-dimensional and clearly structured description of road conditions, including descriptions of weather and task information. The traffic condition image information includes a set of road condition reference images, the camera projection of traffic instance masks, and the camera projection of lane line topology information. Inject the corresponding traffic condition text information into the downsampling features, and inject both the corresponding traffic condition text information and traffic condition image information into the upsampling features.
[0032] See Figure 1 . In the training stage, by performing multi-step noise addition operations on the real autonomous driving data set, and then extracting the latent representation through the encoder ( Figure 1 E in it), and denoising using the noise predicted by the UNet model. After training is completed, the trained Unet model predicts noise under the guidance of preset road traffic condition information, and uses the predicted noise to denoise a random noise image to generate an autonomous driving image scene that meets expectations, realizing accurate and consistent high-quality data generation.
[0033] The above-mentioned latent representation refers to the internal knowledge learned by the model from the training data in the fields of machine learning and artificial intelligence. This knowledge is stored in the architecture of the model and is specifically manifested as weights and biases. The latent representation can be regarded as a compressed form of the input data, containing the most important features required for the model to perform tasks. They are the result of the model's attempt to understand the latent structure or pattern in the data. The latent representation is crucial for the operation of machine learning models. They enable the model to generalize from the training data to unseen data, so that it can make accurate predictions or perform other tasks on new data that has not been trained.
[0034] The above-mentioned encoder compresses the noisy image data into a low-dimensional representation.
[0035] The above traffic examples include vehicles, pedestrians, cyclists, etc.
[0036] The size of the above road condition reference image set is Nr, where Nr is a set value. The Nr road condition reference images are the Nr historical images that are closest in time to the current denoised image P k in terms of time. ; The historical images are real driving images during training and are , , ……, autopilot images generated at time
[0037] The noise denoising predicted by the above UNet model is an iterative process. In this solution, the implementation model of this iterative process is encapsulated as the Figure 1 Unet denoising module in
[0038] (2) Conditional injection method (2.1) Traffic condition text information The traffic condition text information is directly embedded and injected using the cross-attention mechanism. Specifically, queries Query are provided for the latent variables generated in each downsampling and upsampling module in the Unet model: . The traffic condition text information is used as the control condition . After pre-encoding, keys Key are provided: and values Value: and , , , are the weights of the query, key, and value in the attention mechanism respectively. Then the cross-attention output is:
[0039] In the above formula, d is the dimension of the K vector.
[0040] The attention calculation result is used to replace the original latent variable and injected into the next up or downsampling module of the network to enhance the control of the conditional information over the generation process.
[0041] (2.2) Traffic condition image information For traffic condition image information, due to its high-dimensional and complex structure characteristics, it is guided and controlled through the ControlNet structure and injected into the UNet model in an additive or concatenated manner during the upsampling process.
[0042] As can be seen from the above, the traffic condition image information includes a set of road condition reference images, the camera projection of the traffic instance mask, and the camera projection of the lane line topology information.
[0043] The camera projection data processing flow of the traffic instance mask: project the 3D bounding box of the traffic instance into the coordinate system of the camera parameters of the image where it is located. Guide the model to learn about traffic instances through the traffic mask. Similarly, the camera projection of the lane line topology information projects the lane lines into the coordinate system of the camera parameters of the image where they are located.
[0044] At different perspectives, the camera projections of traffic instances are different. To enable the model to learn the knowledge of the same traffic instance from different perspectives, this case uses multiple cameras to obtain traffic instances and / or lane lines from multiple angles, and uses the same traffic instances and / or lane lines obtained from multiple angles as a set of conditions, enabling the model to learn the information of the same traffic instance and / or lane lines from different perspectives, improving the generalization ability of the model and the generation consistency of the same instance between different perspectives, and making the generated driving images closer to the real perspective.
[0045] To learn the consistency information of the same traffic instance and / or lane line information from different perspectives, after projecting the 3D bounding box into 2D, different masks are added to the projections of different instances, and a mask id uniquely related to the instance is assigned to the mask. By obtaining the masks of the 2D regions of the projections of the 3D boxes of the same instance from different camera perspectives, cross-perspective consistency is ensured. To facilitate understanding of this process, the same traffic instances from different perspectives are visualized. The same instance has the same mask identification id in different perspectives, that is, the same color in the visualization. See the Figure 2 bird's-eye view perspective in
[0046] One way to assign the mask identification id is: , where refers to the th visible instance in the current scene, refers to all unoccluded visible instances in the current scene. Perform normalization processing on for convenient model training. Determine the same traffic instance from different perspectives through the mask identification id.
[0047] The method for obtaining the mask shape area is: 3D bounding box (8 corner points) in the radar coordinate system → project onto the image → obtain the corner point convex polygon contour → intersect and clip with the image plane. The formula is as follows:
[0048] The above formula projects the 3D bounding box onto the image, where are the 8 corner points of the 3D bounding box of an instance, and respectively represent the conversion matrices from radar to camera and from camera to image, are the corner points projected onto the image.
[0049]
[0050] The above formula is for obtaining the convex polygon contour of the corner points, where represents the convex polygon area enclosed by the corner points in the image, is the algorithm for obtaining the maximum convex polygon with the specified points as vertices, are the horizontal and vertical coordinates of the corner points where the 8 corner points of the 3D bounding box of an instance are projected onto the camera image, that is , is the dimension perpendicular to the image plane, is where the depth of the third dimension coordinate in
[0051]
[0052] The above formula is for intersecting and clipping the convex polygon contour with the image plane, is the area of the entire image, ensuring a reasonable visible range of the mask. The final is Figure 2 the transparent color area shown in
[0053] Lane line projection To ensure the controllability and accuracy of important road information, such as Figure 3 the lane line projection shown in the bird's-eye view and surround camera view. Lane line projection is used to prompt the model of the key lane line information area. The acquisition method is the same as that of the 3D bounding box of the instance, but only the line segments connected by points need to be calculated without calculating the convex polygon area. The id, i.e., the value of the mask, of the same lane line is the same in multi-camera cross-views, and the colors are the same in visualization.
[0054] ControlNet inputs the above conditions into a replicated learnable Unet downsampling structure to obtain features of corresponding sizes for each layer. In the Unet network, the output sizes of the corresponding layers in the upsampling and downsampling parts are the same. During the upsampling process of the main Unet network, let the upsampling of the Unet at the layer be , and the control features of the same size obtained by the above ControlNet, then the feature fusion is expressed as:
[0055] replace , continue to complete the upsampling process of the backbone Unet network. During the upsampling process, the above-mentioned feature fusion needs to be completed for each layer.
[0056] (III) Training and Inference 3.1 Training Example Use cameras with 6 different perspectives to simultaneously obtain traffic scene images at N consecutive moments. Therefore, 6 images of the same traffic scene from 6 perspectives can be obtained simultaneously at each moment. Such a set of traffic scene images is regarded as a set of samples. As Figure 2 and Figure 3 shown in the perspective of the surround-view camera, the 6 different perspectives are the front left, the front center, and the front right; the rear left, the rear center, and the rear right.
[0057] For each image in each set of samples, obtain the road traffic conditions therein, including traffic condition text information. The traffic condition text information uses text to describe the weather and task description, and the traffic condition text information appears in the form of traffic condition prompt words. For example Figure 1 the traffic condition prompt words such as weather tasks shown in : "Sunny, vehicle turning left, vehicle in the left lane changing lanes", etc.; and existing image processing software can be used to obtain the traffic condition image information in each image. The traffic condition image information includes lane lines and traffic instances, and their projections are obtained based on the camera parameters corresponding to the image. Relative to the moment corresponding to each image , , ……, .
[0058] Perform noise addition processing on each image in each set of samples, such as using Gaussian noise. The noise addition can be multi-step noise addition, and the number of noise addition steps can be the same or different from the number of subsequent denoising steps.
[0059] Process P1: Input the image after noise addition into the Unet model. During downsampling, considering that text encoding is a compact and global semantic vector, directly embed and inject the traffic condition text information corresponding to the image through the cross-attention mechanism, interact with the feature map of UNet through the attention mechanism, and provide semantic guidance during the downsampling stage. Process layer by layer in this way until the last downsampling process is completed.
[0060] After injecting the last downsampling into the traffic condition text information, perform upsampling after convolutional processing. During upsampling, inject both traffic condition text information and traffic condition image information. Injecting the image condition during the upsampling stage can help restore the spatial structure during image restoration, such as contours, edges, etc., while the simultaneously injected text condition can better refine the image details. Finally, predict the noise through the convolution output at the top of the Unet.
[0061] Subtract the predicted noise from the noisy image to obtain the current denoised image, and use it as the new image , iteratively repeat the above process P1 until the iteration stop condition is met, and use the decoder ( Figure 1 D in) to obtain the image The image restored after denoising.
[0062] During this process P1, the set of road condition reference images remains unchanged. For the noisy image at the next moment , add the real autonomous driving image corresponding to the noisy image at time to the corresponding set of road condition reference images, and replace the oldest real autonomous driving image, so that the set of road condition reference images always retains the latest Nr images. That is to say, the generated images from the previous moment to the forward th moment are defined as a group of historical road condition reference images. Through 3D convolution, the sum in the temporal dimension is finally matched with the dimension of the upsampling part of the Unet, so that the image features generated by the model have road condition consistency in time series. Every time the model completes a step, this group of road condition reference images will be updated, push in the latest generated image (i.e., Figure 4 the new generated picture shown in is pushed in), and exit the generated image at the forward th moment (i.e., Figure 4 the old generated picture shown in exits), that is, the closest Nr historical images are , as shown in the schematic diagram Figure 4 .
[0063] 3.2 Loss function 3.2.1 Reconstruction loss (standard diffusion loss):
[0064] Where is the noise added at the th step, is the original image, is the noise predicted according to the current control condition at time step using the noise map at the th step , and the expectation of the difference between the two E This is the standard diffusion loss, denoted as .
[0065] 3.2.2 Conditional Controllability Loss:
[0066] Among them is the image generated by denoising under the control condition , is the condition extracted from the generated image . The condition can utilize a pre-trained instance segmentation model is the expected value calculated by a structural similarity metric, such as IoU, SSIM, etc., which is used to represent the difference between the instance segmentation of the image generated by the above model and the control condition, and is used to characterize the controllability of road traffic conditions.
[0067] 3.2.3 Consistency Loss (Multi-scale Consistency Loss):
[0068] Among them is a pre-trained network with the ability to extract image features, which can extract the features in the image into feature vectors and abstractly represent them is the reference image of historical road conditions, referring to the road condition reference in Figure 1 . The loss reflects the difference in features between the reference image and the generated image in the abstract space. The stronger the consistency between the reference image and the generated image, the closer the two-norms of the respectively extracted feature vectors are. Then this loss is used to characterize the feature consistency with the historical road condition reference image.
[0069] The final optimization objective is:
[0070] Among them, and are the weights of the conditional controllability and multi-scale consistency losses, is the total loss of the model.
[0071] 3.3 Inference Application Using the trained Unet model, under the guidance of the road traffic condition information from a preset perspective, a noisy image is generated into an autonomous driving image. A continuous-time autonomous driving video can be obtained through the continuously generated autonomous driving images.
[0072] During inference, the set of reference images of road traffic conditions in the traffic condition image information is the autonomous driving images generated in the most recent Nr moments, that is, if the current moment is , the generation times corresponding to the images in the traffic condition reference image set are from near to far as , , …, moments.
[0073] Whether for training or inference, the autonomous driving images generated at each moment are predicted by noise for a preset number of steps, and the predicted noise is used to denoise the noise image. Each step of noise is predicted by the Unet model, and the Unet model predicts the noise based on the current noise image and the guidance condition. After the current noise image is denoised by the predicted noise, if the stop denoising condition is not met, the current denoised image is used as the new current noise image.
[0074] (IV) Summary In this solution, by inputting noise and using the trained Unet model to denoise under the guidance of injecting various rich traffic condition information such as preset traffic condition text information and road traffic condition information, an autonomous driving image scene that meets the expectations is generated, providing data support for the training of the autonomous driving system to improve the understanding ability of the autonomous driving system for complex traffic environments.
[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the method of the present disclosure can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, in more cases for the present disclosure, software program implementation is a better implementation manner.
[0076] Although the embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, and these all belong to the scope of protection of the present disclosure.
Claims
1. A method for generating an autonomous driving image with controllable injection of road traffic conditions, characterized in that: Using the trained Unet model, under the guidance of road traffic condition information of a preset perspective, an autonomous driving image is generated based on the noise image, wherein the road traffic condition information includes traffic condition text information and traffic condition image information; The training steps of the Unet model include: obtaining temporally continuous traffic scenes, and for each traffic scene, taking real driving images obtained from multiple perspectives as a group of samples; adding noise to each image in each group of samples, and using the Unet model to remove noise under the guidance of road traffic conditions in each real driving image to restore the real driving images from each perspective.
2. The method according to claim 1, characterized in that The guiding step includes injecting road traffic conditions into each downsampled feature and upsampled feature of the Unet model: The corresponding traffic condition text information is injected into the down-sampled features, and the corresponding traffic condition text information and traffic condition image information are injected into the up-sampled features at the same time.
3. The method according to claim 2, characterized in that The steps of injecting traffic condition text information include: The features to be injected are used as query features of cross-attention, and the text information of traffic conditions is used as keyword features and value features. The cross-attention mechanism is used to perform attention calculation, and the calculation results are used as new down-sampling features.
4. The method according to claim 2, characterized in that: The injection of traffic image information is to add and fuse the road image information features with the features to be injected.
5. The method according to claim 1, characterized in that The traffic condition text information includes descriptions of weather and mission information.
6. The method according to claim 1, characterized in that The traffic condition image information includes a set of traffic condition reference images, a camera projection of a traffic instance mask, and a camera projection of lane line topology information.
7. The method according to claim 6, characterized in that For each group of samples, the same traffic instances under different perspectives in the traffic image information are identified by mask ID.
8. The method according to claim 6, characterized in that The size of the road condition reference image set is Nr, where Nr is a set value, and Nr road condition reference images are the same as the denoised image at the current moment. The Nr historical images closest in time ; The historical images are real driving images during training and real driving images during reasoning. , , …, Autonomous driving images generated moment by moment.
9. The method according to claim 1, characterized in that: The total loss used in training is calculated as follows: Where: and is the weight; Reconstruction loss , For the The noise added in step is the original image, According to the current control conditions At time step Use the Step Noise Graph The predicted noise; E is the expectation; Conditional controllability loss , For control conditions The image generated by denoising is To generate an image from Extract conditions from is the structural similarity measure; Consistency loss , It is a pre-trained network with image feature extraction capability. It is a reference image of historical road conditions.
10. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Traffic network state estimation method based on probability diffusion model
CN117690290A
Automatic driving track planning method and device based on diffusion model and electronic equipment
CN119739150A
Automatic driving scene controllable generation method based on knowledge enhancement
CN119889030A
Image style conversion method based on content and style analyzer
CN119941491A
Cited By
Driver attention prediction method, system and device, storage medium and product
CN120612675A
Driver attention prediction method, system, device, storage medium and product
CN120612675B
High-fidelity lane line image generation method without re-marking
CN121010956A
End-to-end automatic driving real confrontation scene generation and closed loop verification system and method
CN121257331A
End-to-end autonomous driving real adversarial scenario generation and closed-loop verification system and method
CN121257331B