An image generation method for autonomous driving with controllable injection of road traffic conditions
By injecting traffic condition information into the diffusion model and the Unet model in combination with the cross attention mechanism, high-quality autonomous driving images are generated, which solves the problem of high cost and controllability of extreme driving scenarios in the existing technology, and improves the data richness and generalization capabilities of the autonomous driving system.
Patent Information
- Application Number
- CN202510614950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing technology collects extreme driving scenarios on actual roads with high cost and difficulty in achieving controllable extreme conditions, and faces challenges in data privacy, security regulations and label consistency, making it difficult to build high-quality and highly controllable autonomous driving data sets.
Using the diffusion model and Unet model combined with the cross attention mechanism, high-quality autonomous driving images are generated by injecting road traffic condition information, including traffic conditions text and image information, to build a controllable synthetic data set.
It enriches the autonomous driving data set, improves the generalization ability and robustness of the system, solves the problems of high data acquisition costs and controllability, and provides a safe and efficient data generation method.
Smart Images

Figure CN120147995B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the processing or generation of autonomous driving image data, and in particular to a method for generating autonomous driving images with controllable injection of road traffic conditions. Background Art
[0002] The development of autonomous driving systems relies on large amounts of real and accurately labeled multi-modal perception data for model training, verification, and testing. These data not only include images but also cover multi-source perception channels such as semantic segmentation, depth information, optical flow, and sensor trajectories, constituting the core foundation in the perception-decision-control closed loop. High-quality data can significantly improve the model's understanding ability of complex traffic environments, thereby enhancing the system's robustness and generalization ability, and is a key resource for promoting the continuous evolution of autonomous driving technology.
[0003] However, collecting extreme driving scenarios (such as aggressive lane changes by surrounding vehicles, rapid approach of rear vehicles, etc.) in real roads is not only costly but also difficult to achieve controllable combinations of extreme conditions, making it difficult to systematically cover more possible situations. In addition, there are also practical challenges such as data privacy, safety regulations, and annotation consistency during the data collection process. Therefore, constructing a high-quality and highly controllable synthetic dataset has become an important direction for improving the generalization ability of autonomous driving systems. Summary of the Invention
[0004] To solve the above problems existing in the prior art, the present disclosure proposes a method for generating autonomous driving data in real scenarios by restricting road traffic conditions, which uses a diffusion model to achieve the conversion from conditional control to high-quality autonomous driving scenario images, thereby enriching the real autonomous driving dataset. The specific technical solutions are as follows.
[0005] In a first aspect, the present disclosure proposes a method for generating autonomous driving images with controllable injection of road traffic conditions. Using a trained Unet model, under the guidance of road traffic condition information at a preset perspective, an autonomous driving image is generated based on a noise image, where the road traffic condition information includes traffic condition text information and traffic condition image information; wherein: the training steps of the Unet model include: obtaining temporally continuous traffic scenes, and for each traffic scene, taking real driving images obtained from multiple perspectives as a set of samples; adding noise to each image in each set of samples, and using the Unet model to denoise under the guidance of the road traffic conditions in each real driving image to restore the real driving images from each perspective.
[0006] In one implementation of the above technical solution, the guidance is achieved by injecting road traffic conditions into each downsampling feature and upsampling feature of the Unet model. The steps include: injecting the corresponding traffic condition text information into the downsampling feature, and injecting both the corresponding traffic condition text information and traffic condition image information into the upsampling feature.
[0007] In one implementation of the above technical solution, the step of injecting traffic condition text information includes: using the feature to be injected as the query feature of cross-attention, using the traffic condition text information as the key feature and value feature, performing attention calculation using the cross-attention mechanism, and taking the calculation result as the new downsampling feature.
[0008] In one implementation of the above technical solution, the injection of traffic condition image information is to add and fuse the road image information feature with the feature to be injected.
[0009] In one implementation of the above technical solution, the traffic condition text information includes descriptions of weather and task information.
[0010] In one implementation of the above technical solution, the traffic condition image information includes a set of road condition reference images, the camera projection of traffic instance masks, and the camera projection of lane line topology information.
[0011] In one implementation of the above technical solution, for each group of samples, the same traffic instance in different perspectives in the traffic condition image information is identified by a mask id.
[0012] In one implementation of the above technical solution, the size of the set of road condition reference images is Nr, where Nr is a set value, and the Nr road condition reference images are the Nr historical images that are closest in time to the current denoised image P k in time. ; The historical images are real driving images during training and are self-driving images generated at , , …, moments during inference.
[0013] In one implementation of the above technical solution, the total loss used in training is calculated as follows:
[0014]
[0015] In the formula: and are weights; the reconstruction loss , is the noise added at the th step, is the original image, According to the current control conditions At the time step Using the step noise map predicted noise; E is the expectation; conditional controllability loss , is the image generated after denoising under the control condition , is to extract the conditions from the generated image , is the structural similarity metric; consistency loss , is a pre-trained network with the ability to extract image features is the historical road condition reference image
[0016] In a second aspect, the present disclosure provides a computer-readable storage medium storing a computer program that can be loaded and executed by a processor to perform any of the above methods
[0017] Beneficial technical effects of the present disclosure: A high-quality and highly controllable synthetic data set can be constructed to enrich the real autonomous driving data set for the learning and training of the autonomous driving system, and improve the generalization ability of the autonomous driving system BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings
[0019] Figure 1 Schematic diagram of a driving scenario generation model with high road traffic controllability, high time series, and cross-view consistency
[0020] Figure 2 Schematic diagram of the topological camera projection of the surrounding vehicle instances
[0021] Figure 3 Schematic diagram of the topological camera projection of the lane lines
[0022] Figure 4 Schematic diagram of the historical road condition reference image update mechanism DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] As can be seen from the background art, collecting extreme driving scenarios in the actual road in the prior art is not only costly, but also very difficult to achieve a controllable combination of extreme conditions, difficult to systematically cover more possible situations, and there are also real challenges such as data privacy, safety regulations, and annotation consistency during the data collection process
[0024] With the continuous development of generative artificial intelligence technology in recent years, it has become possible to generate real autonomous driving scenario data. This case proposes a method for generating a real-scenario autonomous driving dataset based on a diffusion model, combining the ControlNet structure with a cross-attention conditional injection mechanism to generate a rich autonomous driving dataset through conditionally controllable images and videos, thereby significantly improving the performance of the generated images in terms of road traffic controllability.
[0025] The following clearly and completely describes how the technical solution of this case is implemented. Obviously, the described implementation manners are only part of the implementation manners of this case, rather than all of them. Based on the implementation manners in this case, all other implementation manners obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this application.
[0026] (I) Image generation model
[0027] (1) Diffusion model
[0028] The diffusion model consists of a diffusion process and a denoising process. In the diffusion process, Gaussian noise is gradually added to the input data, and finally it is destroyed into approximately pure Gaussian noise. In the denoising process, through a learnable inverse diffusion operation, it is gradually restored from the noise to the original input data.
[0029] Specifically, the diffusion model generates intermediate states by adding noise to the original image, and learns the reverse process to restore the original image during the training process. Its core modeling is as follows:
[0030]
[0031] where the original image is , and the forward diffusion process perturbs it into a noisy image , is the number of steps for adding noise, is the cumulative product of the noise attenuation coefficients at each step, is the noise attenuation coefficient at the th step, which controls the intensity of noise injection. In the reverse generation process, the model restores the original image by learning , where is the control condition, is the number of denoising steps. Specifically, in the implementation method, the noise prediction function is trained. Subtracting the predicted noise from the noisy image can obtain a clearer image. Therefore, the optimization objective is:
[0032]
[0033] Here is random noise. That is the forward noise addition process, and the control conditions in this method refer to road traffic condition information, including traffic condition image information and traffic condition text information.
[0034] (2) Noise prediction model based on Unet model
[0035] This solution uses a trained Unet model to generate an autonomous driving image from a noise image under the guidance of road traffic condition information at a preset perspective. The road condition information includes traffic condition text information and traffic condition image information.
[0036] Among them, the training steps of the Unet model include: obtaining temporally continuous traffic scenes, and for each traffic scene, using the real driving images obtained from multiple perspectives as a set of samples; adding noise to each image in each set of samples, and using the Unet model to denoise under the guidance of the road traffic conditions in each real driving image to restore the real driving images of each perspective.
[0037] The Unet model consists of a downsampling part and an upsampling part. The traffic condition text information is a low-dimensional and clearly structured description of road conditions, including descriptions of weather and task information. The traffic condition image information includes a set of road condition reference images, the camera projection of traffic instance masks, and the camera projection of lane line topology information. Inject the corresponding traffic condition text information into the downsampling features, and inject both the corresponding traffic condition text information and traffic condition image information into the upsampling features.
[0038] See Figure 1 . In the training stage, by performing multi-step noise addition operations on the real autonomous driving data set, and then extracting the latent representation through the encoder ( Figure 1 E in it), and denoising using the noise predicted by the UNet model. After training, the trained Unet model predicts noise under the guidance of the preset road traffic condition information, and uses the predicted noise to denoise the random noise image to generate an autonomous driving image scene that meets the expectations, realizing accurate and consistent high-quality data generation.
[0039] The above-mentioned latent representation refers to the internal knowledge learned by a model from training data in the fields of machine learning and artificial intelligence. This knowledge is stored in the architecture of the model and is specifically manifested as weights and biases. The latent representation can be regarded as a compressed form of the input data, containing the most important features required for the model to perform tasks. They are the result of the model's attempt to understand the underlying structure or pattern in the data. The latent representation is crucial for the operation of machine learning models. It enables the model to generalize from training data to unseen data, thus enabling accurate predictions or the execution of other tasks on new data that has not been trained on.
[0040] The above-mentioned encoder compresses the noisy image data into a low-dimensional representation.
[0041] The above-mentioned traffic instances include vehicles, pedestrians, cyclists, etc.
[0042] The size of the above-mentioned set of road condition reference images is Nr, where Nr is a set value. The Nr road condition reference images are the Nr historical images that are closest in time to the current denoised image P k and are the Nr historical images closest in time ; the historical images are real driving images during training and are self-driving images generated at , , ……, during inference.
[0043] The above-mentioned noise denoising predicted by the UNet model is an iterative process. In this solution, the implementation model of this iterative process is encapsulated as the Unet denoising module in Figure 1 .
[0044] (2) Conditional injection method
[0045] (2.1) Traffic condition text information
[0046] The traffic condition text information is directly embedded and injected using the cross-attention mechanism. Specifically, queries Query are provided for the latent variables generated in each downsampling and upsampling module in the Unet model: . The traffic condition text information is used as a control condition , and after pre-encoding, keys Key are provided: and values Value: and , , , are the weights of the query, key, and value in the attention mechanism, respectively. Then the cross-attention output is:
[0047]
[0048] In the above formula, d is the dimension of the K vector.
[0049] Replace the original hidden variable with the attention calculation result Inject it into the next upsampling or downsampling module of the network to enhance the control of conditional information over the generation process.
[0050] (2.2) Traffic condition image information
[0051] For traffic condition image information, due to its characteristics of high dimension and complex structure, it is injected into the UNet model in an additive or concatenated manner through the ControlNet structure during the upsampling process for guidance and control.
[0052] As can be seen from the above, traffic condition image information includes a set of road condition reference images, the camera projection of traffic instance masks, and the camera projection of lane line topology information.
[0053] The camera projection data processing flow of traffic instance masks: project the 3D bounding box of traffic instances into the coordinate system of the camera parameters of the image where they are located. Guide the model to learn traffic instances through traffic masks. Similarly, the camera projection of lane line topology information projects lane lines into the coordinate system of the camera parameters of the image where they are located.
[0054] Under different perspectives, the camera projections of traffic instances are different. In order to enable the model to learn the knowledge of the same traffic instance under different perspectives, this case uses multiple cameras to obtain traffic instances and / or lane lines from multiple angles, and uses the same traffic instances and / or lane lines obtained from multiple angles as a set of conditions, enabling the model to learn the information of the same traffic instance and / or lane lines under different perspectives, improving the generalization ability of the model and the generation consistency of the same instance between different perspectives, and making the generated driving images closer to the real perspective.
[0055] In order to learn the consistency information of the same traffic instance and / or lane line information under different perspectives, after projecting the 3D bounding box into 2D, different masks are added to the projections of different instances, and a mask id uniquely related to the instance is assigned to the mask. By obtaining the masks of the 2D regions of the projections of the same instance's 3D box from different camera perspectives, cross-perspective consistency is ensured. For the convenience of understanding this process, the same traffic instance under different perspectives is visualized. The same instance has the same mask identification id under different perspectives, that is, the same color in the visualization. See the Figure 2 bird's-eye view perspective in.
[0056] One way to assign the mask identification id is: , where refers to the th visible instance in the current scene, Refers to all unoccluded visible instances in the current scene. For Normalization processing is performed to facilitate model training. The same traffic instance under different perspectives is determined by the mask identification ID.
[0057] The method for obtaining the mask shape area is as follows: 3D bounding box (8 corner points) in the radar coordinate system → project onto the image → obtain the convex polygon contour of the corner points → make an intersection cut with the image plane. The formula is as shown below:
[0058]
[0059] The above formula projects the 3D bounding box onto the image, where are the 8 corner points of the 3D bounding box of an instance, and represent the transformation matrices from radar to camera and from camera to image respectively, are the corner points projected onto the image.
[0060]
[0061] The above formula is for obtaining the convex polygon contour of the corner points, where represents the convex polygon area enclosed by the corner points in the image, is the algorithm for obtaining the maximum convex polygon with the specified points as vertices, are the horizontal and vertical coordinates of the corner points of the 3D bounding box of an instance projected onto the camera image, that is , is the dimension perpendicular to the image plane, is The third dimension coordinate depth in is screened to filter out the points projected behind the camera.
[0062]
[0063] The above formula is for making an intersection cut of the convex polygon contour with the image plane, is the area of the entire image, ensuring a reasonable visible range of the mask. The final is Figure 2 The transparent color area shown in, which is the camera projection mask area of the topological information of surrounding traffic instances.
[0064] Lane line projection To ensure the controllability and accuracy of important road information, such as Figure 3The lane line projection shown in the bird's-eye view and the surround camera view. The lane line projection is used to prompt the model about the key lane line information area. The acquisition method is the same as that of the 3D bounding box in the example, but only the line segments connected by points need to be calculated, without calculating the convex polygon area. The id, i.e., the value of the mask, of the same lane line is the same in multi-camera cross-views, and the colors are the same in visualization.
[0065] ControlNet inputs the above conditions into a replicated learnable Unet downsampling structure to obtain features of corresponding sizes for each layer. In the Unet network, the output sizes of the corresponding layers in the upsampling and downsampling parts are the same. During the upsampling process of the backbone Unet network, let the upsampling of UNet at the th layer be , and the control features of the same size obtained by the above ControlNet , then the feature fusion is expressed as:
[0066]
[0067] Replace , and continue to complete the upsampling process of the backbone Unet network. During the upsampling process, the above-mentioned feature fusion needs to be completed for each layer.
[0068] (III) Training and Inference
[0069] 3.1 Training Example
[0070] Use cameras with 6 different perspectives to simultaneously obtain traffic scene images at N consecutive moments. Therefore, 6 images of the same traffic scene from 6 perspectives can be obtained simultaneously at each moment, and such a set of traffic scene images is used as a set of samples. As shown in Figure 2 and Figure 3 for the surround camera view, the 6 different perspectives are left front, front, right front, left rear, rear, and right rear.
[0071] For each image in each set of samples, obtain the road traffic conditions therein, including traffic condition text information. The traffic condition text information uses text to describe the weather and task description, and the traffic condition text information appears in the form of traffic condition prompt words. For example, Figure 1 shows traffic condition prompt words such as weather tasks: "sunny, vehicle turning left, vehicle in the left lane changing lanes", etc.; and existing image processing software can be used to obtain the traffic condition image information in each image. The traffic condition image information includes lane lines and traffic instances, and their projections are obtained based on the camera parameters corresponding to the image. Relative to the moment corresponding to each image , obtain the Nr images closest to it as the traffic condition reference image set, that is, the moments corresponding to each image in this traffic condition reference image set are from far to near as , , ……, .
[0072] Add noise to each image in each group of samples, such as using Gaussian noise. The noise addition can be multi-step, and the number of noise addition steps can be the same or different from the subsequent denoising steps.
[0073] Process P1:
[0074] Input the noise-added image into the Unet model. During downsampling, considering that text encoding is a compact and global semantic vector, the corresponding traffic condition text information of the image is directly embedded and injected through the cross-attention mechanism, interacts with the feature map of UNet through the attention mechanism, and provides semantic guidance in the downsampling stage. Process layer by layer in this way until the last downsampling process is completed.
[0075] After injecting the traffic condition text information into the last downsampling, perform upsampling after convolution processing. During upsampling, inject both traffic condition text information and traffic condition image information. Injecting the image condition in the upsampling stage can help restore the spatial structure in the image restoration process, such as contours, edges, etc., while the simultaneously injected text condition can better refine the image details. Finally, the predicted noise is output through the convolution at the top of Unet.
[0076] Subtract the predicted noise from the noise-added image to obtain the current denoised image, and use it as the new image , and iteratively repeat the above Process P1 until the iteration stop condition is met. Use the decoder ( Figure 1 D in it) to obtain the image The image restored after denoising.
[0077] In this Process P1, the set of road condition reference images remains unchanged. For the noise-added image at the next moment , add the real autonomous driving image corresponding to the noise-added image at the moment to the corresponding set of road condition reference images, and replace the oldest real autonomous driving image, so that the set of road condition reference images always retains the latest Nr images. That is to say, the generated images from the previous moment to the previous th moment are defined as a group of historical road condition reference images. Finally, through 3D convolution, the sum in the temporal dimension is added to match the dimension of the upsampling part of Unet, so that the image features generated by the model have road condition consistency in time series. Each time the model completes a step, this group of road condition reference images will be updated, and the latest generated image (i.e., Figure 4 the new generated picture shown in) is pushed in, and the previous The generated image at a certain moment (i.e., Figure 4 the old generated image in exits), that is, the closest Nr historical images are Figure 4 . The schematic diagram is as shown in
[0078] 3.2 Loss Function
[0079] 3.2.1 Reconstruction Loss (Standard Diffusion Loss):
[0080]
[0081] where is the noise added at the step, is the original image, is the noise predicted according to the current control condition at time step using the noise map at the step , and the expectation of the difference between the two is the standard diffusion loss, denoted as E . .
[0082] 3.2.2 Conditional Controllability Loss:
[0083]
[0084] where is the image generated by denoising under the control condition , is the condition extracted from the generated image . The condition can use a pre-trained instance segmentation model, is the structural similarity metric, such as IoU, SSIM, etc., and the calculated expected value is used to represent the difference between the instance segmentation of the image generated by the above model and the control condition, and is used to characterize the controllability of road traffic conditions.
[0085] 3.2.3 Consistency Loss (Multi-scale Consistency Loss):
[0086]
[0087] where is a pre-trained network with the ability to extract image features, which can extract the features in the image into feature vectors and abstractly represent them, is the historical road condition reference image, referring to Figure 1The road condition reference in it. This loss reflects the difference in features between the reference image and the generated image in the abstract space. The stronger the consistency between the reference image and the generated image, the closer the two-norms of the respectively extracted feature vectors are. Then this loss is used to characterize the consistency with the features of the historical road condition reference image.
[0088] The final optimization goal is:
[0089]
[0090] Among them, and are the weights of the conditional controllability and multi-scale consistency losses, is the total loss of the model.
[0091] 3.3 Inference Application
[0092] Using the trained Unet model, under the guidance of the road traffic condition information from a preset perspective, a noise image is generated into an autonomous driving image. A continuous-time autonomous driving video can be obtained through the continuously generated autonomous driving images.
[0093] In inference, the set of road condition reference images in the traffic road condition image information is the autonomous driving images generated in the nearest Nr moments. That is, if the current moment is , then the generation moments of the images in the set of road condition reference images are from the nearest to the farthest as , , …, moment.
[0094] Regardless of training or inference, the autonomous driving images generated at each moment are denoised through noise prediction for a preset number of steps. The predicted noise is used to perform denoising processing on the noise image. Each step of noise is predicted by the Unet model, and the Unet model performs noise prediction based on the current noise image and the guidance conditions. After the current noise image is denoised by the predicted noise, if the stop denoising condition is not met, the current denoised image is used as the new current noise image.
[0095] (IV) Summary
[0096] In this solution, by inputting noise, the trained Unet model performs denoising under the guidance of injecting various rich traffic condition information such as preset traffic road condition text information and road traffic condition information, and generates an autonomous driving image scene that meets expectations, providing data support for the training of the autonomous driving system to improve the understanding ability of the autonomous driving system for complex traffic environments.
[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that the method of the present disclosure can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally speaking, functions accomplished by computer programs can easily be implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, in most cases for the present disclosure, implementation by software programs is a better embodiment.
[0098] Although the embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, those of ordinary skill in the art can also make many forms, and these all fall within the scope of protection of the present disclosure.
Claims
1. A method for generating an autonomous driving image with controllable injection of road traffic conditions, characterized in that: Using a trained Unet model, under the guidance of road traffic condition information from a preset perspective, an autonomous driving image is generated based on a noise image. The road condition information includes traffic condition text information and traffic condition image information. The traffic condition image information includes a set of road condition reference images, a camera projection of a traffic instance mask, and a camera projection of lane line topology information; Among them: the training steps of the Unet model include: obtaining temporally continuous traffic scenes, and for each traffic scene, taking the real driving images obtained from multiple perspectives as a group of samples; adding noise to each image in each group of samples, and using the Unet model to denoise under the guidance of the road traffic conditions in each real driving image to restore the real driving images of each perspective; The guidance is achieved by injecting the road traffic conditions into each downsampling feature and upsampling feature of the Unet model. The steps include: injecting the corresponding traffic condition text information into the downsampling feature, and injecting both the corresponding traffic condition text information and traffic condition image information into the upsampling feature at the same time; among them, the traffic condition text information includes descriptions of weather and task information. The steps for injecting the traffic condition text information include: taking the feature to be injected as the query feature of cross-attention, taking the traffic condition text information as the key feature and value feature, performing attention calculation using the cross-attention mechanism, and taking the calculation result as the new downsampling feature; for the traffic condition image information, it is guided and controlled through the ControlNet structure and injected into the UNet model in an additive or concatenated manner during the upsampling process.
2. The method according to claim 1, wherein For each group of samples, the same traffic instances under different perspectives in the traffic condition image information are identified by a mask id.
3. The method according to claim 1, wherein The size of the set of road condition reference images is K, where K is a set value, and the K road condition reference images are the K historical images {Pk-1, Pk-2,..., Pk-K} that are closest in time to the denoised image Pk at the current moment; the historical images are real driving images during training and autonomous driving images generated at times k-1, k-2,..., k-K during inference.
4. The method according to claim 1, characterized in that, The total loss used in training is calculated as follows: In the formula: and are weights; Reconstruction loss , For the The noise added in step is the original image, According to the existing conditions At time step Use the Step Noise Graph The predicted noise; E is the expectation; Conditional controllability loss , is the generated image of the above-mentioned generation model under conditional control, is to extract the condition from the generated image , is the structural similarity metric; Consistency loss , is a pre-trained network with image feature extraction capabilities, serves as reference information.
5. A computer-readable storage medium, characterized in that: There is a computer program stored that can be loaded and executed by a processor to perform any one of the methods as claimed in claims 1 to 4.
Citation Information
Patent Citations
Automatic driving scene controllable generation method based on knowledge enhancement
CN119889030A
Image style conversion method based on content and style analyzer
CN119941491A