Driving data generation method and device, computer equipment and storage medium
By constructing a diffusion probability model registry and text tag training diffusion model framework, combined with the ControlNet model injecting environmental features, the problem of insufficient complex environment simulation in traditional methods is solved, high-quality driving data is generated, and the adaptability and robustness of the autonomous driving system is improved.
Patent Information
- Application Number
- CN202510572958.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional driving data generation methods are difficult to simulate complex environmental changes, especially under special weather conditions, which cannot generate data that fully reflects the variability and complexity of the real world.
By constructing a diffusion probability model registry, using text tags to train the diffusion model framework, generating target driving data, and combining the ControlNet model to inject environmental features to achieve simulation of complex environments.
Diversified driving data has been generated, which improves the flexibility, accuracy and practicality of data generation, adapts to different scenarios, and enhances the environmental adaptability and robustness of the autonomous driving system.
Smart Images

Figure CN120451706A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of driving technology, and in particular to a driving data generation method, apparatus, computer equipment, storage medium, and computer program product. Background Art
[0002] With the rapid development of autonomous driving technology, data collection in complex driving environments has become increasingly important. This is especially true in special driving scenarios, such as rain, snow, fog, and other adverse weather conditions. Driving data from special scenarios often involves a comprehensive understanding of multiple dimensions, including driving behavior, traffic conditions, vehicle performance, and road conditions. However, the scarcity and high cost of this data limit the feasibility of large-scale data collection. To address the limitations of real-world data collection, driving data generation methods have emerged. By utilizing generative models to simulate driving data in special scenarios, we can provide high-quality data that can be used to train autonomous driving systems without relying on extensive physical data collection.
[0003] In traditional technology, the vehicle's driving process and sensor data are simulated by setting virtual roads, traffic signs, vehicle behavior and weather conditions.
[0004] However, rule-based simulation methods are usually limited by preset rules and scenario models, and are difficult to handle dynamically changing complex driving situations. Especially when simulating special environments such as complex weather conditions and nonlinear traffic flows, the generated data cannot fully reflect the variability and complexity of the real world. Summary of the Invention
[0005] Based on this, it is necessary to provide a driving data generation method, device, computer equipment, computer-readable storage medium and computer program product that can adapt to changes in complex environments, generate real perception data, and solve the problem that traditional methods cannot simulate complex environmental changes.
[0006] In a first aspect, the present application provides a driving data generation method, comprising:
[0007] Writing multiple diffusion probability models into a pre-built model registry to obtain a target database, and generating a dataset generation system based on the target database;
[0008] Based on the data set generation system, the input text instructions are preprocessed to obtain the text keywords corresponding to the text instructions, and the target diffusion probability model is determined based on the text keywords;
[0009] The target diffusion probability model is called in the data set generation system, and the target driving data is generated using the target diffusion probability model.
[0010] In one embodiment, before writing the plurality of diffusion probability models into a pre-built model registry to obtain a target database, the method further includes:
[0011] Acquire a target data set, the target data set including a special scene image and an original image corresponding to the special scene image; wherein the special scene image includes a text label;
[0012] Based on the target dataset, a diffusion model framework with text labels as input is constructed;
[0013] According to the diffusion model framework and the target dataset, multiple diffusion probability models guided by text labels are trained.
[0014] In one embodiment, text labels are converted into high-dimensional embedding vectors through a language-image contrast model; and a diffusion model framework with text labels as input is constructed, including:
[0015] The conditional convolutional layer is introduced into the diffusion model framework; the feature extraction of high-dimensional embedding vector is performed through the conditional convolutional layer.
[0016] In one embodiment, based on the diffusion model framework and the target dataset, multiple diffusion probability models guided by text labels are trained, including:
[0017] In the diffusion model framework, noise is added to the original image through the forward diffusion process to convert the original image into completely noisy data;
[0018] Using the special scene images and text labels corresponding to the original images as guidance, a reverse denoising process is performed on the completely noisy data to obtain a noise prediction model.
[0019] In one embodiment, the method further includes:
[0020] The loss function is used to optimize the noise prediction model to obtain the diffusion probability model.
[0021] In one embodiment, multiple diffusion probability models are written into a pre-built model registry to obtain a target database, including:
[0022] Create a database framework and build a model registry;
[0023] Based on a predefined field structure, multiple diffusion probability models are written into a model registry to obtain a target database; wherein the field structure includes a diffusion probability model identifier, a diffusion probability model classification label, and a diffusion probability model storage path.
[0024] In a second aspect, the present application further provides a driving data generating device, comprising:
[0025] A construction module is used to write multiple diffusion probability models into a pre-built model registry to obtain a target database, and generate a data set generation system based on the target database;
[0026] A determination module is used to pre-process the input text instructions based on the data set generation system, obtain text keywords corresponding to the text instructions, and determine the target diffusion probability model based on the text keywords;
[0027] The generation module is used to call the target diffusion probability model in the data set generation system and generate target driving data using the target diffusion probability model.
[0028] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0029] Writing multiple diffusion probability models into a pre-built model registry to obtain a target database, and generating a dataset generation system based on the target database;
[0030] Based on the data set generation system, the input text instructions are preprocessed to obtain the text keywords corresponding to the text instructions, and the target diffusion probability model is determined based on the text keywords;
[0031] The target diffusion probability model is called in the data set generation system, and the target driving data is generated using the target diffusion probability model.
[0032] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:
[0033] Writing multiple diffusion probability models into a pre-built model registry to obtain a target database, and generating a dataset generation system based on the target database;
[0034] Based on the data set generation system, the input text instructions are preprocessed to obtain the text keywords corresponding to the text instructions, and the target diffusion probability model is determined based on the text keywords;
[0035] The target diffusion probability model is called in the data set generation system, and the target driving data is generated using the target diffusion probability model.
[0036] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0037] Writing multiple diffusion probability models into a pre-built model registry to obtain a target database, and generating a dataset generation system based on the target database;
[0038] Based on the data set generation system, the input text instructions are preprocessed to obtain the text keywords corresponding to the text instructions, and the target diffusion probability model is determined based on the text keywords;
[0039] The target diffusion probability model is called in the data set generation system, and the target driving data is generated using the target diffusion probability model.
[0040] The aforementioned driving data generation method, apparatus, computer device, storage medium, and computer program product write multiple diffusion probability models into a pre-built model registry to obtain a target database, and then generate a dataset generation system based on the target database. The dataset generation system preprocesses input text instructions to obtain text keywords corresponding to the text instructions, and determines a target diffusion probability model based on the text keywords. The target diffusion probability model is then called within the dataset generation system and used to generate target driving data. This application utilizes the diffusion probability model to generate driving data in complex environments, such as diverse weather and dynamic traffic, thereby increasing the diversity of driving data generation. Furthermore, the diffusion probability model is used to generate dynamic driving data that is closer to actual driving conditions, thereby improving the practicality and authenticity of driving data generation. Furthermore, the dataset generation system is able to select an appropriate target diffusion probability model based on different scenario requirements, greatly enhancing the flexibility and scalability of driving data generation. By selecting the most appropriate target diffusion probability model for data generation, the data's relevance and accuracy are also improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 A diagram illustrating an application environment of a driving data generating method according to an embodiment;
[0043] Figure 2 1 is a flow chart of a method for generating driving data in one embodiment;
[0044] Figure 3 A schematic diagram of a process for constructing a diffusion probability model in one embodiment;
[0045] Figure 4 is a structural block diagram of a driving data generating device in one embodiment;
[0046] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0048] The driving data generation method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process. The data storage system can be integrated with server 104, or located in the cloud or on other network servers. Terminal 102 sends a driving data generation request to server 104. Server 104 receives the driving data generation request and writes multiple diffusion probability models into a pre-built model registry to obtain a target database. Based on the target database, a dataset generation system is generated. The dataset generation system pre-processes the input text instruction to obtain text keywords corresponding to the text instruction and determines a target diffusion probability model based on the text keywords. The target diffusion probability model is then called in the dataset generation system and used to generate target driving data. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices can include smart watches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0049] In an exemplary embodiment, Figure 2 As shown, a driving data generation method is provided, which is applied to Figure 1 The server in FIG. 1 is used as an example to illustrate the method, including the following steps 202 to 206. Among them:
[0050] Step 202 : Write multiple diffusion probability models into a pre-built model registry to obtain a target database, and generate a data set generation system based on the target database.
[0051] Exemplarily, multiple diffusion probability models obtained through training are written into a pre-built model registry according to a preset field structure (e.g., diffusion probability model unique identifier, diffusion probability model classification label, etc.), thereby forming a target database. This target database is used to uniformly manage and control the call of multiple diffusion probability models. Furthermore, based on this target database, a dataset generation system is constructed. This system can automatically retrieve and call matching diffusion probability models based on user-entered text instructions, generating synthetic image data (target driving data) corresponding to the text instructions, thereby achieving automated generation of image datasets for various extreme weather conditions.
[0052] Optionally, the entire process can be integrated through FastAPI, integrating the model routing, target database, and cache layer to form a dataset generation system. FastAPI is a web service interface framework that supports asynchronous calls and fast routing matching. Model routing is used to encapsulate and schedule multiple diffusion probability models, providing a unified call entry point for different diffusion probability models. Optionally, model routing can automatically identify the required diffusion probability model based on input text instructions, improving the flexibility of model management and service interaction. The cache layer is used to cache intermediate results and hot-load diffusion probability model states during the diffusion probability model identification process, reducing response time for repeated calls and improving system operational efficiency. Furthermore, the dataset generation system includes a service initialization module, a model loading module, and a routing controller. The service initialization module is used to complete basic parameter configuration and lifecycle management when the dataset generation system starts. The model loading module is used to load diffusion probability models from the target database on demand, providing the required diffusion probability models for the model routing. The routing controller is used to generate the final results.
[0053] Step 204 : pre-processing the input text instruction based on the data set generation system to obtain text keywords corresponding to the text instruction, and determining a target diffusion probability model based on the text keywords.
[0054] In the dataset generation system, user-entered text commands are normalized and keyword extracted. Exemplarily, normalization includes preprocessing the text commands, including character set unification, converting full-width characters to half-width characters, standardizing English uppercase and lowercase characters, and removing interfering symbols, noise characters, and meaningless punctuation, thereby improving the accuracy of subsequent semantic recognition. After normalization, the text commands are segmented and tagged with parts of speech. Combined with rule-based intent recognition methods, semantic analysis is performed on the normalized text commands to extract the corresponding text keywords. These keywords include, but are not limited to, weather type keywords (e.g., "smog," "rainstorm," "dust"), visual attribute keywords (e.g., "blur," "clarity," "occlusion"), and numerical parameter keywords (e.g., "50%," "medium concentration"), among other semantically specific terms. Keyword extraction enables accurate recognition of user intent, providing a reliable semantic foundation for subsequent matching with the diffusion probability model, thereby enabling automatic scheduling of diffusion probability model calls and precise configuration of image generation tasks.
[0055] Step 206 : calling the target diffusion probability model in the data set generation system, and using the target diffusion probability model to generate target driving data.
[0056] In the dataset generation system, based on the aforementioned text keyword recognition and matching results with multiple diffusion probability models, the target diffusion probability model corresponding to the target keyword is automatically invoked. Optionally, during this invocation, the dataset generation system loads the target diffusion probability model from the target database into runtime memory via a model loading module. Combining this with pre-set baseline clear image data and identified weather labels as generation input, the target diffusion probability model generates target driving data that conforms to specific weather characteristics (such as rain, snow, fog, etc.) and visual blur levels. This target driving data includes synthetic image sequences from high-speed driving perspectives under various extreme weather conditions, enhancing the model's training sample coverage in scarce scenarios and improving the environmental adaptability and robustness of the autonomous driving system.
[0057] In the aforementioned driving data generation method, multiple diffusion probability models are written into a uniformly structured model registry to construct a target database. This enables centralized management of different categories of diffusion probability models, facilitating subsequent on-demand retrieval and rapid loading, improving the organizational efficiency and scalability of diffusion probability model resources, and supporting automatic switching across multiple scenarios. Text instructions are preprocessed and keyword extracted, enabling the dataset generation system to identify key intent (such as weather type and degree of ambiguity) from natural language and accurately match the most appropriate diffusion probability model accordingly. This achieves automatic semantic alignment between input text instructions and diffusion probability model calls, significantly enhancing the intelligence and interactive friendliness of the dataset generation system. By invoking diffusion probability models that match the text keywords of the text instructions, driving data for specified weather conditions or complex environments can be automatically generated based on actual needs, improving the adaptability and reliability of driving data generation.
[0058] In the previous exemplary embodiment, multiple diffusion probability models are written into a pre-built model registry to obtain a target database, including: creating a database framework and building a model registry; based on a pre-defined field structure, multiple diffusion probability models are written into the model registry to obtain the target database; wherein the field structure includes a diffusion probability model identifier, a diffusion probability model classification label, and a diffusion probability model storage path.
[0059] For example, creating a SQLite database and establishing a model registry (models) provides an efficient solution for managing and calling multiple diffusion probability models. Multiple field structures are defined in the SQLite database, including diffusion probability model identifiers, diffusion probability model classification labels, and diffusion probability model storage paths. The field structures can also include the deep learning framework used, GPU memory requirements, and other extended parameters. These field structures enable detailed recording of relevant information for each diffusion probability model, ensuring that the appropriate model can be quickly located and selected in actual operations. Multiple pre-trained diffusion probability models are written to the model registry to obtain a target database. When the corresponding diffusion probability model needs to be retrieved for operation, the dataset generation system connects to the target database and, based on text keywords, uses multi-level matching logic to query and obtain the target diffusion probability model. This allows the selection of the optimal model that best meets current needs, ensuring efficient utilization of the diffusion probability model and optimizing the performance of the dataset generation system.
[0060] Optionally, to further optimize the use of dataset generation system resources, a dynamic memory management and model loading optimization mechanism has been introduced. Specifically, the dataset generation system automatically detects and unloads idle diffusion probability models that have not been used for a long time, freeing up GPU memory resources. When a diffusion probability model needs to be loaded, a model loading method based on CPU offloading technology and a memory-efficient attention mechanism is adopted. Through intelligent resource scheduling and memory management technology, the required diffusion probability model is efficiently loaded, reducing time delays and memory usage during the loading process, thereby improving the operating efficiency and responsiveness of the entire dataset generation system.
[0061] In the above-described embodiment, ControlNet conditional parameter injection technology is applied to achieve more accurate image (target driving data) generation by incorporating external control parameters into the diffusion probability model generation process. Specifically, a pre-trained ControlNet model is first loaded. This ControlNet model has been trained to handle image generation tasks under various edge conditions and adverse weather scenarios, and is capable of adapting to changes in various environmental conditions. When loading the ControlNet model, the data type used is specified as half-precision floating-point numbers to improve computational efficiency and reduce resource consumption. The loaded ControlNet model is then passed to the diffusion probability model generation pipeline, and the two work together. During the target driving data generation process, the ControlNet model performs object detection on the input image based on external control conditions and generates a conditional map containing environmental feature information. This conditional map includes, but is not limited to, environmental feature information such as sky color, image blur, and weather conditions. By injecting this environmental feature information into the diffusion probability model, the details and style of the image generated by the diffusion model are controlled. The ControlNet model provides fine-grained control of environmental features during the diffusion probability model generation process, ensuring that the generated image is not simply based on a single change in the input image, but dynamically adapts to external environmental conditions. For example, in scenarios like heavy rain or haze, ControlNet can provide a detailed description of weather and environmental conditions. By controlling the generated intensity, it can adjust the diffusion model's performance on image details such as blur and color saturation. In this way, the ControlNet model significantly enhances the performance of the diffusion probability model in these specific scenarios, generating images that better reflect adverse weather conditions. Ultimately, by incorporating external environmental information, the diffusion probability model can generate images that fully reflect the specified weather conditions, providing high-quality and accurate driving data for subsequent dataset generation.
[0062] In an exemplary embodiment, Figure 3As shown, before writing multiple diffusion probability models into a pre-built model registry to obtain a target database, constructing a diffusion probability model includes steps 302 to 306.
[0063] Step 302: Acquire a target data set, where the target data set includes special scene images and original images corresponding to the special scene images; wherein the special scene images include text labels.
[0064] Step 304: Based on the target dataset, a diffusion model framework is constructed with the text label as input.
[0065] Step 306 : Based on the diffusion model framework and the target dataset, multiple diffusion probability models guided by text labels are trained.
[0066] Since training a diffusion probability model requires both original images, corresponding special scene images, and corresponding text labels for text training, the NH-Haze and RainSnow datasets were selected. These datasets contain nearly a thousand pairs of images of extreme weather scenes of varying concentrations and their corresponding clear images, meeting the required training dataset criteria. Special scene images refer to images from weather conditions such as haze, heavy rain, and snowstorms. After determining the target dataset, the images in the target dataset were cropped and resized to 1600 x 900 pixels to ensure that all input images had the same size, enabling better training of the diffusion probability model. Text labels describe the content of the special scene images, specifically including the type of weather condition presented, such as "haze," "heavy rain," or "sunny." Manually annotating text labels provides clear guidance for subsequent training of the diffusion probability model, enabling the model to understand the relationship between image content and weather conditions during learning.
[0067] Based on the target dataset described above, a diffusion model framework was constructed that takes text labels as input. This diffusion model framework transforms text labels into conditional information, guiding the diffusion probability model to generate images that meet specific scene or environmental conditions during the image generation process. Specifically, the text labels, as input conditions, are converted through specific preprocessing and encoding methods into a format that the diffusion probability model can understand and process. This guides the diffusion probability model to generate images that meet the predefined conditions based on the text labels.
[0068] Based on the aforementioned diffusion model framework and the target dataset, training is performed to generate multiple diffusion probability models guided by text labels. During the diffusion probability training process, images that match specific scene descriptions are generated based on the text labels. For example, during training, if the text label indicates "smog," the diffusion probability model adjusts image generation properties such as blur and color saturation to produce images consistent with smog conditions. Through this process, multiple diffusion probability models are trained.
[0069] In this embodiment, by constructing a target dataset containing special scene images and corresponding original images, the diversity of the diffusion probability model training data is ensured to be consistent with real-world scenarios. Text labels provide clear semantic information for each image, significantly improving the accuracy of data annotation and facilitating the construction of a more accurate and flexible diffusion probability model. Furthermore, by combining text labels with special scene images as input, the semantic understanding capability of the diffusion probability model is effectively enhanced, enabling the diffusion probability model to accurately generate images that meet specific conditions based on the text labels. This allows the diffusion probability model to adapt to changes in complex environments, resolving the problem of traditional methods being unable to simulate complex environmental changes.
[0070] In an exemplary embodiment, text labels are converted into high-dimensional embedding vectors through a language-image contrast model; a diffusion model framework with text labels as input is constructed, including: introducing a conditional convolutional layer into the diffusion model framework; and performing feature extraction on the high-dimensional embedding vector through the conditional convolutional layer.
[0071] Among them, the language-image contrast model is the CLIP (Contrastive Language-Image Pretraining) model.
[0072] The above-mentioned diffusion model framework can accept text labels as input. For example, the input text labels are mapped to a high-dimensional vector space through the CLIP model to generate embedding vectors corresponding to the semantics of the text labels. These embedding vectors capture key information in the text labels, such as scene features such as weather, time, and traffic conditions. A conditional convolution layer is added to the diffusion model framework as a conditional network. This conditional convolution layer extracts features from the embedding vectors and effectively combines the input high-dimensional embedding vectors with the image generation process of the diffusion probability model. Through the conditional convolution layer, the diffusion probability model can effectively utilize the text label information to guide each stage of the diffusion probability model's image generation process to ensure that the generated target driving data conforms to the given text labels.
[0073] In an exemplary embodiment, based on a diffusion model framework and a target data set, multiple diffusion probability models guided by text labels are trained, including: within the diffusion model framework, adding noise to the original image through a forward diffusion process to convert the original image into completely noisy data; using the special scene image and text label corresponding to the original image as a guide, performing a reverse denoising process on the completely noisy data to obtain a noise prediction model.
[0074] During the training process of the diffusion probability model, noise is gradually added to the original image through the forward diffusion process, and the original image is gradually transformed into a noise image. For example, the training data (Original image) In the form of a Markov chain, according to the pre-set noise scheduling sequence , gradually add a small amount of Gaussian noise , until completely noisy data is formed The above noise scheduling sequence It is used to control the degree of noise attenuation in each step. The above process can be expressed by the following formula:
[0075]
[0076] This forward diffusion process provides the basis for subsequent denoising and image restoration, and can effectively simulate the degradation characteristics of images in complex environments.
[0077] Using the special scene images corresponding to the original images and their associated text labels as guidance, the diffusion probability model is trained through a reverse denoising process. This process aims to recover a qualified image from the completely noisy data formed after forward diffusion, gradually learning the data distribution patterns under specified conditions and generating a noise prediction model. Optionally, the specified conditions include text labels and the noise difference between the original image and its corresponding special scene image.
[0078] In the reverse denoising process, text labels not only provide additional semantic information for the model, but also guide the diffusion probability model to generate images that conform to the text description according to specific scene requirements. For example, when the text label describes a certain weather condition (such as "smog" or "heavy rain"), the diffusion probability model will adjust the features of the generated image according to these descriptions during the denoising process to ensure that the final image can accurately reflect the environmental conditions of the text label. The goal of this reverse denoising process is to predict the noise at the current step, denoted as By minimizing the difference between the predicted noise and the actual noise, the noise distribution is learned, and the noise prediction at each step is gradually optimized with the help of the loss function, and the image is updated according to the predicted noise. Until a clear image is restored, the formula is as follows:
[0079]
[0080] The reverse denoising process generates a specific noise prediction model by learning the noise differences between special scene images in the target dataset and the original images corresponding to the special scene images, and combines the scene conditions provided by the text labels, further improving the generation ability of the diffusion probability model under complex weather conditions.
[0081] In the previous exemplary embodiment, the method further includes: optimizing the noise prediction model using a loss function to obtain a diffusion probability model.
[0082] During the diffusion probability model training process, the noise prediction model is optimized by constructing a loss function to obtain a more accurate and generalizable diffusion probability model. With prediction noise The covariance between them is determined to quantify the accuracy of the noise prediction, and the formula is:
[0083]
[0084] In order to prevent overfitting of the diffusion probability model and improve the quality of the generated images (target dataset), L2 regularization technology is also used to constrain the diffusion probability model, further improving the generalization ability and generation effect of the diffusion probability model.
[0085] In this embodiment, the original image is converted into completely noisy data through a forward diffusion process, providing the diffusion probability model with a complete noise evolution path, helping to capture the characteristic distribution of images under different levels of degradation. In the reverse denoising process, text labels and special scene images are introduced as conditional inputs, improving the diffusion probability model's ability to understand special scenes and enhancing the match between the generated image and the text instruction input description. This enables the diffusion probability model to adapt to changes in complex environments and generate realistic perception data. The noise prediction model is optimized through a loss function, effectively improving the generalization ability and generation accuracy of the diffusion probability model under different climate or special scene conditions.
[0086] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0087] Based on the same inventive concept, embodiments of the present application also provide a driving data generation device for implementing the aforementioned driving data generation method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more driving data generation device embodiments provided below can be found in the above-described limitations of the driving data generation method and will not be further elaborated here.
[0088] In an exemplary embodiment, Figure 4 As shown, a driving data generating device is provided, including: a construction module 402, a determination module 404 and a generation module 406, wherein:
[0089] A construction module 402 is used to write multiple diffusion probability models into a pre-built model registry to obtain a target database, and generate a data set generation system based on the target database;
[0090] The determination module 404 is configured to pre-process the input text instruction based on the data set generation system to obtain text keywords corresponding to the text instruction, and determine a target diffusion probability model based on the text keywords;
[0091] The generating module 406 is configured to call the target diffusion probability model in the data set generating system and generate target driving data using the target diffusion probability model.
[0092] In an exemplary embodiment, the driving data generating device further includes:
[0093] The diffusion model generation module is used to obtain a target dataset, which includes special scene images and original images corresponding to the special scene images; wherein the special scene images include text labels; based on the target dataset, a diffusion model framework with text labels as input is constructed; based on the diffusion model framework and the target dataset, multiple diffusion probability models guided by text labels are trained.
[0094] In an exemplary embodiment, text labels are converted into high-dimensional embedding vectors through a language-image contrast model; the diffusion model generation module is further used to introduce a conditional convolutional layer into the diffusion model framework; and feature extraction is performed on the high-dimensional embedding vector through the conditional convolutional layer.
[0095] In an exemplary embodiment, the diffusion model generation module is also used to add noise to the original image through a forward diffusion process under the diffusion model framework to convert the original image into completely noisy data; using the special scene image and text label corresponding to the original image as a guide, a reverse denoising process is performed on the completely noisy data to obtain a noise prediction model.
[0096] In an exemplary embodiment, the diffusion model generation module is further configured to optimize the noise prediction model using a loss function to obtain a diffusion probability model.
[0097] In an exemplary embodiment, the construction module 402 is also used to create a database framework and build a model registry; based on a predefined field structure, multiple diffusion probability models are written into the model registry to obtain a target database; wherein the field structure includes a diffusion probability model identifier, a diffusion probability model classification label, and a diffusion probability model storage path.
[0098] Each module in the aforementioned driving data generation device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor within a computer device in the form of hardware, or may be stored in a memory within the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0099] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store target data sets. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a driving data generation method is implemented.
[0100] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0101] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0102] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0103] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0104] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0105] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0106] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0107] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A driving data generation method, characterized in that: The method comprises: Writing multiple diffusion probability models into a pre-built model registry to obtain a target database, and generating a data set generation system based on the target database; Preprocessing the input text instruction based on the data set generation system to obtain text keywords corresponding to the text instruction, and determining a target diffusion probability model based on the text keywords; The target diffusion probability model is called in the data set generation system, and the target driving data is generated using the target diffusion probability model.
2. The method according to claim 1, characterized in that Before writing the plurality of diffusion probability models into a pre-built model registry to obtain a target database, the method further includes: Acquire a target data set, the target data set including a special scene image and an original image corresponding to the special scene image; wherein the special scene image includes a text label; Based on the target data set, constructing a diffusion model framework with the text label as input; According to the diffusion model framework and the target data set, a plurality of diffusion probability models guided by the text labels are trained.
3. The method according to claim 2, characterized in that The text labels are converted into high-dimensional embedding vectors through a language-image contrast model; and the construction of a diffusion model framework with the text labels as input includes: A conditional convolution layer is introduced into the diffusion model framework; and features are extracted from the high-dimensional embedding vector through the conditional convolution layer.
4. The method according to claim 2, characterized in that The step of training a plurality of diffusion probability models guided by the text labels based on the diffusion model framework and the target data set includes: Adding noise to the original image through a forward diffusion process under the diffusion model framework to convert the original image into completely noisy data; Using the special scene image and text label corresponding to the original image as a guide, a reverse denoising process is performed on the completely noisy data to obtain a noise prediction model.
5. The method according to claim 4, characterized in that The method further comprises: The noise prediction model is optimized using a loss function to obtain the diffusion probability model.
6. The method according to claim 1, characterized in that Multiple diffusion probability models are written into a pre-built model registry to obtain the target database, including: Create a database framework and build a model registry; Based on a predefined field structure, a plurality of the diffusion probability models are written into the model registry to obtain a target database; wherein the field structure includes a diffusion probability model identifier, a diffusion probability model classification label and a diffusion probability model storage path.
7. A driving data generating device, characterized in that: The device comprises: A construction module is used to write multiple diffusion probability models into a pre-built model registry to obtain a target database, and generate a data set generation system based on the target database; a determination module, configured to pre-process the input text instruction based on the data set generation system, obtain text keywords corresponding to the text instruction, and determine a target diffusion probability model based on the text keywords; A generation module is used to call the target diffusion probability model in the data set generation system and generate target driving data using the target diffusion probability model.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.