Image generation method and device, electronic equipment and storage medium
By acquiring map topology information, traffic flow simulation configuration data, and user semantic description information, and using a large language model and ControlNet to generate autonomous driving model images, the problem of scarce long-tail scene data and poor semantic consistency in autonomous driving systems is solved, and image generation with logical consistency and controllability is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-07
AI Technical Summary
Autonomous driving systems face challenges such as scarce data, insufficient controllability in generating data, and poor semantic consistency in long-tail scenarios. Existing technologies struggle to efficiently generate diverse long-tail scenario data that is realistic, logically consistent, and highly controllable.
By acquiring map topology information, traffic flow simulation configuration data, and user semantic description information, we use a large language model to generate autonomous driving model images. Combined with ControlNet's spatial adapter and channel adapter, we perform structural and semantically constrained image generation to achieve road image generation for various special scenarios.
The system generates autonomous driving model images with strong logical consistency and high controllability, solving the problems of scarce data and poor semantic consistency in long-tail scenarios, and meeting the training requirements of autonomous driving systems.
Smart Images

Figure CN121810863A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and more particularly to an image generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] In the field of autonomous driving, to ensure safe and stable driving in real-world road environments, autonomous driving systems need to accurately perceive and understand various complex and ever-changing traffic scenarios. To achieve this goal, the industry generally relies on large-scale, diverse training datasets to train autonomous driving perception, prediction, and decision-making models. Currently, training data is mainly collected through methods such as fleet road testing and simulation playback, but these methods are costly and time-consuming. Furthermore, data for long-tail scenarios such as extreme weather, special road conditions, and rare accidents is severely lacking.
[0003] To address this issue, existing technologies typically employ data generation techniques. However, the realism of data synthesized using traditional 3D simulation or image generation methods is limited, exhibiting discrepancies with real data in terms of texture, lighting, and material properties. Furthermore, it is difficult to achieve high-fidelity and controllable synthesis of multi-dimensional factors such as climate, lighting, and urban style. Additionally, most image generation methods cannot automatically and accurately update corresponding bounding boxes, semantic segmentation, or instance segmentation annotations after modifying the image's appearance, increasing the cost of manual correction. Therefore, how to efficiently generate realistic, logically consistent, highly controllable data covering various long-tail scenarios without relying on large-scale field data collection has become a critical technical problem urgently needing to be solved in this field. Summary of the Invention
[0004] This invention provides an image generation method, apparatus, electronic device, and storage medium to address the problems of scarce data, insufficient controllability of generation, and poor semantic consistency in the long-tail scenario of autonomous driving in the prior art.
[0005] According to one aspect of the present invention, an image generation method is provided, the method comprising:
[0006] Acquire map topology information, traffic flow simulation configuration data, and user semantic description information;
[0007] The system generates autonomous driving model images corresponding to the user semantic description information, the traffic flow simulation configuration data, and the map topology information based on the large language model.
[0008] According to another aspect of the present invention, an image generation apparatus is provided, the apparatus comprising:
[0009] The information acquisition module is used to acquire map topology information, traffic flow simulation configuration data, and user semantic description information;
[0010] The image generation module is used to generate autonomous driving model images corresponding to the user semantic description information, the traffic flow simulation configuration data, and the map topology information based on the large language model.
[0011] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0012] At least one processor; and
[0013] A memory communicatively connected to the at least one processor; wherein,
[0014] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image generation method according to any embodiment of the present invention.
[0015] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute an image generation method according to any embodiment of the present invention.
[0016] The technical solution of this invention acquires map topology information, traffic flow simulation configuration data, and user semantic description information, and then generates autonomous driving model images corresponding to the user semantic description information, traffic flow simulation configuration data, and map topology information based on a large language model. This technical solution, by acquiring map topology information, traffic flow simulation configuration data, and user semantic description information, can maintain the consistency of the geometric positions, scales, and road structure logic of traffic participants. By generating autonomous driving model images corresponding to the user semantic description information, traffic flow simulation configuration data, and map topology information based on a large language model, it can generate road images for various special scenarios, solving the problems of scarce long-tail scenario data, insufficient controllability in generation, and poor semantic consistency in existing technologies for autonomous driving.
[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1This is a flowchart of an image generation method provided according to Embodiment 1 of the present invention;
[0020] Figure 2 This is a flowchart of an image generation method provided according to Embodiment 2 of the present invention;
[0021] Figure 3 This is a flowchart of an image generation method provided according to Embodiment 3 of the present invention;
[0022] Figure 4 This is a schematic diagram of an image generation device according to Embodiment 4 of the present invention;
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the image generation method of Embodiment 5 of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] Example 1
[0027] Figure 1 This is a flowchart of an image generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to the generation of long-tail images for autonomous driving. The method can be executed by an image generation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0028] S110: Obtain map topology information, traffic flow simulation configuration data, and user semantic description information.
[0029] The map topology information refers to the set of information used to describe the road network structure and the connections between elements. Map topology information can include road geometry, number of lanes, lane connections, and the locations of traffic signs and traffic lights. Traffic flow simulation configuration data refers to the set of configuration data used to generate the simulated traffic environment. Traffic flow simulation configuration data can include vehicle information, traffic participant information, static traffic element information, and traffic light information. Vehicle information refers to the size, model, driving style, average speed, origin, and destination of the vehicle executing the autonomous driving actions. Traffic participant information refers to the type, quantity, behavior patterns, and movement trajectories of other dynamic entities besides the vehicle. The types of dynamic entities can include surrounding vehicles, pedestrians, and non-motorized vehicles. Static traffic element information refers to the information of fixed, non-dynamic entities in the simulated road network structure. Static traffic element information can include the location and status of traffic signs, streetlights, guardrails, traffic cones, and fire hydrants. Traffic light information refers to the color, symbol, and time status of traffic lights at various intersections and roadside locations in the simulated road network structure. User semantic description information refers to the descriptive information about the characteristics of the scene from which the image is generated, input in the form of natural language or structured text. User semantic description information may include information such as time, weather, lighting, road surface material, and city style.
[0030] Specifically, the traffic simulation platform acquires map topology information, traffic flow simulation configuration data, and user semantic description information through data interfaces and remote transmission.
[0031] S120. Generate autonomous driving model images corresponding to user semantic description information, traffic flow simulation configuration data, and map topology information based on the large language model.
[0032] Among them, large language models refer to artificial intelligence models trained on a large amount of text and scene data, possessing cross-modal understanding and generation capabilities. Large language models can generate model images that meet the requirements based on input data, configuration information, etc. Autonomous driving model images refer to model images generated based on multi-source information fusion for specific autonomous driving scenarios.
[0033] Specifically, the user's semantic description information, traffic flow simulation configuration data, and map topology information are input into the trained large language model to generate an autonomous driving model image corresponding to the user's semantic description information, traffic flow simulation configuration data, and map topology information.
[0034] The technical solution of this invention acquires map topology information, traffic flow simulation configuration data, and user semantic description information, and then generates autonomous driving model images corresponding to the user semantic description information, traffic flow simulation configuration data, and map topology information based on a large language model. This technical solution, by acquiring map topology information, traffic flow simulation configuration data, and user semantic description information, can maintain the consistency of the geometric positions, scales, and road structure logic of traffic participants. By generating autonomous driving model images corresponding to the user semantic description information, traffic flow simulation configuration data, and map topology information based on a large language model, it can generate road images for various special scenarios, solving the problems of scarce long-tail scenario data, insufficient controllability in generation, and poor semantic consistency in existing technologies for autonomous driving.
[0035] Example 2
[0036] Figure 2 This is a flowchart of an image generation method provided in Embodiment 2 of the present invention. This embodiment further refines the above embodiment:
[0037] like Figure 2 As shown, the method includes:
[0038] S210. Obtain the uploaded map data and generate a simulated road network structure for the map data.
[0039] The uploaded map data refers to the original map data file describing the real or virtual road environment. The simulated road network structure refers to a digital road model extracted and reconstructed from the original map data, containing complete topological connections. The simulated road network structure may include at least one of lanes, lane lines, road boundaries, and drivable areas.
[0040] Specifically, the traffic simulation platform acquires uploaded original map data files describing the real or virtual road environment through data interfaces, remote transmission, and other means, and generates a simulated road network structure from the acquired map data. The simulated road network structure includes at least one of lanes, lane lines, road boundaries, and drivable areas.
[0041] S220. Obtain the user-configured traffic flow simulation data.
[0042] Traffic flow simulation configuration data refers to the set of configuration data used to generate the simulated traffic environment. This data can include vehicle information, traffic participant information, static traffic element information, and traffic light information. Vehicle information refers to the size, model, driving style, average speed, origin, and destination of the vehicle executing the autonomous driving actions. Traffic participant information refers to the type, quantity, behavior patterns, and movement trajectories of other dynamic entities besides the vehicle. Dynamic entity types can include surrounding vehicles, pedestrians, and non-motorized vehicles. Static traffic element information refers to the information of fixed, non-dynamic entities in the simulated road network structure. This information can include the location and status of traffic signs, streetlights, guardrails, traffic cones, and fire hydrants. Traffic light information refers to the color, symbol, and timing status of traffic lights at various intersections and roadside locations in the simulated road network structure.
[0043] Specifically, the traffic simulation platform obtains real-time or pre-configured traffic flow simulation configuration data from users through data interfaces, remote transmission, and other means. The traffic flow simulation configuration data includes vehicle information, traffic participant information, traffic static element information, and traffic signal information.
[0044] S230, Collect user semantic description information.
[0045] Among them, user semantic description information refers to the descriptive information of the scene to which the image is generated, which is input in the form of natural language or structured text. User semantic description information may include information such as time, weather, lighting, road surface material and city style.
[0046] Specifically, users generate corresponding semantic description statements based on the actual autonomous driving model training scenario data required. The traffic simulation platform collects the user's semantic description information, which can be in the form of natural language or structured text.
[0047] S240. Input the user semantic description information, traffic flow simulation configuration data and map topology information into the input layer of the large language model, and obtain the structural constraint vector and semantic constraint vector generated by the input layer.
[0048] The input layer refers to the part of the large language model responsible for receiving and parsing input data. This input layer can be a dedicated encoder for different data types. The structural constraint vector is a numerical vector representing the structural relationships in the simulation scene, obtained by the input layer of the large language model based on the parsed input data. These structural relationships can include geometric relationships between entities, physical rules, and logical constraints. The semantic constraint vector is a numerical vector representing the semantic relationships in the simulation scene, obtained by the input layer of the large language model based on the user's semantic description information.
[0049] Specifically, the traffic simulation platform inputs user semantic description information, traffic flow simulation configuration data, and map topology information into the input layer of the large language model. The input layer parses the user semantic description information, traffic flow simulation configuration data, and map topology information and generates structural constraint vectors containing geometric relationships between entities, physical rules, and logical constraints, as well as semantic constraint vectors representing semantic relationships in the simulation scenario.
[0050] S250. The ControlNet spatial adapter is called to gradually inject the structural constraint vector into the noisy image generated by the generative layer of the large language model to generate the initial traffic image.
[0051] ControlNet refers to a neural network framework that adds precise conditional control to artificial intelligence image generation models. The spatial adapter is the core module of ControlNet, responsible for receiving structural constraint vectors and encoding them into spatial control signals. Noisy images refer to image data that has not been fully denoised during the image generation process guided by the large language model. The generation of noisy images can include random generation, selection of existing similar images, etc.
[0052] Specifically, the spatial adapter of ControlNet is called on the traffic simulation platform. The structural constraint vector generated by inputting the user's semantic description information, traffic flow simulation configuration data and map topology information into the input layer of the large language model is injected into the generation layer of the large language model to generate image data that has not yet been fully denoised, so as to generate the initial traffic image.
[0053] S260. Call the ControlNet channel adapter to fuse the semantic constraint vector with the initial traffic image to obtain a dual-constraint traffic image.
[0054] The channel adapter is a core module in ControlNet, responsible for receiving semantic constraint vectors and encoding them into channel control signals. A dual-constraint traffic image refers to an output traffic image that satisfies both structural and semantic constraint vectors.
[0055] Specifically, the channel adapter of ControlNet is called on the traffic simulation platform to fuse the semantic constraint vector generated by the input layer of the large language model based on the user's semantic description information, traffic flow simulation configuration data and map topology information with the initial traffic image, so as to obtain a dual-constraint traffic image that combines structural constraint vector and semantic constraint vector. The fusion method can include weighted fusion, staged fusion, etc.
[0056] S270. The dual-constraint traffic image is backsampled and updated according to the large language model to obtain a single-frame driving image.
[0057] Backsampling update refers to the process of progressively removing noise from an image along the opposite direction of noise addition, thereby refining a noisy input image into a clear output image. A single-frame driving image refers to a static scene image obtained after backsampling update, which can be used in an autonomous driving simulation test environment.
[0058] Specifically, after obtaining a dual-constraint traffic image combining structural constraint vectors and semantic constraint vectors, the traffic simulation platform performs progressively refined backsampling update operations on the dual-constraint traffic image according to the large language model, and finally generates a single-frame driving image suitable for testing autonomous driving models.
[0059] S280. The traffic flow simulation configuration data is labeled onto a single-frame driving image, and each single-frame driving image labeled with the traffic flow simulation configuration data is used as an autonomous driving model image.
[0060] Annotation refers to the process of establishing a precise mapping relationship between traffic flow simulation configuration data and corresponding object elements in a single frame of a driving image using specific annotation methods. Specific annotation methods can include bounding box annotation, keypoint annotation, and semantic segmentation mask annotation. Autonomous driving model images refer to traffic simulation images that contain user semantic description information, traffic flow simulation configuration data, and map topology information. Autonomous driving model images can be used for testing and verification experiments of autonomous driving algorithms.
[0061] Specifically, after obtaining a single-frame driving image, the traffic simulation platform annotates the traffic flow simulation configuration data onto the single-frame driving image, resulting in a single-frame driving image with traffic flow simulation configuration data annotations. Then, each single-frame driving image annotated with traffic flow simulation configuration data is used as an autonomous driving model image.
[0062] The technical solution of this invention involves acquiring uploaded map data and generating a simulated road network structure from the map data. This is followed by acquiring user-configured traffic flow simulation configuration data, collecting user semantic description information, and inputting the user semantic description information, traffic flow simulation configuration data, and map topology information into the input layer of a large language model. The structural constraint vector and semantic constraint vector generated by the input layer are then obtained. The spatial adapter of ControlNet is then invoked to progressively inject the structural constraint vector into the noisy image generated by the generative layer of the large language model to generate an initial traffic image. The channel adapter of ControlNet is then invoked to fuse the semantic constraint vector with the initial traffic image to obtain a dual-constraint traffic image. The dual-constraint traffic image is then backsampled and updated according to the large language model to obtain a single-frame driving image. The traffic flow simulation configuration data is then labeled onto the single-frame driving image, and each single-frame driving image labeled with the traffic flow simulation configuration data is used as an autonomous driving model image. The above technical solution, by acquiring uploaded map data and generating a simulated road network structure for the map data, can maintain the road structure in line with the real road scenario. By acquiring user-configured traffic flow simulation configuration data and collecting user semantic description information, it can maintain the geometric position, scale, and logical consistency of road structure of traffic participants. By back-sampling and updating the dual-constraint traffic image according to the large language model, a single-frame driving image can be obtained, which can generate road images for various special scenarios. This solves the problems of scarce long-tail scenario data, insufficient controllability of generation, and poor semantic consistency in existing technologies for autonomous driving.
[0063] Furthermore, based on the above embodiments of the invention, obtaining user-configured traffic flow simulation configuration data includes:
[0064] The system acquires user-defined vehicle driving information and / or vehicle control information as vehicle information. The vehicle driving information includes at least the starting point, ending point, average speed, and driving style. The vehicle control information includes trajectory planning algorithm, motion control algorithm, and custom control logic.
[0065] The system acquires user-configured traffic participant types, number of traffic participants, participant locations, participant movement trajectories, participant postures, and participant interaction perception capabilities as vehicle information.
[0066] Obtain the static element types, number of static elements, static element positions, and static element postures configured by the user as traffic static element information;
[0067] Obtain the user-configured traffic light position and size as traffic light information;
[0068] Simulation predictions are performed on at least one of the following: vehicle information, traffic participant information, traffic static element information, and traffic signal light information, to obtain prediction simulation configuration data at different times.
[0069] Add the predicted simulation configuration data to the traffic flow simulation configuration data, and export the traffic flow simulation configuration data as time-series 3D bounding box data.
[0070] Simulation prediction refers to the process of extrapolating and calculating the motion state, interaction behavior, and environmental changes of traffic elements from their initial or current state over a future period based on pre-set information. Temporal 3D bounding box data refers to serialized data organized in chronological order, where each timestamp corresponds to one or more sets of 3D bounding boxes. Each set of bounding boxes describes the position, size, and orientation of a traffic element in 3D space.
[0071] Specifically, the traffic simulation platform acquires user-defined vehicle driving information and / or vehicle control information as vehicle information; it acquires user-configured traffic participant types, numbers, locations, trajectories, postures, and interactive perception capabilities as vehicle information; it acquires user-configured static element types, numbers, locations, and postures as traffic static element information; and it acquires user-configured traffic light locations and sizes as traffic signal information. The vehicle driving information includes at least the start point, end point, average speed, and driving style, while the vehicle control information includes trajectory planning algorithms, motion control algorithms, and custom control logic. The traffic simulation platform then performs simulation predictions on at least one of the vehicle information, traffic participant information, traffic static element information, and traffic signal information to obtain predicted simulation configuration data at different times. This predicted simulation configuration data is then added to the traffic flow simulation configuration data, and the traffic flow simulation configuration data is exported as temporal 3D bounding box data.
[0072] Furthermore, based on the above embodiments of the invention, user semantic description information, traffic flow simulation configuration data, and map topology information are input into the input layer of the large language model, and the structural constraint vector and semantic constraint vector generated by the input layer are obtained, including:
[0073] The map topology information and traffic flow simulation configuration data are aligned in coordinate system and time sequence according to the input layer, and the aligned map topology information and traffic flow simulation configuration data are converted into structured tensor information as structural constraint vectors.
[0074] According to the input layer, the user's semantic description information is converted into semantic feature vectors as semantic constraint vectors through semantic decomposition and multimodal encoding.
[0075] Coordinate system alignment refers to the process of uniformly transforming the spatial location information in map topology information and traffic flow simulation configuration data to the same coordinate system. Temporal alignment refers to the process of synchronizing the static and dynamic temporal information in map topology information and traffic flow simulation configuration data. Structured tensor information refers to integrating the aligned map topology information and traffic flow simulation configuration data to generate multidimensional array information with fixed dimensions.
[0076] Specifically, the input layer of the large language model receives map topology information and traffic flow simulation configuration data. It collects the corresponding coordinates of feature points in the coordinate systems of the map topology information and traffic flow simulation configuration data, and solves for transformation parameters using least squares or calibration tools to complete the coordinate transformation operation. By unifying the device clock and timestamp format, it uses interpolation, nearest neighbor, and resampling methods to match as needed, completing the timestamp matching operation. This transforms the spatial location information in the map topology information and traffic flow simulation configuration data to the same coordinate system and synchronizes the static and dynamic temporal information in the map topology information and traffic flow simulation configuration data, completing coordinate system alignment and temporal alignment operations. The aligned map topology information and traffic flow simulation configuration data are integrated and generated as structured tensor information as structural constraint vectors through coordinate transformation parameter encoding and timestamp matching parameter encoding. Simultaneously, according to the input layer, the user's semantic description information is converted into semantic feature vectors as semantic constraint vectors through semantic decomposition and multimodal encoding. The multimodal encoding methods can include word embedding encoding, pre-trained language model encoding, and speech modality encoding. The input layer of a large language model bears the crucial task of integrating multi-source heterogeneous information into a unified representation that the model can understand. This process first requires rigorous standardization of the objective environmental data. Specifically, the input layer receives topological information (such as lane networks and curb locations) from high-precision maps and temporal traffic flow data (such as vehicle trajectories and traffic light status) from the simulation engine. Since these two types of data are usually based on different spatial reference systems and time bases, the input layer performs coordinate system alignment and temporal alignment.
[0077] Furthermore, based on the above embodiments of the invention, traffic flow simulation configuration data is annotated onto a single-frame driving image, including:
[0078] Extract the temporal 3D bounding box data of traffic flow simulation configuration data;
[0079] Each time-series 3D bounding box data is labeled onto a single frame driving image according to time and location.
[0080] Traffic flow simulation configuration data refers to the set of configuration data used to generate simulated traffic environments. This data can include vehicle information, traffic participant information, static traffic element information, and traffic light information. Temporal 3D bounding box data refers to serialized data organized in chronological order, where each timestamp corresponds to one or more sets of 3D bounding boxes. Each set of bounding boxes describes the position, size, and orientation of a traffic element in 3D space. Single-frame driving images refer to static scene images obtained after backsampling and updating, suitable for use in autonomous driving simulation testing environments.
[0081] Specifically, temporal 3D bounding box data is parsed from traffic flow simulation configuration data containing information on vehicles, traffic participants, static traffic elements, and traffic lights. Each timestamp corresponds to one or more sets of 3D bounding boxes, and each set of bounding boxes describes the position, size, and orientation of a traffic element in 3D space. Then, each temporal 3D bounding box data is labeled onto a single frame of driving image according to time and location.
[0082] Furthermore, based on the above embodiments, the invention also includes:
[0083] Input the autonomous driving model image into the preset autonomous driving model for model training or model validation.
[0084] The pre-defined autonomous driving model refers to a machine learning model architecture used to achieve autonomous driving functions, possessing capabilities such as environmental perception, path planning, or motion control. Model training refers to the process of using labeled images and their corresponding annotation information as training data, updating model parameters through forward and backward propagation, enabling the model to learn and optimize its performance in autonomous driving-related tasks. Model validation refers to the process of using labeled images and their corresponding annotation information as test data, inputting them into a fully trained or training autonomous driving model, and quantifying model performance by evaluating the consistency metrics between the model's output predictions and the actual annotation information.
[0085] Example 3
[0086] Figure 3 This is a flowchart of an image generation method provided in Embodiment 3 of the present invention.
[0087] For details, see Figure 3 The flowchart of an image generation method described in this embodiment shows that the image generation system mainly includes two core modules to realize the image generation action:
[0088] The traffic simulator module is used to generate dynamic and static scene information that conforms to traffic rules based on the map data and traffic configuration input by the user, and output the temporal state data of traffic participants, static elements, traffic lights, etc. in the form of 3D bounding boxes.
[0089] The diffusion renderer module is used to input the state information of map structure, traffic participants, static elements, etc., along with natural language scene descriptions, as multimodal control signals into the visual language diffusion model to achieve high-fidelity image generation of multi-dimensional long-tail scenes.
[0090] The specific steps for the traffic simulator module to operate are as follows:
[0091] First, users upload a map JSON file to the traffic simulator. The simulator parses the Waymo dataset map and the user-defined JSON map, then constructs a topological structure including lanes, lane lines, boundaries, and drivable areas based on the map data. Second, users set their vehicle's start point, end point, average speed, or driving style. The traffic simulator's built-in vehicle control algorithm performs trajectory planning and motion control based on the user-defined data. The simulator also provides an interface for users to integrate custom control logic. Next, traffic participant behavior is edited, with two modes: batch generation and individual editing. In batch generation, the total number and category ratio of traffic participants are adjusted via parameters; the start and end points can be randomly generated with a random seed to ensure reproducibility. In individual editing, the user specifies the category, start / end point, and speed, and the system automatically plans the path and controls the movement of traffic participants. Furthermore, all dynamic traffic participants in the simulation have interactive perception capabilities, preventing collisions with other vehicles, other traffic participants, and static facilities. Finally, static elements are edited, with two modes: batch generation and individual editing. The batch generation mode allows for batch adjustment of the quantity and proportion of static elements, which are then automatically arranged according to rules. In the individual editing mode, users can freely add and adjust the type, quantity, position, and orientation of static elements. A 3D bounding box is then used to represent the position and size of traffic lights for traffic light control. The position and size information of the traffic lights can be exported in JSON format. Finally, simulation output is performed in standardized JSON format for direct use by downstream rendering modules. The simulation output allows users to select the content to export according to actual needs.
[0092] The specific steps of the diffusion renderer module are as follows:
[0093] First, the input data is constructed, comprising three parts: structural input, semantic input, and control signal integration. The structural input data consists of map JSON and BBox JSON, providing structural constraints such as road topology, location, and attitude of traffic participants. Semantic input refers to user descriptions of features such as time, weather, lighting, road surface material, and city style using natural language, encoded into multimodal feature vectors via VLM. Control signal integration involves combining ControlNet to use both structural and semantic inputs as conditional vectors for the diffusion model.
[0094] Secondly, a diffusion generation process is carried out. At time step t, iterative backsampling starts from Gaussian noise. Combined with structural constraints, the geometric consistency of roads and traffic participants is maintained. Combined with semantic control signals, multi-dimensional controllable generation is achieved. The model maintains the spatial consistency and temporal continuity of objects during the generation process, which is suitable for autonomous driving training requirements.
[0095] Finally, a high-fidelity, consistent-style, and controllable long-tail scene image sequence is output, which contains synchronous annotation information for corresponding traffic participants, static elements, and traffic lights.
[0096] The steps for implementing image generation based on the image generation system are as follows:
[0097] First, the map JSON file is imported into the traffic manager. The system can parse the custom map data and construct the simulated road network structure. Then, the EGO status of vehicles, traffic participants, static obstacles, and traffic lights are set. The setting of traffic participants and static obstacles can be done in batches or individually. Based on the user-defined vehicle and traffic participant configurations, and combined with a deep learning traffic behavior prediction model, the system generates temporal 3D bounding box data for various traffic participants during the simulation period. After the simulation, a bounding box JSON file is exported. This file includes both a map JSON file and a bounding box JSON file. The map JSON file contains static structural information such as roads, lane lines, and lane boundaries; the bounding box JSON file contains the type, location, posture, size, and timestamp information of traffic participants during the simulation. The map JSON file and the bounding box JSON file are then imported into the renderer. Using these as structural constraint inputs, sensor settings and viewpoint layouts are adjusted, and natural language text descriptions are input to form multimodal control signals. Utilizing the diffusion model embedded in ControlNet, a series of high-fidelity scene images conforming to the set conditions are generated while maintaining the logical and geometric consistency of traffic elements. Finally, the generated images and corresponding temporal labels are output for use in training or validation of autonomous driving models. The temporal labels include 3D bounding boxes, segmentation masks, etc.
[0098] The technical solution of this invention can maintain the road structure in line with the real road scene by importing a map JSON file into the traffic manager. By setting the vehicle status, traffic participants, static obstacles and traffic light information, it can maintain the geometric position, scale and logical consistency of the road structure of traffic participants. By adjusting the sensor settings and view layout and inputting natural language text descriptions, it can generate road images of various special scenarios, solving the problems of scarce long-tail scenario data, insufficient controllability of generation and poor semantic consistency in the prior art for autonomous driving.
[0099] Example 4
[0100] Figure 4 This is a schematic diagram of an image generation device provided in Embodiment 4 of the present invention. Figure 4 As shown, the device includes: an information acquisition module 410 and an image generation module 420; wherein,
[0101] The information acquisition module 410 is used to acquire map topology information, traffic flow simulation configuration data, and user semantic description information.
[0102] The image generation module 420 is used to generate autonomous driving model images corresponding to user semantic description information, traffic flow simulation configuration data, and map topology information based on the large language model.
[0103] The technical solution of this invention acquires map topology information, traffic flow simulation configuration data, and user semantic description information through an information acquisition module. An image generation module generates autonomous driving model images corresponding to the user semantic description information, traffic flow simulation configuration data, and map topology information based on a large language model. This technical solution, by acquiring map topology information, traffic flow simulation configuration data, and user semantic description information, can maintain the consistency of geometric positions, scales, and road structure logic of traffic participants. By generating autonomous driving model images corresponding to the user semantic description information, traffic flow simulation configuration data, and map topology information based on a large language model, it can generate road images for various special scenarios, solving the problems of scarce long-tail scenario data, insufficient controllability in generation, and poor semantic consistency in existing autonomous driving technologies.
[0104] Optionally, the information acquisition module 410 is specifically used for:
[0105] Acquire map topology information, traffic flow simulation configuration data, and user semantic description information.
[0106] Optionally, map topology information, traffic flow simulation configuration data, and user semantic description information can be obtained, including:
[0107] The system acquires uploaded map data and generates a simulated road network structure for the map data. The simulated road network structure includes at least one of the following: lanes, lane lines, boundaries, and drivable areas.
[0108] Obtain the user-configured traffic flow simulation configuration data, which includes vehicle information, traffic participant information, traffic static element information, and traffic light information.
[0109] Collect user semantic description information, which includes time, weather, lighting, road surface material, and city style.
[0110] Optionally, obtain user-configured traffic flow simulation configuration data, including:
[0111] The system acquires user-defined vehicle driving information and / or vehicle control information as vehicle information. The vehicle driving information includes at least the starting point, ending point, average speed, and driving style. The vehicle control information includes trajectory planning algorithm, motion control algorithm, and custom control logic.
[0112] The system acquires user-configured traffic participant types, number of traffic participants, participant locations, participant movement trajectories, participant postures, and participant interaction perception capabilities as vehicle information.
[0113] Obtain the static element types, number of static elements, static element positions, and static element postures configured by the user as traffic static element information;
[0114] Obtain the user-configured traffic light position and size as traffic light information;
[0115] Simulation predictions are performed on at least one of the following: vehicle information, traffic participant information, traffic static element information, and traffic signal light information, to obtain prediction simulation configuration data at different times.
[0116] Add the predicted simulation configuration data to the traffic flow simulation configuration data, and export the traffic flow simulation configuration data as time-series 3D bounding box data.
[0117] Optionally, the image generation module 420 is specifically used for:
[0118] Based on the large language model, generate autonomous driving model images corresponding to user semantic description information, traffic flow simulation configuration data, and map topology information.
[0119] Optionally, autonomous driving model images are generated based on the large language model, corresponding to user semantic description information, traffic flow simulation configuration data, and map topology information, including:
[0120] User semantic description information, traffic flow simulation configuration data, and map topology information are input into the input layer of the large language model, and the structural constraint vector and semantic constraint vector generated by the input layer are obtained.
[0121] The spatial adapter of ControlNet is invoked step by step to inject structural constraint vectors into the noisy image generated by the generative layer of the large language model in order to generate the initial traffic image;
[0122] The ControlNet channel adapter is called to fuse the semantic constraint vector with the initial traffic image to obtain a dual-constraint traffic image;
[0123] The dual-constraint traffic image is backsampled and updated according to the large language model to obtain a single-frame driving image;
[0124] Traffic flow simulation configuration data is labeled onto single-frame driving images, and each single-frame driving image labeled with traffic flow simulation configuration data is used as an image for the autonomous driving model.
[0125] Optionally, user semantic description information, traffic flow simulation configuration data, and map topology information are input into the input layer of the large language model, and the structural constraint vector and semantic constraint vector generated by the input layer are obtained, including:
[0126] The map topology information and traffic flow simulation configuration data are aligned in coordinate system and time sequence according to the input layer, and the aligned map topology information and traffic flow simulation configuration data are converted into structured tensor information as structural constraint vectors.
[0127] According to the input layer, the user's semantic description information is converted into semantic feature vectors as semantic constraint vectors through semantic decomposition and multimodal encoding.
[0128] Optionally, traffic flow simulation configuration data can be annotated onto single-frame driving images, including:
[0129] Extract the temporal 3D bounding box data of traffic flow simulation configuration data;
[0130] Each time-series 3D bounding box data is labeled onto a single frame driving image according to time and location.
[0131] Optional, also includes:
[0132] Input the autonomous driving model image into the preset autonomous driving model for model training or model validation.
[0133] The image generation apparatus provided in the embodiments of the present invention can execute the image generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0134] It is worth noting that the modules included in the above-mentioned image generation device are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional module are only for easy differentiation and are not used to limit the protection scope of the embodiments of the present invention.
[0135] Example 5
[0136] Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0137] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0138] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0139] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image generation methods.
[0140] In some embodiments, the image generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image generation method by any other suitable means (e.g., by means of firmware).
[0141] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0142] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0143] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0146] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0147] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the image generation method provided in any embodiment of this invention.
[0148] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0149] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0150] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An image generation method, characterized in that, The method includes: Acquire map topology information, traffic flow simulation configuration data, and user semantic description information; The system generates autonomous driving model images corresponding to the user semantic description information, the traffic flow simulation configuration data, and the map topology information based on the large language model.
2. The method according to claim 1, characterized in that, The acquisition of map topology information, traffic flow simulation configuration data, and user semantic description information includes: The uploaded map data is acquired, and a simulated road network structure of the map data is generated, wherein the simulated road network structure includes at least one of lanes, lane lines, boundaries, and drivable areas; Obtain user-configured traffic flow simulation configuration data, wherein the traffic flow simulation configuration data includes vehicle information, traffic participant information, traffic static element information, and traffic signal light information; The user semantic description information is collected, which includes time, weather, lighting, road surface material, and city style.
3. The method according to claim 2, characterized in that, The process of obtaining user-configured traffic flow simulation configuration data includes: The vehicle information includes user-defined autonomous vehicle driving information and / or autonomous vehicle control information, wherein the autonomous vehicle driving information includes at least the starting point, the ending point, the average speed, and the driving style, and the autonomous vehicle control information includes a trajectory planning algorithm, a motion control algorithm, and custom control logic. The vehicle information is obtained by acquiring the user-configured traffic participant types, number of traffic participants, participant locations, participant movement trajectories, participant postures, and participant interaction perception capabilities. The static element type, number, position, and posture configured by the user are obtained as the traffic static element information. The user-configured traffic light position and size are obtained as the traffic light information. Simulation prediction is performed on at least one of the vehicle information, traffic participant information, traffic static element information and traffic light information to obtain prediction simulation configuration data at different times. The predicted simulation configuration data is added to the traffic flow simulation configuration data, and the traffic flow simulation configuration data is exported as time-series three-dimensional bounding box data.
4. The method according to claim 1, characterized in that, The process of generating the autonomous driving model image corresponding to the user semantic description information, the traffic flow simulation configuration data, and the map topology information based on the large language model includes: The user semantic description information, the traffic flow simulation configuration data, and the map topology information are input into the input layer of the large language model, and the structural constraint vector and semantic constraint vector generated by the input layer are obtained. The spatial adapter of ControlNet is invoked step by step to inject the structural constraint vector into the noisy image generated by the generative layer of the large language model to generate the initial traffic image; The semantic constraint vector is fused with the initial traffic image by calling the ControlNet channel adapter to obtain a dual-constraint traffic image; The dual-constraint traffic image is backsampled and updated according to the large language model to obtain a single-frame driving image. The traffic flow simulation configuration data is labeled onto the single-frame driving image, and each single-frame driving image labeled with the traffic flow simulation configuration data is used as the autonomous driving model image.
5. The method according to claim 4, characterized in that, The step of inputting the user semantic description information, the traffic flow simulation configuration data, and the map topology information into the input layer of the large language model, and obtaining the structural constraint vector and semantic constraint vector generated by the input layer, includes: The map topology information and traffic flow simulation configuration data are aligned in coordinate system and in time according to the input layer, and the aligned map topology information and traffic flow simulation configuration data are converted into structured tensor information as the structure constraint vector. According to the input layer, the user semantic description information is converted into a semantic feature vector through semantic decomposition and multimodal encoding, which serves as the semantic constraint vector.
6. The method according to claim 4, characterized in that, The step of annotating the traffic flow simulation configuration data onto the single-frame driving image includes: Extract the temporal three-dimensional bounding box data of the traffic flow simulation configuration data; The temporal 3D bounding box data are labeled onto the single-frame driving image according to time and location.
7. The method according to claim 1, characterized in that, Also includes: The autonomous driving model image is input into a preset autonomous driving model for model training or model verification.
8. An image generation apparatus, characterized in that, The device includes: The information acquisition module is used to acquire map topology information, traffic flow simulation configuration data, and user semantic description information; The image generation module is used to generate autonomous driving model images corresponding to the user semantic description information, the traffic flow simulation configuration data, and the map topology information based on the large language model.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image generation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to perform the image generation method according to any one of claims 1-7.