Automatic driving scene image generation method and device, equipment, medium and product
By parsing the text of autonomous driving scenarios and rendering them using atomic capability library tools, high-fidelity autonomous driving scene images are generated, solving the problems of low efficiency and poor quality in generating complex scenes in existing technologies, and enabling efficient training and testing of autonomous driving systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to efficiently and realistically generate complex and rare autonomous driving scene images, failing to meet the massive training and testing needs of autonomous driving systems.
By parsing the text describing autonomous driving scenarios, determining filtering conditions, retrieving basic scene images from the database, and using tools in the atomic capability library to render corresponding elements at set locations, an autonomous driving scene image matching the text is generated.
It enables the efficient and flexible generation of high-fidelity autonomous driving scene images, improving the efficiency and quality of scene generation and meeting the training and testing needs of autonomous driving systems.
Smart Images

Figure CN121808089A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, device, medium and product for generating images of autonomous driving scenarios. Background Technology
[0002] With the rapid development of intelligent driving and autonomous driving technologies, ensuring vehicle safety in various complex and even extreme situations has become crucial. The safety verification of autonomous driving systems heavily relies on large-scale simulation testing, especially the simulation of rare scenarios (also known as long-tail scenarios). Long-tail scenarios refer to critical scenarios that occur very infrequently in the real world but could lead to serious consequences if they do occur, such as sudden fog, complex construction zones, unusual obstacles on the road, or abnormal behavior of pedestrians / non-motorized vehicles. Therefore, efficiently and realistically reproducing these long-tail scenarios is a core challenge in training and verifying the robustness of autonomous driving systems.
[0003] Currently, methods for generating long-tail scenes mainly include manual 3D modeling and traditional procedural generation. Manual 3D modeling relies on professional 3D modelers to manually build, arrange, and render the entire scene in simulation software. The drawbacks of this method are its extremely high cost, long development cycle, inability to achieve large-scale generation, and the fact that the quality and diversity of the scene heavily depend on the modeler's experience, lacking unified standards and efficiency. Traditional procedural generation attempts to automate the generation of scene elements by using fixed algorithms or scripts, such as using scripts to cyclically place traffic cone models within a specific coordinate range of a road to improve generation efficiency. However, its core logic is rigid, heavily reliant on manually set parameters, and is essentially still a "human brain planning, script execution" model. In addition, because the generation process lacks context awareness, it cannot understand the scene to be generated and cannot guarantee that the generated scene conforms to physical and logical consistency. For example, obstacles may be incorrectly placed in the air or inside walls.
[0004] In summary, existing methods lack intelligence, cannot fully understand the autonomous driving scenarios to be generated, and are unable to efficiently generate high-fidelity autonomous driving scenario images according to complex requirements, thus failing to meet the needs of massive training and testing of autonomous driving systems. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and product for generating autonomous driving scene images, so as to achieve efficient generation of high-fidelity autonomous driving scene images.
[0006] In a first aspect, embodiments of this application provide a method for generating images of an autonomous driving scene, including:
[0007] The text used to describe autonomous driving scenarios is parsed to determine the filtering conditions;
[0008] Based on the filtering conditions, retrieve basic scene images from the database;
[0009] Based on the text, the corresponding tool is called from the atomic capability library to render the corresponding elements at the set position of the basic scene image through the called tool, so as to obtain an autonomous driving scene image that matches the text;
[0010] The atomic capability library includes a variety of tools, each of which is used to render the corresponding element.
[0011] Secondly, embodiments of this application also provide an autonomous driving scene image generation device, comprising:
[0012] The parsing module is used to parse the text used to describe autonomous driving scenarios to determine filtering conditions;
[0013] The retrieval module is used to retrieve basic scene images from the database based on the filtering conditions;
[0014] The generation module is used to call the corresponding tool from the atomic capability library according to the text, so as to render the corresponding elements at the set position of the basic scene image through the called tool, and obtain an autonomous driving scene image that matches the text.
[0015] The atomic capability library includes a variety of tools, each of which is used to render the corresponding element.
[0016] Thirdly, embodiments of this application provide an electronic device, including:
[0017] One or more processors;
[0018] Storage device for storing one or more programs;
[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the autonomous driving scene image generation method as described in the first aspect.
[0020] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the autonomous driving scene image generation method as described in the first aspect.
[0021] Fifthly, embodiments of this application also provide a computer program product, including a computer program and / or instructions, which, when executed by a processor, implement the autonomous driving scene image generation method as described in any of the above embodiments.
[0022] This application provides a method, apparatus, device, medium, and product for generating autonomous driving scene images. The method includes: parsing text describing an autonomous driving scene to determine filtering conditions; retrieving a basic scene image from a database based on the filtering conditions; and calling a corresponding tool from an atomic capability library based on the text to render corresponding elements at predetermined positions in the basic scene image using the called tool, thereby obtaining an autonomous driving scene image matching the text. The atomic capability library includes multiple tools, each used to render corresponding elements. This technical solution, by parsing text to determine filtering conditions, can fully understand the required autonomous driving scene and retrieve a suitable basic scene image accordingly. Furthermore, by calling tools from the atomic capability library to render corresponding elements for the autonomous driving scene, it efficiently generates high-fidelity autonomous driving scene images, thus meeting the needs of massive training and testing of autonomous driving systems. Attached Figure Description
[0023] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0024] Figure 1 A flowchart illustrating an autonomous driving scene image generation method provided in this application embodiment;
[0025] Figure 2 A schematic diagram of an autonomous driving scene image provided in an embodiment of this application;
[0026] Figure 3 A schematic diagram of another autonomous driving scenario image provided in an embodiment of this application;
[0027] Figure 4 A schematic diagram of another autonomous driving scenario image provided in an embodiment of this application;
[0028] Figure 5 This is a schematic diagram of the structure of an autonomous driving scene image generation device provided in an embodiment of this application;
[0029] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0030] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present application, not the entire structure.
[0031] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. The process can be terminated when its operation is complete, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0032] It should be noted that the concepts of "first" and "second" mentioned in the embodiments of this application are only used to distinguish different devices, modules, units or other objects, and are not used to limit the order of functions performed by these devices, modules, units or other objects or their interdependencies.
[0033] Furthermore, the embodiments and features described in this application may be combined with each other, unless otherwise specified.
[0034] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.
[0035] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the relevant content of the solution.
[0036] Figure 1 This is a flowchart illustrating a method for generating autonomous driving scene images according to an embodiment of this application. This embodiment is applicable to situations where autonomous driving scene images are generated based on text descriptions. Specifically, this method for generating autonomous driving scene images can be executed by an autonomous driving scene image generation device. This device can be implemented through software and / or hardware and integrated into an electronic device. The electronic device includes, but is not limited to, devices with control functions such as computers, smartphones, host computers, or servers.
[0037] like Figure 1 As shown, the method specifically includes the following steps:
[0038] S110. Parse the text used to describe the autonomous driving scenario to determine the filtering conditions;
[0039] The text can be entered by the user and contains information describing the autonomous driving scenario, such as time, location, weather, lighting, roads, lane markings, traffic lights, buildings, obstacles, pedestrians, and / or vehicles. Furthermore, the text can contain one or more keywords, as well as one or more sentences. Autonomous driving scenarios can include uncommon scenarios, such as long-tail scenarios.
[0040] The process of parsing text can be implemented based on Natural Language Processing (NLP). For example, artificial intelligence models can be used to parse text to understand its content and clarify user intent. During this process, filtering conditions can be determined; these conditions can also be understood as retrieval criteria and serve as the basis for filtering or retrieving the basic scene images needed to generate the autonomous driving scenario.
[0041] S120. Retrieve basic scene images from the database according to the filtering conditions;
[0042] The database stores a large number of basic scene images, covering various common scenarios. These images can be real-life photos or images collected and uploaded by vehicles with autonomous driving capabilities. Each basic scene image has corresponding formatted data annotations, such as time, location, weather, lighting, and / or road information, with the format consistent with the filtering criteria. Basic scene images that meet the filtering criteria can be used as the original images for generating autonomous driving scenarios. If the text specifies that a certain number of autonomous driving scene images need to be generated, the corresponding number of basic scene images can be retrieved from the database.
[0043] For example, if the text reads "Generate a cardboard box obstacle on a cloudy city road", then the filtering conditions include: cloudy weather and city road environment. Based on this, a basic scene image that meets the filtering conditions is retrieved. S130: According to the text, the corresponding tool is called from the atomic capability library to render the corresponding elements at the set position of the basic scene image through the called tool, thereby obtaining an autonomous driving scene image that matches the text;
[0044] The Atomic Capability Library can be understood as a toolbox, containing various tools (tools can also be understood as atomic capabilities or modules). Each tool has 3D rendering capabilities and can contain parameterized, complex scripts for rendering corresponding elements. Each tool's script can include corresponding code, algorithms, models, and / or functions. For example, a tool can render specific elements on a base scene image, such as rendering a cardboard box on a road surface, rendering fog in a distant view area, or changing the lighting effects or color tone of the base scene image. The Atomic Capability Library can also include documentation describing the function, parameters, and usage of each tool, serving as a knowledge base for artificial intelligence models.
[0045] Based on the text, we can analyze the requirements for generating autonomous driving scene images and call appropriate tools from the atomic capability library to generate corresponding commands. For example, in the example above, "generate a cardboard box obstacle on a cloudy city road," the retrieved base scene image satisfies the conditions of cloudy weather and city road environment. Therefore, we can call a tool for rendering cardboard boxes and use that tool to render a cardboard box at a predetermined location within the base scene image.
[0046] The location can be determined based on the content of the text and / or the constraints provided by the tool. For example, in the example above, "generate a cardboard box obstacle on a cloudy city road," the location must be the road surface area in the base scene image, not the area of buildings or the sky. This ensures that the rendered elements conform to the rules of the real scene, avoiding the generation of erroneous or meaningless autonomous driving scene images.
[0047] The autonomous driving scene image generation method in this embodiment achieves automated batch generation driven by natural language. Users only need to describe their input requirements using natural text to automatically and in batches generate the required autonomous driving scene images, greatly improving the efficiency and scale of scene generation. Furthermore, this method is highly scalable and flexible. The atomic capability library design decouples scene editing capabilities from the intelligent orchestration core, allowing developers to dynamically add new long-tail scene generation code (such as new weather effects, new obstacle types, etc.) to the capability library at any time without requiring major modifications to the intelligent core or scene analysis module, significantly improving the system's iteration speed and maintainability.
[0048] In one embodiment, the method further includes:
[0049] S1210. Perform semantic segmentation on the basic scene image to obtain the mask constraint corresponding to the basic scene image;
[0050] S1220. Determine the set position according to the mask constraint.
[0051] Semantic segmentation can be achieved by combining artificial intelligence models (such as the Grounded SAM model) with camera intrinsic and extrinsic parameters. Its main purpose is to achieve accurate scene perception by analyzing instances and spatial information in the basic scene image, obtain mask constraints, and provide precise position and region restrictions for scene editing. This allows the corresponding elements to be rendered to the set positions, ensuring the authenticity of the generated autonomous driving scene image and its logical consistency with the real world. It avoids errors such as element misalignment, floating, and clipping that do not conform to physical laws in traditional methods, thereby improving the quality of the generated autonomous driving scene image.
[0052] In one embodiment, the text includes descriptive information about the autonomous driving scenario, descriptive information about the rendered elements, and the number of autonomous driving scenario images generated.
[0053] The text used to describe autonomous driving scenarios is parsed to determine filtering conditions, including generating filtering conditions in JSON format based on the description information of the autonomous driving scenarios.
[0054] Based on filtering conditions, retrieve basic scene images from the database, including: retrieving basic scene images from the database based on filtering conditions in JSON format;
[0055] Based on the text, the corresponding tools are invoked from the atomic capability library, including: based on the description information of the rendered elements, the corresponding tools are invoked from the atomic capability library to obtain a corresponding number of autonomous driving scene images through the invoked tools.
[0056] For example, the text can include descriptive information about autonomous driving scenarios, used to describe common or basic scene images. Based on this descriptive information, filtering conditions can be determined and used as the basis for retrieving basic scene images. The text can also include descriptive information about rendered elements, used to describe the differences between the rendered elements or the common autonomous driving images to be generated and the common or basic scene images. Based on this descriptive information, the tools to be invoked can be determined to render the basic scene images accordingly. For example, in the example above, "generate a cardboard box obstacle on a cloudy city road," "cloudy" and "city road" are information used to describe common or basic scene images, while "generate a cardboard box obstacle" is difference information. Furthermore, the text can also include information specifying the required or generated number of basic scene images. For example, the text can specify generating multiple autonomous driving scene images, or generating an autonomous driving scene video containing multiple autonomous driving scene images, specifically, specifying a frame length of 10 or generating 10 autonomous driving scene images in the text. Optionally, multiple autonomous driving scene images can be generated from the same base scene image or from different base scene images, depending on the method described in the text or pre-configured. Based on this, flexible, diverse, and high-fidelity autonomous driving scene images can be generated.
[0057] In one embodiment, the step of rendering corresponding elements at predetermined locations in the base scene image using a invoked tool to obtain an autonomous driving scene image matching the text includes:
[0058] By using the invoked tool, corresponding elements are added to the set positions of the base scene image, and ambient occlusion (AO) is calculated to generate shadows of the elements, and anti-aliasing is applied to the elements.
[0059] For example, the invoked tools can calculate ambient occlusion (AO) to generate soft shadows for elements, such as the ground shadow and self-occlusion shadow of a cardboard box. Furthermore, 4x supersampling anti-aliasing (4x SSAA) can be implemented. This involves rendering the image to a buffer four times higher than the screen's native resolution, then downsampling it back to the original resolution. During this process, edge pixels can be blurred to smooth out jagged edges, making diagonal lines and curves appear more natural. Based on this, the fidelity and realism of the generated autonomous driving scene images can be improved.
[0060] In one embodiment, the method is implemented based on Large Multimodal Models (LMMs). As the intelligent core, the LMM integrates automated task scheduling and execution algorithms, responsible for understanding user natural language requirements, generating task plans, retrieving data, and intelligently orchestrating and calling appropriate tools. Using artificial intelligence models for text parsing, image retrieval, and tool invocation ensures the authenticity of the generated autonomous driving scene images and their logical consistency with the real world, avoiding errors in traditional methods such as element misalignment, floating, and clipping that do not conform to physical laws, thus improving the quality of the generated autonomous driving scene images.
[0061] In one embodiment, the filtering conditions are JSON filtering conditions. By adopting JSON-formatted filtering conditions and JSON-formatted annotations of the basic scene images recorded in the database, complex filtering logic is supported, improving the flexibility and expressiveness of retrieval. It can quickly retrieve the required basic scene images from large-scale data, reduce unnecessary data transmission and processing overhead, and also make the code logic clearer, easier to understand and maintain, thus reducing code complexity.
[0062] The following describes the process of generating the above-mentioned autonomous driving scene image through a specific example. For example, the process of generating the autonomous driving scene image includes: Step 1: Task Input
[0063] Using LMM as the intelligent core, it receives natural text descriptions input by users, specifying the number of scene frames to be generated (e.g., "Generate a cardboard box obstacle on a cloudy city road, with a frame length of 10, generate 10 frames.").
[0064] Step 2: Scene Retrieval
[0065] The intelligent core first understands the user's intent, parses the natural text into structured JSON filtering conditions and the sorted natural text after removing the number of images, and generates a scene description in JSON format, such as {"Weather": "Dusk", "Environment": "City Road", "Text": "Generate a cardboard box obstacle on a cloudy city road, frame length 10"}; then, the intelligent core retrieves basic scene images that meet the filtering conditions from the database according to the JSON conditions, until the number reaches the number of images specified by the user.
[0066] Step 3: Intelligent Orchestration and Perceptual Constraints
[0067] By consulting the atomic capability library, the intelligent core identified the "cardboard box obstacle," selected the `generate_box.py` tool from the library, and then generated the following shell command: `python generate_box.py \`
[0068] --scene-path "example_scene" \
[0069] --output-dir ". / output_longtail" \
[0070] --num-frames 10 \
[0071] --sam_result ". / sam_result" \
[0072] --box-center-x 20.0 \
[0073] --box-center-y -0.5 \
[0074] --box-yaw-degrees 30.0 \
[0075] --ao-samples 128
[0076] Depending on the segmentation requirements of the tool, the scene analysis module (Grounded SAM) can be invoked to analyze the basic scene image, obtain high-precision masks (such as sky and roads), and then save the results to the sam_result folder.
[0077] Step 4: Scene Generation
[0078] The core function of `generate_box.py` is to batch render a high-fidelity 3D cardboard box onto an autonomous driving video sequence (which can contain multiple base scene images). Specifically, scene data and camera intrinsics and extrinsics can be loaded via `--scene-path`, and the number of frames specified by `--num-frames` can be processed. Parameters such as `--box-length`, `--box-center-x`, and `--box-yaw-degrees` can be used to precisely define the static 3D position and orientation of the cardboard box. When rendering each frame (`process_single_frame`), the script loads a pre-calculated road mask from `sam_result` and uses this mask as a physical constraint to ensure that the back projection and shadow calculations in the `project_box_onto_frame` function only occur on these "road" pixels. To achieve realism, the script can also calculate dense ambient occlusion (AO) soft shadows (including ground shadows and cardboard box self-occlusion) and 4x supersampling anti-aliasing (SSAA) via the `--ao-samples` parameter. Finally, the synthesized autonomous driving scene image and the stitched autonomous driving video sequence are saved to the directory specified by `--output-dir`.
[0079] Step 5: Output
[0080] Output the location of the saved autonomous driving scene image or autonomous driving video sequence and notify the user.
[0081] Figures 2 to 4 The images show three different autonomous driving scenarios: "a cardboard box obstacle is generated on a cloudy city road", "damaged lane markings are on a cloudy city road", and "heavy fog is on a cloudy city road".
[0082] Figure 5 This is a schematic diagram of the structure of an autonomous driving scene image generation device provided in an embodiment of this application. Figure 5 As shown, the autonomous driving scene image generation device provided in this embodiment includes:
[0083] Parsing module 210 is used to parse the text used to describe autonomous driving scenarios to determine filtering conditions;
[0084] The retrieval module 220 is used to retrieve basic scene images from the database according to the filtering conditions;
[0085] The generation module 230 is used to call the corresponding tool from the atomic capability library according to the text, so as to render the corresponding element at the set position of the basic scene image by the called tool, and obtain an autonomous driving scene image that matches the text.
[0086] The atomic capability library includes a variety of tools, each of which is used to render the corresponding element.
[0087] This device determines filtering conditions by parsing text, which can fully understand the required autonomous driving scenario and retrieve appropriate base scene images accordingly. By calling tools in the atomic capability library, it can render the corresponding elements of the autonomous driving scenario, thereby efficiently generating high-fidelity autonomous driving scenario images, thus meeting the needs of massive training and testing of autonomous driving systems.
[0088] Based on any of the above embodiments, the device further includes:
[0089] The scene analysis module is used to perform semantic segmentation on the basic scene image to obtain the mask constraints corresponding to the basic scene image;
[0090] The position determination module is used to determine the set position based on the mask constraint.
[0091] Based on any of the above embodiments, the text includes descriptive information about the autonomous driving scenario, descriptive information about the rendered elements, and the number of autonomous driving scenario images generated.
[0092] The parsing module 210 is specifically used for: generating JSON-formatted filtering conditions based on the description information of the autonomous driving scenario;
[0093] The retrieval module 220 is specifically used to: retrieve basic scene images from the database according to the filtering conditions in the JSON format;
[0094] The generation module 230 is specifically used to: call the corresponding tools from the atomic capability library according to the description information of the rendered elements, so as to obtain a corresponding number of autonomous driving scene images through the called tools.
[0095] Based on any of the above embodiments, the generation module 230 is specifically used to: add corresponding elements at a set position in the basic scene image by calling a tool, calculate ambient light occlusion to generate the shadow of the elements, and perform anti-aliasing processing on the elements.
[0096] Based on any of the above embodiments, the method is implemented based on a multimodal large model (LMM).
[0097] Based on any of the above embodiments, the filtering conditions are JSON filtering conditions.
[0098] The autonomous driving scene image generation apparatus provided in this application embodiment can be used to execute the autonomous driving scene image generation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0099] Figure 6 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 10 may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, user equipment, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0100] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0101] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks and wireless networks.
[0102] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above.
[0103] In some embodiments, the methods described above can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the methods of any of the embodiments described above by any other suitable means (e.g., by means of firmware).
[0104] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0105] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0106] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device 10, which includes: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device 10. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0108] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0109] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0110] This application also provides a computer program product, including a computer program and / or instructions, which, when executed by a processor, implement the autonomous driving scene image generation method as described in any of the above embodiments.
[0111] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0112] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for generating images of autonomous driving scenes, characterized in that, include: The text used to describe autonomous driving scenarios is parsed to determine the filtering conditions; Based on the filtering conditions, retrieve basic scene images from the database; Based on the text, the corresponding tool is called from the atomic capability library to render the corresponding elements at the set position of the basic scene image through the called tool, so as to obtain an autonomous driving scene image that matches the text; The atomic capability library includes a variety of tools, each of which is used to render the corresponding element.
2. The method according to claim 1, characterized in that, Also includes: Semantic segmentation is performed on the base scene image to obtain the mask constraints corresponding to the base scene image; The set position is determined based on the mask constraint.
3. The method according to claim 1, characterized in that, The text includes descriptive information about the autonomous driving scenario, descriptive information about the rendered elements, and the number of autonomous driving scenario images generated. The step of parsing the text used to describe autonomous driving scenarios to determine filtering conditions includes: The filter conditions in JSON format are generated based on the description information of the autonomous driving scenario; Based on the filtering conditions, retrieve basic scene images from the database, including: Based on the filtering conditions in the JSON format, retrieve basic scene images from the database; The step of calling the corresponding tool from the atomic capability library based on the text includes: Based on the description information of the rendered elements, the corresponding tools are called from the atomic capability library to obtain a corresponding number of autonomous driving scene images through the called tools.
4. The method according to claim 1, characterized in that, The process of rendering corresponding elements at predetermined locations in the base scene image using the invoked tool to obtain an autonomous driving scene image matching the text includes: The tool is invoked to add corresponding elements at set positions in the base scene image, calculate ambient occlusion to generate shadows for the elements, and perform anti-aliasing on the elements.
5. The method according to claim 1, characterized in that, The method is based on a multimodal large model (LMM).
6. The method according to claim 1, characterized in that, The filtering conditions are JSON filtering conditions.
7. A method for generating images of autonomous driving scenes, characterized in that, include: The parsing module is used to parse the text used to describe autonomous driving scenarios to determine filtering conditions; The retrieval module is used to retrieve basic scene images from the database based on the filtering conditions; The generation module is used to call the corresponding tool from the atomic capability library according to the text, so as to render the corresponding elements at the set position of the basic scene image through the called tool, and obtain an autonomous driving scene image that matches the text. The atomic capability library includes a variety of tools, each of which is used to render the corresponding element.
8. An electronic device, characterized in that, include: At least one processor; A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the autonomous driving scene image generation method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the autonomous driving scene image generation method as described in any one of claims 1-6.
10. A computer program product comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by the processor, they implement the autonomous driving scene image generation method as described in any one of claims 1-6.