Image generation method, device, electronic device and storage medium
By using a game engine to generate rendering images and auxiliary enhancement information with SD and ControlNet models, the method addresses the limitations of existing automatic driving simulation technologies, enhancing the accuracy and flexibility of image style transfer and enhancement for improved simulation testing.
Patent Information
- Application Number
- CN202311763211.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-12-20
AI Technical Summary
In the existing technology, in autonomous driving simulation, traditional rendering solutions have problems such as single migration style and low accuracy. The NeRF reconstruction solutions lack flexibility and editing capabilities, and cannot handle unacquisitive working conditions, resulting in inaccurate control of simulated images.
The game engine obtains the rendered image and auxiliary enhancement information of the original image, combines the target prompt word and inputs the target model, and uses the SD model and the ControlNet plug-in model to migrate and enhance images to achieve accurate control of the image.
It improves the accuracy and flexibility of autonomous driving simulation testing, and can generate multi-style simulation images to meet the diverse needs of autonomous driving simulation.
Smart Images

Figure CN117830580B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving technology, and in particular to technical fields such as simulation testing and image processing. Background Art
[0002] In the field of autonomous driving, simulation synthetic data mainly includes traditional rendering schemes mainly based on computer graphics and game engines and neural radiance fields (NeRF) reconstruction schemes mainly based on deep learning. Although traditional rendering schemes have strong flexibility and editing capabilities, due to the domain difference between rendered images and real sensors, a style transfer or image enhancement model often needs to be added after rendering the images, and accurate image control cannot be directly achieved. On the other hand, the NeRF reconstruction scheme uses real sensor data to reconstruct multi-view images, cannot handle uncollected working conditions, and is extremely lacking in flexibility and editing capabilities. Therefore, how to achieve precise control over the migration and enhancement of simulation images is the key to realizing autonomous driving simulation. Summary of the Invention
[0003] The present disclosure provides an image generation method, apparatus, electronic device, and storage medium.
[0004] According to a first aspect of the present disclosure, there is provided an image generation method applied to the field of autonomous driving simulation, including:
[0005] Determine an original image and a target prompt word that matches the target style;
[0006] Obtain a rendered image of the original image and auxiliary enhancement information obtained through a game engine;
[0007] Input the target prompt word, the rendered image, and the auxiliary enhancement information into a target model;
[0008] Obtain a simulation image of the target style of the original image generated by the target model.
[0009] According to a second aspect of the present disclosure, there is provided an image generation apparatus applied to the field of autonomous driving simulation, including:
[0010] A determination module for determining an original image and a target prompt word that matches the target style;
[0011] A first acquisition module for obtaining a rendered image of the original image and auxiliary enhancement information obtained through a game engine;
[0012] An input module for inputting the target prompt word, the rendered image, and the auxiliary enhancement information into a target model;
[0013] A second acquisition module, configured to acquire a simulation image of a target style of an original image generated by a target model.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] At least one processor;
[0016] A memory communicatively connected to the at least one processor;
[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any embodiment in the present disclosure.
[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method of any embodiment in the present disclosure.
[0019] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program stored on a storage medium, and the computer program, when executed by a processor, implements the method of any embodiment in the present disclosure.
[0020] According to the solution of the present disclosure, a rendered image and auxiliary enhancement information of an original image are obtained through a game engine, and the target prompt word, the rendered image and the auxiliary enhancement information provided by the game engine are input into the target model, so as to realize precise control of image migration and enhancement, to meet the requirements of multi-style image migration of autonomous driving simulation data, and further improve the accuracy of autonomous driving simulation testing.
[0021] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present application will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in the present application and should not be regarded as limiting the scope of the present application.
[0023] Figure 1 It is a schematic flowchart of an image generation method according to an embodiment of the present disclosure;
[0024] Figure 2 It is a schematic architecture diagram of an image generation method according to an embodiment of the present disclosure;
[0025] Figure 3 It is a schematic diagram of the process for determining auxiliary enhancement information according to an embodiment of the present disclosure;
[0026] Figure 4 It is a schematic diagram of obtaining a target prompt word according to an embodiment of the present disclosure;
[0027] Figure 5 It is a schematic diagram of the structure of an image generation device according to an embodiment of the present disclosure;
[0028] Figure 6 It is a schematic diagram of a scenario of an image generation method according to an embodiment of the present disclosure;
[0029] Figure 7 It is a schematic diagram of the structure of an electronic device for implementing the image generation method according to an embodiment of the present disclosure. Detailed implementation manners
[0030] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described here without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] Terms such as "first", "second", and "third" in the description embodiments, claims, and the above-mentioned accompanying drawings of the present disclosure are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0032] In the field of autonomous driving, the synthetic data for autonomous driving simulation mainly includes traditional rendering solutions mainly based on computer graphics and game engines and NeRF reconstruction solutions mainly based on deep learning. Although the rendering solutions of traditional game engines have strong flexibility and editing capabilities, due to the domain difference between the rendered images and real sensors, it is often necessary to add a style transfer or image enhancement model after the rendered images, such as Convolutional Neural Networks (CNN), Cycle-consistent Generative Adversarial Networks (CycleGan), etc. However, the NeRF reconstruction solution needs to use real sensor data to reconstruct multi-view images and naturally cannot handle uncollected working conditions. Moreover, the NeRF reconstruction solution lacks flexibility and editing capabilities extremely and is unable to cope well with the data of corner cases.
[0033] In the prior art, for the traditional rendering solution mainly based on computer graphics and game engines, and adding a CNN or CycleGan model for style transfer or image enhancement after the rendered images, the following problems exist:
[0034] 1. The transferred style is single; specifically, new models need to be trained for different style transfer requirements.
[0035] 2. The accuracy is not high; specifically, for the Generative Adversarial Networks (GAN) model, the image processing effects are extremely polarized. In practical applications, it is impossible to manually screen the results for a large number of image generation tasks.
[0036] In order to at least partially solve one or more of the above problems and other potential problems, the present disclosure proposes an image generation method. By using the game engine to provide perfect auxiliary enhancement information such as depth, normal, contour, and semantics, the rendered image of the original image is obtained through the game engine. The target prompt word, the rendered image provided by the game engine, and the auxiliary enhancement information are input into the target model to achieve precise control of image transfer and enhancement, so as to meet the multi-style image transfer requirements of autonomous driving simulation data, and further improve the accuracy of autonomous driving simulation tests.
[0037] An embodiment of the present disclosure provides an image generation method, which is applied to the field of autonomous driving simulation. Figure 1FIG. 0 is a schematic flowchart of an image generation method according to an embodiment of the present disclosure. The image generation method can be applied to an image generation device. The image generation device is located in an electronic device. The electronic device includes, but is not limited to, a fixed device and / or a mobile device. For example, the fixed device includes, but is not limited to, a server, and the server can be a cloud server or a general server. For example, the mobile device includes, but is not limited to, a vehicle terminal, a laptop computer, a tablet computer, etc. In some possible implementation manners, the image generation method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, the image generation method includes:
[0038] S101: Determine an original image and a target prompt word that matches the target style;
[0039] S102: Obtain a rendered image of the original image and auxiliary enhancement information obtained by a game engine;
[0040] S103: Input the target prompt word, the rendered image, and the auxiliary enhancement information into a target model;
[0041] S104: Obtain a simulation image of the target style of the original image generated by the target model.
[0042] In some embodiments, the original image can be image data collected by an autonomous vehicle. The original image can be a road condition image collected by a collection device of the autonomous vehicle. The original image can also be a traffic image collected by a collection device of the autonomous vehicle. The original image can also be an environmental image collected by a collection device of the autonomous vehicle. The above is only an exemplary illustration and does not limit all possible contents of the original image, but only does not list them all here.
[0043] In some embodiments, for the style requirements of the simulation image of the target style, the target style can include style requirements for the weather of the simulation image, such as a highway in a heavy rainstorm; the target style can also include style requirements for the lighting of the simulation image, such as an intersection under natural light. The above is only an exemplary illustration and does not limit all possible contents of the target style, but only does not list them all here.
[0044] In some embodiments, the target prompt word can be obtained from a pre-established prompt word library (also called a Prompt library).
[0045] To establish a Prompt library, the following steps can be performed:
[0046] Determine the goals and scope of the library: First, clarify the topics and domains you hope the Prompt library will cover. For example, you can choose to create a Prompt library for programming problems or one for autonomous driving.
[0047] Collect suitable Prompts: Based on the goals and scope, start collecting relevant Prompts. This can include common questions, tips, example sentences, etc. You can collect Prompts from various resources such as documents, websites, books, discussion forums, etc.
[0048] Organize and classify Prompts: Organize and classify the collected Prompts. You can create different topics or categories to better organize and use these Prompts. For example, you can classify programming problems by programming language or difficulty level.
[0049] Write and collate Prompt descriptions: Write descriptions for each Prompt so that users can better understand what each Prompt covers. These descriptions can include relevant example questions, usage tips, expected answer formats, etc.
[0050] Continuously update and maintain: A good Prompt library should be continuously updated and maintained. You can add, update, and modify Prompts based on user feedback, new questions, and topics.
[0051] It should be noted that building a Prompt library requires a certain amount of time and effort, and continuous updating and maintenance are needed to ensure its practicality and accuracy.
[0052] In practical applications, based on various simulation style requirements of simulation tests, multiple prompt words that match the various simulation style requirements can be generated and stored in the prompt word library. This prompt word library can be stored in the data source of the simulation device.
[0053] In some embodiments, the game engine refers to the core components of some pre-written editable computer game systems or some interactive real-time image application programs. The game engine can include a rendering engine, that is, a "renderer", which includes a two-dimensional image engine and a three-dimensional image engine, responsible for rendering the scenes in the game to achieve realistic visual effects, a physics engine, a collision detection system, a script engine, scene management, etc.
[0054] In some embodiments, the game engine may be the Unity engine (UnityEngine), the Unreal engine (UnrealEngine), the Cocos2d-x game engine (Cocos2d-x Game Engine), or the Unreal Engine 4 (UnrealEngine4, UE4). Here, the present disclosure does not specifically limit the type of game engine, and it can be flexibly configured according to actual needs.
[0055] In some embodiments, the auxiliary enhancement information may include: Depth, Normal, Contour, Semantics, segmentation, Canny, and other information.
[0056] In some embodiments, the depth is used to determine the hierarchical relationship of objects in the scene. When the depth value is larger, the object is closer to the observer; when the depth value is smaller, the object is farther from the observer. Depth plays an important role in rendering. Depth can affect the objects in the scene and the way the objects are presented. In short, depth plays an important role in image rendering. Depth is used to determine the hierarchical relationship of objects in the scene, hide occluded objects, determine the drawing order, etc. When performing rendering, it is necessary to accurately calculate the depth value and correctly use the depth value to ensure high-quality rendering effects.
[0057] In some embodiments, the normal refers to a vector perpendicular to the surface. The normal is used to represent the direction and curvature of the surface. The normal is an important property of the model surface and can be used to achieve various rendering effects, such as lighting, shadows, reflections, etc. In short, in image rendering, the normal is one of the properties of the model surface and is used to represent the surface direction and curvature. The calculation and application of the normal are crucial for achieving high-quality rendering effects. By using the normal information, a rendered image with clear edges and details can be generated.
[0058] In some embodiments, the contour refers to the line of the object's edge and is used to define the shape and boundary of the object. The recognition and extraction of the contour line are crucial for the clarity and details of the rendered image. In short, in image rendering, the contour refers to the line of the object's edge and is used to define the shape and boundary of the object. The recognition and extraction of the contour line are crucial for the clarity and details of the rendered image. Commonly used contour extraction methods include methods based on color edges, gradient directions, and depth information. By using these methods and techniques, a rendered image with clear edges and details can be generated.
[0059] In some embodiments, the semantics refer to the understanding and interpretation of image content. Semantics involve the recognition and analysis of information such as objects, scenes, actions, etc. in the image, as well as the associations and relationships between this information. In short, in image rendering, semantics refer to the understanding and interpretation of image content. Through semantic analysis, we can better recognize and understand the information in the image, and thus better present this information during the rendering process. This is of great significance for achieving higher levels of image processing and applications.
[0060] In some embodiments, the block refers to dividing the pixels in an image into multiple regions or categories, that is, segmenting the image into different parts. This process is also called image segmentation. In short, the block is an important step in image rendering, which can help us better understand and process different parts of the image, thus achieving a more refined and accurate rendering effect.
[0061] In some embodiments, the edge refers to the boundary between regions where the brightness or color in the image changes sharply. The edge is one of the most basic features of an image and is also one of the important research objects in the fields of computer vision and image processing. In short, in image rendering, the edge refers to the boundary between regions where the brightness or color in the image changes sharply. By recognizing and extracting the edge, we can better understand and process the information in the image, thus achieving higher levels of image processing and applications. At the same time, by processing and optimizing the edge, we can improve the visual effect and clarity of the image, making the image more beautiful and easier to observe.
[0062] In some embodiments, the target model includes the StableDiffusion (SD) model and the ControlNet plug-in model.
[0063] In some embodiments, the simulation image is used in the field of autonomous driving simulation and can be used as data for autonomous driving simulation synthesis. Exemplarily, the simulation image obtained by the above method can be used to simulate and test a certain autonomous driving function of an autonomous vehicle; the simulation image obtained by the above method can also be used to simulate and test the safety of an autonomous vehicle under different weather conditions; the simulation image obtained by the above method can also be used to simulate and test the road environment where the autonomous vehicle is located. The simulation image obtained by the above method can also be used to simulate and test the route planning and path planning of an autonomous vehicle. The above is only an exemplary illustration and does not limit all possible simulation tests applicable to the simulation image. There is no exhaustive list here.
[0064] Figure 2 The schematic diagram of the architecture of the image generation method is shown, as Figure 2As shown, the original image is rendered based on a game engine to obtain a rendered image; the original image is classified based on the game engine to obtain auxiliary enhancement information such as edges, depth, normals, blocks, contours, semantics, etc.; the auxiliary enhancement information is input into the SD model through the ControlNet plug-in model, and the rendered image and multiple prompt words are input into the SD model, so that the SD model performs image enhancement and style transfer on the original image, and finally obtains a simulation image.
[0065] In the technical solution of the embodiment of the present disclosure, the original image and the target prompt word matching the target style are determined; the rendered image and the auxiliary enhancement information of the original image obtained through the game engine are obtained; the target prompt word, the rendered image and the auxiliary enhancement information are input into the target model; the simulation image of the target style of the original image generated by the target model is obtained. In this way, first, the rendered image and the auxiliary enhancement information of the original image are obtained through the game engine, and then the target prompt word and the rendered image and the auxiliary enhancement information provided by the game engine are input into the target model. Since the game engine usually has more advanced and efficient rendering technologies, the obtained rendered image has significant improvements in image quality and performance, and the auxiliary enhancement information is more realistic and authentic. By inputting the target prompt word and the rendered image and the auxiliary enhancement information provided by the game engine into the target model for image generation, precise control of image migration and enhancement can be achieved to meet the multi-style image migration requirements of autonomous driving simulation data, thereby improving the accuracy of autonomous driving simulation tests.
[0066] In the embodiment of the present disclosure, the target model includes an SD model and a ControlNet plug-in model.
[0067] In some embodiments, the SD model, also known as the StableDiffusion model. The structure of the SD model includes multiple components and models, and a pre-trained model based on Contrastive Language-Image Pre-training (CLIP). The CLIP is used to convert the input text into a digital representation form. The model has two main parts: an image information creator and an image decoder. The image information creator is used to generate image information from the text description, and the image decoder is used to decode the image information into an actual image. In the SD model, a diffusion model based on the latent space is also used. The diffusion model first compresses the image into the latent space, and then uses the diffusion model to generate the latent space representation of the image. This latent space-based method can improve the computational efficiency because the latent space of the image is usually smaller than the pixel space. Finally, through the use of an autoencoder, the generated latent space representation can be decoded into an actual image. It should be noted that in the SD model, stability constraints are also used to ensure that the generated image is more stable and closer to the real image.
[0068] In some embodiments, the SD model has the following advantages in image rendering: high fidelity, that is, by processing images through deep learning technology, it can improve the fidelity of images and better display the details and information of the original images; high-quality requirements, that is, by using a high-quality image rendering engine, it can output images with high clarity and realism, meeting users' requirements for perfect visual effects; high accuracy, that is, the algorithm has high accuracy and can better maintain the features and forms of images when processing images, reducing problems such as image distortion or blurring; strong scalability, that is, it supports users to input text descriptions, liberating users' creative imaginations, fully meeting users' requirements for perfect visual effects, and helping users better express their creative ideas; good stability, that is, during the training process, by using stability constraints, it can ensure that the generated images are more stable and closer to real images.
[0069] In some embodiments, the ControlNet plug-in model is a plug-in model for controlling the generation of artificial intelligence (AI) images. This model uses the technology of conditional generative adversarial networks (CGANs) to generate images. ControlNet allows users to have fine control over the generated images.
[0070] In some embodiments, the ControlNet plug-in model has the following advantages in image rendering: precise control, that is, ControlNet allows users to have fine control over the generated images through conditional generative adversarial network technology. This enables users to more accurately control the details and features of the images, thus obtaining better rendering effects; high flexibility, that is, the ControlNet plug-in model can be flexibly applied to different image rendering scenarios and requirements. Whether it is for discrete control and process control, or for application scenarios such as image registration, shape change, and autonomous driving simulation synthetic data, ControlNet can provide highly flexible control solutions; high efficiency, that is, when implementing image rendering control, ControlNet has high computational efficiency and performance. It reduces the consumption of computational resources and memory by optimizing the network structure and parameters, improving the rendering efficiency.
[0071] In the technical solution of the embodiments of the present disclosure, the target model includes the SD model and the ControlNet plug-in model. Since both the SD model and the ControlNet plug-in model have powerful capabilities in image rendering, by inputting the accurate auxiliary enhancement information provided by the game engine into the ControlNet plug-in model, it helps to improve the SD model to generate simulation images with perfect visual effects.
[0072] In the embodiments of the present disclosure, inputting the target prompt, the rendered image, and the auxiliary enhancement information into the target model includes: inputting the target prompt and the rendered image into the SD model; inputting the auxiliary enhancement information into the ControlNet plug-in model.
[0073] In practical applications, the auxiliary enhancement information can be input into the ControlNet plug-in model first, and then the target prompt and the rendered image can be input into the SD model; alternatively, the target prompt and the rendered image can be input into the SD model first, and then the auxiliary enhancement information can be input into the ControlNet plug-in model; or the auxiliary enhancement information can be input into the ControlNet plug-in model while the target prompt and the rendered image are input into the SD model. It should be noted that the present disclosure does not limit the execution order of inputting the auxiliary enhancement information into the ControlNet plug-in model and inputting the target prompt and the rendered image into the SD model.
[0074] In some embodiments, if the autonomous driving test project is to test the function of "automatically entering or exiting a highway ramp" of an autonomous driving vehicle under different weather conditions, and the original image is an image of a certain highway ramp taken on a sunny day, then the original image is input into a game engine for rendering to obtain a rendered image and auxiliary enhancement information; target prompt 1 (such as heavy rain weather), target prompt 2 (such as foggy weather), target prompt 3 (such as windy weather), and target prompt 4 (such as sunny weather) are obtained; the auxiliary enhancement information is input into the ControlNet plug-in model, and at the same time, the four target prompts and the rendered image are input into the SD model; the simulation images 1 in heavy rain weather, simulation image 2 in foggy weather, simulation image 3 in windy weather, and simulation image 4 in sunny weather generated by the target model are obtained.
[0075] The embodiments of the present disclosure do not change or limit the specific structure within the SD model and the ControlNet plug-in model.
[0076] The technical solution of the embodiments of the present disclosure inputs the auxiliary enhancement information into each ControlNet plug-in model in the target model; inputs the target prompt and the rendered image into the SD model in the target model to obtain the simulation image output by the target model. In this way, by inputting the auxiliary enhancement information provided by the game engine into the ControlNet plug-in model, since the auxiliary enhancement information provided by the game engine is more accurate than the auxiliary enhancement information provided by the ControlNet plug-in model itself, it helps to achieve precise control of image migration and enhancement, and thus improve the accuracy of autonomous driving simulation testing.
[0077] In the embodiments of the present disclosure, after the auxiliary enhancement information is input into the ControlNet plug-in model, it further includes: masking the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model; or replacing the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model with the auxiliary enhancement information.
[0078] Here, the pre-processor of the ControlNet plug-in model is a pre-processor for generating a scribble or sketch effect. It can convert the input image into a form similar to a scribble or sketch, making the image look more concise and abstract.
[0079] In some embodiments, the initial enhancement information refers to the auxiliary information of the original image output by the pre-processor of the ControlNet plug-in model. The initial enhancement information may include: information such as color, shape, texture, contour, depth, etc.
[0080] In some embodiments, the enhancement information obtained by masking the pre-processor of the ControlNet plug-in model may include: turning off the pre-processor or not triggering the pre-processor.
[0081] Exemplarily, when it is detected that auxiliary enhancement information, such as auxiliary enhancement information 1 and auxiliary enhancement information 2, is input into the ControlNet plug-in model, the pre-processor is turned off; the ControlNet plug-in model inputs the auxiliary enhancement information 1 and the auxiliary enhancement information 2 into the SD model.
[0082] Another example, when it is detected that auxiliary enhancement information, such as auxiliary enhancement information 1 and auxiliary enhancement information 2, is input into the ControlNet plug-in model, the pre-processor is not triggered; the ControlNet plug-in model inputs the auxiliary enhancement information 1 and the auxiliary enhancement information 2 into the SD model.
[0083] In some embodiments, replacing the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model with the auxiliary enhancement information includes: triggering the pre-processor, but using the auxiliary enhancement information to replace the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model.
[0084] Yet another example, when it is detected that auxiliary enhancement information, such as auxiliary enhancement information 1 and auxiliary enhancement information 2, is input into the ControlNet plug-in model, the pre-processor is triggered; the initial enhancement information is obtained through the pre-processor of the ControlNet plug-in model, and the auxiliary enhancement information 1 and the auxiliary enhancement information 2 are used to replace the initial enhancement information; the ControlNet plug-in model inputs the auxiliary enhancement information 1 and the auxiliary enhancement information 2 into the SD model.
[0085] The technical solution of the embodiment of the present disclosure shields the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model; or replaces the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model with the auxiliary enhancement information. In this way, it is ensured that the ControlNet plug-in model in the target model does not use the initial enhancement information obtained by its own preprocessor, but adopts the auxiliary enhancement information obtained based on the game engine. Since the accuracy of the auxiliary enhancement information is better than that of the initial enhancement information, it can improve the accuracy of the simulation images generated by the SD model combined with the ControlNet plug-in model, thereby helping to improve the accuracy of autonomous driving simulation and the safety of autonomous driving.
[0086] Figure 3 The schematic diagram of the determination process of the auxiliary enhancement information is shown, as Figure 3 shown, after the auxiliary enhancement information is input into the ControlNet plug-in model, it further includes: obtaining the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model; combining the initial enhancement information to adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information is adapted to the target style.
[0087] Here, combining the initial enhancement information to adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information is adapted to the target style may include:
[0088] By extracting the features between the initial enhancement information and the target style, and the features between the auxiliary enhancement information and the target style, and recombining these features to adjust the style of the auxiliary enhancement information. This may include extracting features in aspects such as grammar, vocabulary, and sentiment, and recombining and adjusting according to the features of the target style.
[0089] Here, combining the initial enhancement information to adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information is adapted to the target style may include:
[0090] By screening and filtering the parts of the initial enhancement information that do not conform to the target style and the parts of the auxiliary enhancement information that do not conform to the target style, so as to adjust the styles of the initial enhancement information and the auxiliary enhancement information to be adapted to the target style, and combine the content of the initial enhancement information and the auxiliary enhancement information that is adapted to the target style. This can be achieved through techniques such as semantic analysis and sentiment analysis of the auxiliary enhancement information.
[0091] Here, combining the initial enhancement information to adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information is adapted to the target style may include:
[0092] Determine the style conversion technology in combination with the initial enhancement information; use the style conversion technology to convert the style of the auxiliary enhancement information into the target style. This can be achieved through a neural network model, which corresponds and converts the input auxiliary enhancement information and the target style to generate auxiliary enhancement information that conforms to the target style.
[0093] Here, in combination with the initial enhancement information, adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information adapts to the target style, which may include:
[0094] Utilize the existing target style samples and templates to compare and match the initial enhancement information and the auxiliary enhancement information with these samples respectively, thereby adjusting their styles to adapt to the target style, and combine the content in the initial enhancement information and the auxiliary enhancement information that adapts to the target style. This may include using predefined language templates, sentence structures and other reference samples.
[0095] Here, in combination with the initial enhancement information, adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information adapts to the target style, which may include:
[0096] Output the initial enhancement information and the auxiliary enhancement information in the form of an interactive interface;
[0097] Through interaction with the user, adjust and modify the initial enhancement information and the auxiliary enhancement information according to the user's feedback and requirements to meet the user's requirements for the target style; and combine the content in the initial enhancement information and the auxiliary enhancement information that adapts to the target style.
[0098] The above methods can be used alone or in combination, and adjust the style of the auxiliary enhancement information according to specific requirements and scenarios to make it adapt to the target style.
[0099] In the embodiments of the present disclosure, the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model includes initial enhancement information 1 and initial enhancement information 2; when the auxiliary enhancement information, that is, auxiliary enhancement information 1 and auxiliary enhancement information 2, is detected and input into the ControlNet plug-in model, in combination with initial enhancement information 1 and initial enhancement information 2, adjust auxiliary enhancement information 1 and auxiliary enhancement information 2, and after adjustment, obtain auxiliary enhancement information 1' and auxiliary enhancement information 2' that adapt to the target style. The ControlNet plug-in model inputs auxiliary enhancement information 1' and auxiliary enhancement information 2' into the SD model.
[0100] The technical solution of the embodiment of the present disclosure is to obtain the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model; in combination with the initial enhancement information, adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information is adapted to the target style. In this way, the auxiliary enhancement information obtained through the game engine is combined with the initial enhancement information obtained through the pre-processor, so that the adjusted auxiliary enhancement information is more matched with the target style, further improving the accuracy of the auxiliary enhancement information, and thus improving the accuracy of the finally generated simulation image.
[0101] Figure 4 The schematic diagram for obtaining the target prompt is shown, as Figure 4 shown, determining the target prompt matching the target style includes: in response to detecting that there is a prompt in the prompt library that matches the target style, determining the prompt in the prompt library that matches the target style as the target prompt that matches the target style.
[0102] In the embodiment of the present disclosure, the prompt library refers to an instruction library for controlling and guiding the AI model to generate images. The prompt library can help users guide the model to generate images that meet the requirements through natural language descriptions. The prompt library can be constructed according to the image style requirements, and the prompt library includes Prompts for various image styles. Specifically, the prompt library usually contains a series of pre-designed instructions or templates, and these instructions or templates can generate corresponding images according to the text descriptions input by users. Exemplarily, the user can input the text description of "an autonomous driving vehicle driving on a highway", and the prompt library can generate a corresponding image according to this description.
[0103] The technical solution of the embodiment of the present disclosure is to, in response to detecting that there is a prompt in the prompt library that matches the target style, determine the prompt in the prompt library that matches the target style as the target prompt that matches the target style. In this way, the direction of the target style can be macro-controlled through the prompt library, improving the accuracy of the finally generated simulation image, helping to improve the accuracy of autonomous driving simulation tests, and thus improving the safety of autonomous driving vehicles.
[0104] In the embodiment of the present disclosure, determining the target prompt matching the target style includes: obtaining the simulation requirements; in combination with the simulation requirements, generating the target prompt matching the target style.
[0105] In the embodiments of the present disclosure, the simulation requirements may include simulation objects, simulation scenarios, simulation special conditions, etc. The simulation requirements may be generated according to the simulation test project name or according to the simulation test purpose. Exemplarily, if the name of the autonomous driving test project is to test the "automatic entry and exit of highway ramps" function of an autonomous driving vehicle under different weather conditions, the simulation requirement is to solve the problem of the autonomous driving vehicle safely entering and exiting the highway ramp under different weather conditions, driving safely, and thus the target prompt words may include heavy snow, heavy rain, strong wind, hail, sunny day, foggy day, highway ramp scenario, and safe driving. Exemplarily, if the purpose of the autonomous driving test is to accurately identify traffic lights, the simulation requirement is to solve the problem of the autonomous driving vehicle accurately identifying traffic lights, and thus the target prompt words may include autonomous driving vehicle, traffic signal, environmental perception, and traffic light recognition. The above is only an exemplary illustration and does not limit all possible acquisition methods included in the target prompt words. Only, there is no exhaustive list here.
[0106] In the technical solution of the embodiments of the present disclosure, the simulation requirements are obtained; in combination with the simulation requirements, target prompt words matching the target style are generated. In this way, the generated target prompt words can meet both the simulation requirements and the target style, making the finally generated simulation image meet the simulation requirements, which helps to improve the accuracy of autonomous driving simulation.
[0107] The embodiments of the present disclosure provide an image generation device, which is applied to the field of autonomous driving simulation, such as Figure 5 shown, the image generation device may include: a determination module 501, configured to determine an original image and target prompt words matching the target style; a first acquisition module 502, configured to acquire a rendered image of the original image and auxiliary enhancement information obtained through a game engine; an input module 503, configured to input the target prompt words, the rendered image, and the auxiliary enhancement information into a target model; a second acquisition module 504, configured to acquire a simulation image of the target style of the original image generated by the target model.
[0108] In some embodiments, the target model includes an SD model and a ControlNet plug-in model.
[0109] In some embodiments, the input module 503 includes: a first input sub-module, configured to input the target prompt words and the rendered image into the SD model; a second input sub-module, configured to input the auxiliary enhancement information into the ControlNet plug-in model.
[0110] In some embodiments, the input module 503 further includes: a shielding sub-module, configured to shield the initial enhancement information obtained by a pre-processor of the ControlNet plug-in model; a replacement sub-module, configured to use the auxiliary enhancement information to replace the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model.
[0111] In some embodiments, the input module 503 further includes: a first acquisition sub-module, configured to acquire initial enhancement information obtained by a pre-processor of the ControlNet plug-in model; and an adjustment sub-module, configured to adjust the auxiliary enhancement information in combination with the initial enhancement information, so that the adjusted auxiliary enhancement information is adapted to the target style.
[0112] In some embodiments, the determination module 501 includes: a determination sub-module, configured to, in response to detecting that there is a prompt word in the prompt word library that matches the target style, determine the prompt word in the prompt word library that matches the target style as the target prompt word that matches the target style.
[0113] In some embodiments, the determination module 501 includes: a second acquisition sub-module, configured to acquire simulation requirements; and a generation sub-module, configured to generate a target prompt word that matches the target style in combination with the simulation requirements.
[0114] Those skilled in the art should understand that the functions of the various processing modules in the image generation device according to the embodiments of the present disclosure can be understood with reference to the relevant descriptions of the foregoing image generation method. The various processing modules in the image generation device according to the embodiments of the present disclosure can be implemented by a generation circuit that implements the functions of the embodiments of the present disclosure, or can be implemented by software that executes the functions of the embodiments of the present disclosure running on an electronic device.
[0115] The image generation device according to the embodiments of the present disclosure obtains a rendered image of the original image and auxiliary enhancement information through a game engine, and inputs the target prompt word, the rendered image, and the auxiliary enhancement information provided by the game engine to a target model, so as to achieve precise control of image migration and enhancement, to meet the requirements of multi-style image migration of autonomous driving simulation data, and further improve the accuracy of autonomous driving simulation testing.
[0116] The embodiments of the present disclosure provide a schematic diagram of a scenario for image generation, as Figure 6 shown.
[0117] As described above, the image generation method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as in-vehicle digital assistants, in-vehicle telephones, and other similar computing devices.
[0118] Specifically, the electronic device can specifically perform the following operations:
[0119] Determine the original image and the target prompt word that matches the target style;
[0120] Obtain the rendered image of the original image obtained through the game engine and the auxiliary enhancement information;
[0121] Input the target prompt word, the rendered image, and the auxiliary enhancement information into the target model;
[0122] Obtain the simulation image of the target style of the original image generated by the target model.
[0123] Among them, the original image and the target prompt word matching the target style can be obtained from the data source of the autonomous driving vehicle. The data source can be various forms of data storage devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The map data source can also represent various forms of mobile devices, such as in-vehicle digital assistants, in-vehicle telephones, and other similar computing devices. In addition, the data source can be located at the vehicle end or in the cloud.
[0124] It should be understood that Figure 6 The scene graph shown is merely illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 6 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0125] In the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0126] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0127] Figure 7 The schematic block diagram of an example electronic device 700 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0128] As Figure 7As shown, device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 702 or computer programs loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0129] Multiple components in device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0130] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the image generation method. For example, in some embodiments, the image generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the image generation method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the image generation method in any other appropriate way (e.g., by means of firmware).
[0131] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), system on a chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0132] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0133] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0134] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0135] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or in a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0136] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0137] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0138] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An image generation method, applied to the field of autonomous driving simulation, includes: Determine an original image and a target prompt word that matches the target style; Obtain a rendered image of the original image and auxiliary enhancement information obtained through a game engine; Input the target prompt word, the rendered image, and the auxiliary enhancement information into a target model; the target model includes a Stable Diffusion (SD) model and a ControlNet plug-in model; Obtain a simulation image of the target style of the original image generated by the target model; The step of inputting the target prompt word, the rendered image, and the auxiliary enhancement information into the target model includes: Input the target prompt word and the rendered image into the SD model; Input the auxiliary enhancement information into the ControlNet plug-in model; After inputting the auxiliary enhancement information into the ControlNet plug-in model, it further includes: Mask the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model; or Replace the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model with the auxiliary enhancement information; Or After inputting the auxiliary enhancement information into the ControlNet plug-in model, it further includes: Obtain the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model; Combine the initial enhancement information to adjust the auxiliary enhancement information so that the adjusted auxiliary enhancement information adapts to the target style.
2. The method according to claim 1, wherein Determining a target prompt word that matches the target style includes: In response to detecting that there is a prompt word in the prompt word library that matches the target style, determine the prompt word in the prompt word library that matches the target style as the target prompt word that matches the target style.
3. The method according to claim 1, wherein Determining a target prompt word that matches the target style includes: Obtain simulation requirements; Combine the simulation requirements to generate the target prompt word that matches the target style.
4. An image generation device, applied to the field of autonomous driving simulation, includes: A determination module, configured to determine an original image and a target prompt word that matches the target style; A first acquisition module, configured to obtain a rendered image of the original image and auxiliary enhancement information obtained through a game engine; An input module, configured to input the target prompt word, the rendered image, and the auxiliary enhancement information into a target model; The target model includes an SD model and a ControlNet plug-in model; A second acquisition module, configured to obtain a simulation image of the target style of the original image generated by the target model; The input module includes: A first input sub-module, configured to input the target prompt word and the rendered image into the SD model; A second input sub-module, configured to input the auxiliary enhancement information into the ControlNet plug-in model; The input module further includes: A masking sub-module, configured to mask the initial enhancement information obtained by the preprocessor of the ControlNet plug-in model; or A replacement sub-module, configured to replace the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model with the auxiliary enhancement information; Or The input module further includes: A first acquisition sub-module, configured to acquire the initial enhancement information obtained by the pre-processor of the ControlNet plug-in model; An adjustment sub-module, configured to adjust the auxiliary enhancement information in combination with the initial enhancement information, so that the adjusted auxiliary enhancement information is adapted to the target style.
5. The apparatus according to claim 4, wherein The determination module includes: A determination sub-module, configured to, in response to detecting that there is a prompt word in the prompt word library that matches the target style, determine the prompt word in the prompt word library that matches the target style as the target prompt word that matches the target style.
6. The device according to claim 4, wherein The determination module includes: A second acquisition sub-module, configured to acquire the simulation requirements; A generation sub-module, configured to generate the target prompt word that matches the target style in combination with the simulation requirements.
7. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the method according to any one of claims 1-3.
8. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause a computer to execute the method according to any one of claims 1-3.
9. A computer program product, comprising a computer program stored on a storage medium, where the computer program, when executed by a processor, implements the method according to any one of claims 1-3.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN116993902A
Image processing method and device, electronic equipment and storage medium
CN117252791A