Image generation method based on simulation and electronic device

By extracting features and generating diffusion data from simulated images using a target restoration model, the problem of insufficient data for autonomous vehicles in extreme scenarios is solved, generating high-quality, realistic restored images and improving environmental perception and safety.

CN122391438APending Publication Date: 2026-07-14EACON TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EACON TECHNOLOGY CO LTD
Filing Date
2026-03-03
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently generate high-quality, spatiotemporally consistent long-tail scenario data, limiting the generalization ability of autonomous vehicles in extreme interaction scenarios and extreme weather conditions, thus affecting driving safety and system reliability.

Method used

By extracting features from the simulation image and the foreground mask image through the encoding layer of the target restoration model, and combining scene condition features, a diffusion generation layer is used to generate a realistic restoration image, thereby achieving targeted restoration of the target object and improving visual realism.

Benefits of technology

The generated realistic restored images can significantly improve the environmental perception robustness and safety of autonomous driving systems in extreme scenarios, and solve the problems of inconsistency between foreground and background and low realism in simulation images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391438A_ABST
    Figure CN122391438A_ABST
Patent Text Reader

Abstract

The disclosure provides a simulation-based image generation method and an electronic device, and relates to the technical field of automatic driving and unmanned vehicles. The simulation-based image generation method comprises the following steps: obtaining simulation image data in a target scene, and inputting the simulation image data into a target repair model; the simulation image data comprises a simulation image, a background image corresponding to the simulation image, and a foreground mask image corresponding to the simulation image; through an encoding layer of the target repair model, feature extraction is performed on the simulation image and the foreground mask image to obtain image features, and feature extraction is performed on the background image to obtain scene condition features; through a diffusion generation layer of the target repair model, a real repair image corresponding to the simulation image is generated by taking the image features as input and taking the scene condition features as diffusion conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of autonomous driving and driverless vehicle technology, specifically to a simulation-based image generation method and electronic device. Background Technology

[0002] High-quality training data is crucial for improving the environmental perception capabilities of autonomous vehicles. However, it is difficult to obtain sufficient data for long-tail scenarios such as extreme interaction scenarios and extreme weather through actual collection. This limits the generalization ability of autonomous vehicles when dealing with such scenarios, affecting driving safety and system reliability.

[0003] Currently, the industry uses image simulation synthesis methods to generate long-tail scene data. However, these methods still suffer from problems such as slow rendering speed, insufficient image realism, and difficulty in ensuring spatiotemporal consistency. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a simulation-based image generation method and an electronic device.

[0005] Firstly, a simulation-based image generation method is provided, comprising: acquiring simulation image data of a target scene and inputting the simulation image data into a target restoration model, wherein the simulation image data includes a simulation image, and a background image and a foreground mask image corresponding to the simulation image; extracting features from the simulation image and the foreground mask image through the encoding layer of the target restoration model to obtain image features, and extracting features from the background image to obtain scene condition features; and generating a real restored image corresponding to the simulation image through the diffusion generation layer of the target restoration model, using the image features as input and the scene condition features as diffusion conditions.

[0006] In conjunction with the first aspect, in some implementations of the first aspect, the encoding layer includes a first encoding module and a second encoding module. Through the encoding layer of the target restoration model, feature extraction is performed on the simulation image and the foreground mask image to obtain image features, including: encoding the simulation image through the first encoding module to obtain a first image sub-feature corresponding to the simulation image; encoding the foreground mask image through the second encoding module to obtain a second image sub-feature corresponding to the foreground mask image; and concatenating the first image sub-feature and the second image sub-feature to obtain the image features.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, the diffusion generation layer includes a prediction module and a decoding module. Through the diffusion generation layer of the target restoration model, image features are used as input, and scene condition features are used as diffusion conditions to generate a realistic restored image corresponding to the simulated image. This includes: using the prediction module, based on the image features and scene condition features, predicting the overall features, foreground image features, and foreground mask features corresponding to the realistic restored image; fusing the overall features, foreground image features, and foreground mask features to obtain fused features; and decoding the fused features through the decoding module to obtain the realistic restored image.

[0008] In conjunction with the first aspect, in some implementations of the first aspect, the target restoration model is trained by the following method: acquiring training data pairs, which include simulated sample images and real sample images, wherein the simulated sample images have non-realistic visual features; determining the simulated background image and simulated foreground mask image corresponding to the simulated sample image; extracting features from the simulated sample image and simulated foreground mask image through the encoding layer of the target restoration model to obtain sample image features, and sampling the simulated background image to obtain sample scene condition features; generating sample restoration image features by taking the sample image features as input and the sample scene condition features as diffusion conditions through the diffusion generation layer of the target restoration model; encoding the real sample image through the target encoding module to obtain real sample image features; calculating the loss value based on the difference between the sample restoration image features and the real sample image features, and adjusting the encoding layer and diffusion generation layer of the target restoration model based on the loss value.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the loss value is calculated based on the difference between the sample-restored image features and the real sample image features, including: determining the overall prediction loss, foreground image prediction loss, and foreground mask prediction loss between the sample-restored image features and the real sample image features, respectively; and weighting and summing the overall prediction loss, foreground image prediction loss, and foreground mask prediction loss to obtain the loss value.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, both the sample restoration image features and the real sample image features are feature representations of the latent feature space. A weighted sum of the overall prediction loss, foreground image prediction loss, and foreground mask prediction loss is performed to obtain the loss value, including: determining the divergence loss, which measures the difference between the latent variable probability distribution output by the diffusion generation layer and the latent variable probability distribution output by the target encoding module; and performing a weighted sum of the divergence loss, overall prediction loss, foreground image prediction loss, and foreground mask prediction loss to obtain the loss value. Optionally, the weight of the overall prediction loss is greater than the weight of the foreground image prediction loss and the weight of the foreground mask prediction loss, and both the weight of the foreground image prediction loss and the weight of the foreground mask prediction loss are greater than the weight of the divergence loss.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, image generation further includes: acquiring a real sample image and semantic annotation information corresponding to the real sample image; generating a real foreground mask image corresponding to the real sample image based on the semantic annotation information; segmenting the real sample image into a real foreground image and a real background image based on the real foreground mask image; performing non-realistic transformation processing on the real foreground image, and fusing the non-realistic transformed real foreground image with the real background image to generate a simulated sample image. Optionally, the non-realistic transformation processing includes at least one of brightness adjustment, contrast adjustment, blurring, and geometric transformation.

[0012] In conjunction with the first aspect, in some implementations of the first aspect, the simulation image data includes multiple simulation images with spatiotemporal consistency. Acquiring simulation image data of the target scene includes: acquiring a background image of the target scene; extracting a 3D model of the target object from a 3D model library, and performing multi-view rendering on the 3D model of the target object to obtain multiple foreground target images from different viewpoints and foreground mask images corresponding to each of the multiple foreground target images; using the foreground mask image corresponding to each foreground target image, fusing the foreground target image with the background image to generate a simulation image from the viewpoint of the foreground target image. Optionally, the target object includes dynamic targets and / or static targets; optionally, the target scene includes a mining scene; optionally, the target object includes at least one of vehicles, pedestrians, crushing stations, and tunnels.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the simulation image data includes multiple simulation images with spatiotemporal consistency. Acquiring simulation image data of the target scene includes: acquiring multiple initial images from different perspectives of the target scene; performing 3D reconstruction on the multiple initial images to obtain a 3D model of the target scene; extracting at least one target object from the 3D model; editing the target scene by changing the pose of the vehicle and / or the pose of the target object, and performing multi-view rendering on the edited target scene to generate multiple simulation images.

[0014] In a second aspect, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the method provided in the first aspect by executing the executable instructions.

[0015] The simulation-based image generation method disclosed herein, through structured input of simulation images, background images, and foreground mask images, enables the target inpainting model to accurately distinguish and utilize scene context and target priors to achieve targeted inpainting of the target object, effectively solving the problems of inconsistency between foreground and background and low realism in simulation images. Secondly, by using scene condition features as diffusion conditions for the diffusion generation layer, the visual realism of the target object can be improved end-to-end while maintaining the original background environment features. Attached Figure Description

[0016] Figure 1 The diagram shown is a system architecture diagram of a simulation-based image generation method provided in an embodiment of this disclosure.

[0017] Figure 2 The diagram shown is a flowchart of a simulation-based image generation method provided in an embodiment of this disclosure.

[0018] Figure 3 The diagram shown is a structural schematic of a target repair model provided in an embodiment of this disclosure.

[0019] Figure 4 The diagram shown is a flowchart illustrating the steps of extracting features from a simulation image and a foreground mask image through the coding layer of a target repair model, according to an embodiment of this disclosure, to obtain image features.

[0020] Figure 5 The diagram shows a flowchart illustrating the steps of generating a real restored image corresponding to a simulated image through a diffusion generation layer of a target restoration model, using image features as input and scene condition features as diffusion conditions, according to an embodiment of this disclosure.

[0021] Figure 6 The diagram shown is a flowchart illustrating the steps for acquiring simulated image data of a target scene according to an embodiment of this disclosure.

[0022] Figure 7 The diagram shown is a flowchart illustrating the steps for obtaining simulated image data of a target scene according to another embodiment of this disclosure.

[0023] Figure 8 The diagram shown is a flowchart illustrating a training method for a target repair model provided in an embodiment of this disclosure.

[0024] Figure 9 The diagram shown is a schematic diagram of the structure of the target repair model provided in an embodiment of this disclosure during the training phase.

[0025] Figure 10 The diagram shown is a flowchart illustrating the steps of calculating the loss value based on the difference between the features of the repaired image and the features of the real sample image, according to an embodiment of this disclosure.

[0026] Figure 11 The diagram shown is a flowchart illustrating the steps of weighted summation of the overall prediction loss, the foreground image prediction loss, and the foreground mask prediction loss to obtain the loss value, according to an embodiment of this disclosure.

[0027] Figure 12 The diagram shown is a flowchart illustrating the generation process of training data pairs provided in an embodiment of this disclosure.

[0028] Figure 13 The diagram shown is a schematic diagram of the structure of a simulation-based image generation device provided in an embodiment of this disclosure.

[0029] Figure 14 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0031] In autonomous vehicles, the perception model of the autonomous driving system can perform perception modeling of the surrounding environment, 3D object detection, and semantic segmentation, providing fundamental support for subsequent path planning and decision-making control. It is evident that the perception accuracy and robustness of the perception model directly determine the safety and practicality of autonomous vehicles. However, data collection in long-tail scenarios such as extreme weather and extreme interaction scenarios faces significant challenges, limiting the collection of large-scale, diverse data.

[0032] On the one hand, abnormal interaction scenarios (such as heavy mining vehicles overturning out of control, parking in non-standard postures, or sudden intrusion by personnel or animals) and extreme weather scenarios (such as insufficient nighttime lighting, dust, rain, snow, or fog) are difficult to reproduce, and it is also difficult to collect a large amount of data suitable for training perception models under extreme weather conditions. On the other hand, in new operational scenarios, perception models often perform poorly on new, untrained scenario data, and manually labeling the new scenario data is a very lengthy process. For example, in mining scenarios, there may be unconventional obstacles such as temporary sheds, slag falling from mining trucks, ruts formed by mining trucks, and crushing stations.

[0033] Currently, commonly used data simulation methods for long-tail scene data include Neural Radiance Fields (NeRF), 3D Gaussians (3DGS), and WorldModel methods. However, these methods all have certain limitations in practical applications. For example, while NeRF can generate high-quality images from new perspectives, its rendering speed is slow, making it difficult to apply to scenes requiring real-time data generation. 3DGS uses a large number of Gaussian spheres to represent the scene, each with position, orientation, color, and transparency parameters, and generates images through projection rendering. While this improves rendering efficiency, it is more suitable for rendering static scenes and has weak modeling capabilities for dynamic targets (such as pedestrians and vehicles). Furthermore, images generated from new perspectives are prone to distortion. WorldModel methods can generalize to more long-tail scene data through prompts, but when faced with the joint generation of multi-sensor data or time-series data, it often struggles to guarantee spatiotemporal consistency between frames and perspectives, and also suffers from slow generation speed. These limitations make it difficult for existing technologies to efficiently and effectively produce long-tail image data with spatiotemporal consistency, thus restricting further improvements in the performance of autonomous driving systems.

[0034] To address the aforementioned technical problems, this disclosure provides a simulation-based image generation method, comprising: acquiring simulation image data of a target scene and inputting the simulation image data into a target restoration model, wherein the simulation image data includes a simulation image and a corresponding background image and foreground image; extracting features from the simulation image and foreground image through the encoding layer of the target restoration model to obtain image features, and extracting features from the background image to obtain scene condition features; and generating a real restored image corresponding to the simulation image through the diffusion generation layer of the target restoration model, using the image features as input and the scene condition features as diffusion conditions, thereby generating long-tail image data of a specific target scene efficiently and with high quality.

[0035] The following is a combination of... Figure 1 This disclosure introduces a system architecture provided by one embodiment.

[0036] Figure 1 The diagram shown is a system architecture schematic of a simulation-based image generation method provided in an embodiment of this disclosure. Figure 1 As shown, the system architecture includes a terminal 110 and a server 120. The terminal 110 and the server 120 are communicatively connected, for example, via a wired or wireless network.

[0037] Terminal 110 may include electronic devices such as vehicles, mobile phones, laptops, or personal computers, or it may include vehicle-mounted or fixed data acquisition devices. Terminal 110 is used for data acquisition and preprocessing operations in the target scene. For example, terminal 110 can acquire raw image data from multiple perspectives in the target scene and generate simulated image data of the target scene based on the raw image data; or, the user can directly upload the simulated image data to terminal 110 through the interactive interface of terminal 110. Then, terminal 110 sends the simulated image data to server 120.

[0038] Server 120 receives simulated image data from terminal 110, processes the simulated image data, and outputs a highly realistic restored image corresponding to the simulated image data. Finally, the result is sent back to terminal 110 for user use. Server 120 can be a single server, a server cluster composed of multiple servers, or a cloud server; alternatively, server 120 can be integrated into terminal 110.

[0039] Those skilled in the art will know that Figure 1 The number of terminals and servers shown is merely illustrative. Depending on actual needs, there may be any number of terminals and servers, and this disclosure does not impose any limitation on this.

[0040] The following is combined Figures 2 to 12 This disclosure describes in detail the simulation-based image generation method provided in the embodiments of this disclosure.

[0041] Figure 2 The diagram shows a flowchart of a simulation-based image generation method according to an embodiment of this disclosure. Exemplarily, the simulation-based image generation method provided in this embodiment consists of... Figure 1 The server 120 shown is executing. (As shown) Figure 2 As shown in the embodiments of this disclosure, the simulation-based image generation method includes the following steps.

[0042] S210: Acquire simulation image data of the target scene and input the simulation image data into the target repair model.

[0043] The target scene refers to the specific environment in which image data is to be generated. For example, target scenes include extreme weather scenes such as insufficient lighting at night, dust, or rain, snow, and fog, as well as non-standard road scenes such as mining scenes and construction scenes. Alternatively, target scenes may include abnormal vehicle postures (such as overturning or getting stuck in a deep pit) that are difficult to actually capture, or abnormal interaction scenes such as people or animals suddenly entering the scene.

[0044] The simulation image data includes the simulation image, as well as the corresponding background image and foreground mask image.

[0045] Optionally, the simulated image is generated by adding long-tailed target objects to the background image acquired from the actual vehicle using image simulation synthesis methods such as 3D target rendering and image fusion. The target objects include dynamic targets and / or static targets; optionally, dynamic targets include vehicles and pedestrians, and static targets include crushing stations, tunnels, etc.

[0046] The background image corresponding to the simulation image is the background portion of the simulation image excluding the target object. The foreground mask image corresponding to the simulation image is a mask representation used to identify the foreground image in the simulation image; the foreground image is the portion of the image corresponding to the target object. The foreground mask image can be represented as a binary image, a probability image, etc., where different pixel values ​​are used to distinguish the foreground image from the background image. Taking a simulation image corresponding to a mining scene as an example, the foreground mask image can use white to mark the area where the target object (such as a crushing station, mining vehicles, etc.) is located, and black to mark the background area.

[0047] Optionally, the simulation image data includes multiple simulation images with spatiotemporal consistency. For example, the multiple simulation images can be multi-view simulation images corresponding to different vehicle poses in the same scene, where the multi-view simulation images collectively cover the complete spatial range of the target scene and maintain geometric and semantic consistency in overlapping areas. Alternatively, the multiple simulation images can also be an image sequence simulating the motion process of a vehicle or target object in a scene over a continuous time period, where multiple frames in the image sequence are temporally coherent. Accordingly, the simulation image data also includes the background image and foreground mask image corresponding to each simulation image.

[0048] Since simulated images are not actual, collected data, they may suffer from poor realism. Therefore, it is necessary to input the simulated image data into the target restoration model to perform realistic restoration.

[0049] The target repair model consists of an encoding layer and a diffusion generation layer, which will be introduced below.

[0050] S220 extracts features from the simulation image and the foreground mask image through the coding layer of the target repair model to obtain image features, and extracts features from the background image to obtain scene condition features.

[0051] The encoding layer is used to map the input image into a semantic feature vector. Specifically, through the encoding layer, feature extraction and feature fusion are performed on the simulation image and the foreground mask image, respectively, to obtain a comprehensive feature representation of the overall visual and semantic information of the simulation image, as well as the location and structural information of the target object.

[0052] Through the encoding layer, features are extracted from the background image to obtain scene condition features. Scene condition features are feature representations of the global contextual information (such as lighting, hue, weather, etc.) of the background environment in the simulated image. These scene condition features will serve as conditional information to guide the image restoration process.

[0053] S230 uses the diffusion generation layer of the target restoration model as input, takes image features as input, and takes scene condition features as diffusion conditions to generate a real restored image corresponding to the simulation image.

[0054] The diffusion generation layer takes image features as input and scene condition features as diffusion conditions to predict the target object in the image features, obtaining predicted features. Then, the predicted features are decoded to obtain the realistic restored image. In the realistic restored image, the foreground image corresponding to the target object has high realism.

[0055] During this process, the diffusion generation layer can simultaneously repair the target object based on the constraints of visual and structural details and scene conditions in the simulated image. This helps to maintain the coordination and consistency between the target object and the background environment in terms of lighting, tone, and weather, thereby improving the scene rationality of the realistically repaired image.

[0056] Optionally, the diffusion generation layer can employ a diffusion model architecture to predict the target object. The core of this architecture is a denoising network with an encoder-decoder structure. During inference, image features are used as the starting point for generation, scene condition features are used as diffusion conditions, and spatial attention mechanisms are injected into each layer of the denoising network to guide the generation process. This allows global contextual information from the background environment to dynamically guide the prediction and generation of details of the target object, improving the realism of the target object and ultimately outputting predicted features consistent with the dimensions of the image features.

[0057] In this embodiment, through structured input of simulated images, background images, and foreground mask images, the target inpainting model can accurately distinguish and utilize scene context and target priors to achieve targeted inpainting of the target object, effectively solving the problems of inconsistency between foreground and background and low realism in simulated images. Secondly, by using scene condition features as diffusion conditions for the diffusion generation layer, the visual realism of the target object can be improved end-to-end while maintaining the original background environment features.

[0058] Furthermore, the aforementioned simulation-based image generation method can be used to generate long-tail scene data. Its output, highly realistic restored images can be directly used in the training process of perception models in autonomous driving systems, thereby significantly improving the environmental perception robustness and safety of autonomous driving systems under extreme and rare conditions, and has important engineering application value.

[0059] The following section will introduce the specific structure of the coding layer in the target restoration model, as well as the process of encoding the simulated image data.

[0060] Figure 3 The diagram shown is a structural schematic of a target repair model provided in an embodiment of this disclosure. Figure 3 As shown, the coding layer includes a first coding module E1 and a second coding module E2.

[0061] Figure 4 The diagram shown is a flowchart illustrating the steps of extracting features from a simulated image and a foreground mask image using the encoding layer of a target restoration model, according to an embodiment of this disclosure. Figure 4 As shown in this embodiment, the steps for extracting features from the simulation image and the foreground mask image through the coding layer of the target repair model to obtain image features include the following steps.

[0062] S410, the simulation image is encoded by the first encoding module to obtain the first image sub-feature corresponding to the simulation image.

[0063] The first encoding module is a branch network in the encoding layer used to process the simulated image. It can adopt an autoencoder or a convolutional neural network structure and is responsible for encoding the target object and background environment in the simulated image and performing downsampling operations to obtain the first image sub-features. The first image sub-features are the feature representations obtained after the simulated image is encoded by the first encoding module.

[0064] Optionally, the first encoding module employs a variational autoencoder (VAE). A VAE maps the simulated image to a lower-dimensional latent feature space to capture meaningful information contained within the simulated image. This allows subsequent processing to focus on the meaningful information contained in the simulated image, improving the efficiency and accuracy of subsequent processing. Furthermore, compared to other autoencoders, the VAE does not encode the input data as discrete points in the latent feature space, but rather as a continuous range of possibilities represented by a probability distribution.

[0065] S420, the foreground mask image is encoded by the second encoding module to obtain the second image sub-feature corresponding to the foreground mask image.

[0066] The second encoding module is a branch network in the encoding layer used to process the foreground mask image. The structure and specific parameters of the second encoding module may be the same as or different from those of the first encoding module. The second image sub-feature is the feature representation obtained after the foreground mask image is encoded by the second encoding module.

[0067] Since the foreground mask image provides the precise location and contour information of the target object in the simulation image, the second encoding module can process this information to generate features that emphasize the target structure (such as boundaries and shape) and spatial location, providing key structural constraints for subsequent processing.

[0068] Optionally, the second encoding module also employs a variational autoencoder.

[0069] S430: The first image sub-feature and the second image sub-feature are concatenated to obtain the image features.

[0070] By concatenating the first image sub-feature carrying visual information with the second image sub-feature carrying structural information along the channel dimension, a more comprehensive image feature with both visual and structural priors can be formed. This provides the diffusion generation layer with reference information on the visual attributes of the target object, as well as constraints on the spatial extent and geometric shape of the target object.

[0071] In this embodiment, independent first and second encoding modules are used to extract visual information from the simulated image and structural information from the target object, respectively, avoiding feature confusion and interference that may occur with a single encoding module. Furthermore, by performing feature stitching along the channel dimension, a correlation is established between the sub-features of the first and second images, thereby providing prior knowledge about the visual and structural aspects of the target object for the subsequent restoration process. This effectively guides the diffusion generation layer to perform detail reconstruction while adhering to geometric constraints, enhancing the visual realism of the restored image.

[0072] Continue to refer to Figure 3In some embodiments, the coding layer further includes a third coding module E3, which is used to extract features from the background image to obtain scene condition features.

[0073] Specifically, the background image carries environmental context information such as global illumination, overall color tone, weather atmosphere, material texture, and spatial layout of the target scene. Scene conditional features are feature representations of the background environmental context information. As conditional embeddings, they are used to guide and constrain the feature visual style and environmental attributes of the prediction process in the diffusion generation layer, thereby improving the realism of the prediction results.

[0074] The above section introduced the structure of the coding layer in the target inpainting model. The following section will describe the structure of the diffusion generation layer and the process of generating a realistic inpainted image. (Continue to refer to...) Figure 3 The diffusion generation layer includes a prediction module D1 and a decoding module D2.

[0075] Figure 5 The diagram illustrates a flowchart of an embodiment of this disclosure, showing how a diffusion generation layer of a target restoration model generates a realistic restored image corresponding to a simulated image by taking image features as input and scene condition features as diffusion conditions. Figure 5 As shown in this embodiment, the steps for generating a real restored image corresponding to the simulated image by using the diffusion generation layer of the target restoration model as input and scene condition features as diffusion conditions include the following steps.

[0076] S510, through the prediction module, predicts the overall features, foreground image features, and foreground mask features corresponding to the real restored image based on image features and scene condition features.

[0077] The prediction module can adopt a DNet architecture, including a shrinking path (encoder) and an expanding path (decoder). The shrinking path consists of a series of convolutional layers, activation functions, and pooling layers. The expanding path consists of a series of upsampling layers, convolutional layers, and activation functions. Image features and scene condition features are injected into multiple layers of the prediction module through a spatial attention mechanism to guide the prediction process.

[0078] The output layer of the prediction module uses three independent prediction heads: an overall feature prediction head, a foreground image feature prediction head, and a foreground mask feature prediction head. These will be described in detail below.

[0079] The holistic feature prediction head is used to predict global features, including both the target object and the background environment, to obtain holistic features. Holistic features contain complete visual information fused from the target object and the background environment, focusing on expressing the overall semantics of the target scene.

[0080] The foreground image feature prediction head is used to predict the features of the target object, resulting in foreground image features. These foreground image features contain visual information about the target object, ensuring its visual realism.

[0081] The foreground mask feature prediction head is used to predict the precise shape and spatial location of a target object, resulting in foreground mask features. These features contain structural information such as the geometric contours and location of the target object, providing explicit structural constraints.

[0082] S520 fuses the overall features, foreground image features, and foreground mask features to obtain fused features.

[0083] The overall features, foreground image features, and foreground mask features are fused into a unified feature representation, i.e., fused features. Optionally, the overall features, foreground image features, and foreground mask features can be concatenated to obtain fused features.

[0084] The S530 decodes the fused features through a decoding module to obtain a realistic restored image.

[0085] The decoding module is used to decode the fused features. Through multi-level upsampling operations, the fused features are decoded into a real restored image with the same size as the simulation image.

[0086] In this embodiment of the disclosure, a multi-head prediction architecture is used to achieve decoupled prediction of overall features, foreground image, and foreground mask. This allows the prediction process to focus on different dimensions, which not only improves the structural accuracy and appearance realism of the target object, but also makes it highly integrated with the background environment, significantly improving the visual realism and scene rationality of the real restored image.

[0087] The following is combined Figure 6 , Figure 7 This paper introduces a method for generating simulated image data using image simulation synthesis. Optionally, the simulated image data includes multiple simulated images with spatiotemporal consistency.

[0088] Figure 6 The diagram shown is a flowchart illustrating the steps for acquiring simulated image data of a target scene according to an embodiment of this disclosure. Figure 6 As shown in the embodiments of this disclosure, the steps for obtaining simulation image data in the target scene include the following steps.

[0089] S610, acquires background images of the target scene.

[0090] Background images are images that do not contain the target object in the target scene, and are used to provide a realistic scene context in the target scene.

[0091] The background image can be a single still image or a sequence of multiple frames extracted from a continuous video stream. It can be acquired by a fixed camera or vehicle-mounted capture device deployed at the target scene, or extracted from video captured from a real vehicle.

[0092] S620 extracts the 3D model of the target object from the 3D model library and performs multi-view rendering on the 3D model of the target object to obtain multiple foreground target images from different viewpoints and the foreground mask images corresponding to each of the multiple foreground target images.

[0093] A 3D model library is a digital asset database formed by collecting the required target objects using 3D reconstruction equipment and performing 3D reconstruction on each target object. Optionally, the target scene includes a mining scene; in the mining scene, the target objects include dynamic targets and / or static targets, wherein dynamic targets include vehicles and pedestrians, and static targets include crushing stations and tunnels.

[0094] Multiple virtual camera poses are set around the 3D model of the target object. Using a computer graphics rendering pipeline, the 3D model of the target object is rendered from multiple perspectives, generating multiple 2D foreground target images from different viewpoints. A foreground mask image is also generated for each foreground target image; the foreground mask image is used to identify the precise contours and pixel positions of the target object.

[0095] S630 uses the foreground mask image corresponding to each foreground target image to fuse the foreground target image with the background image to generate a simulation image from the perspective of the foreground target image.

[0096] For each foreground target image, based on its corresponding foreground mask image, the foreground target image is fused into a specified area of ​​the background image. The rendered virtual foreground and the real background are then combined pixel-level to obtain a simulated image from the perspective of the foreground target image.

[0097] Since the foreground target image is obtained through multi-view rendering and the background image remains unchanged during the fusion process, the spatiotemporal consistency of the simulation image under multiple views is guaranteed.

[0098] In this embodiment, multiple foreground target images with consistency in geometric structure, texture, and scale are obtained by rendering the 3D model of the target object from multiple perspectives. The foreground target images are then fused with the realistically acquired background images, preserving the high realism of the background environment while maintaining spatiotemporal consistency in the perspective dimension, thus efficiently generating simulation image data with spatiotemporal consistency.

[0099] However, since the target objects in the 3D model library and the acquired background images are collected separately, differences in factors such as lighting, acquisition time, weather conditions, and acquisition equipment can lead to significant differences between the foreground and background in the synthesized simulation images. Directly using these simulation images to train the perception model of an autonomous driving system can easily degrade the perception model. Therefore, it is necessary to perform realism restoration through a target restoration model.

[0100] The following section introduces another method for generating simulated image data.

[0101] Figure 7 The diagram shown is a flowchart illustrating the steps for acquiring simulated image data of a target scene according to another embodiment of this disclosure. Figure 7 As shown in the embodiments of this disclosure, the steps for obtaining simulation image data in the target scene include the following steps.

[0102] The S710 acquires multiple initial images from different perspectives of the target scene.

[0103] Multiple initial images are a collection of images captured from multiple different spatial locations and angles within the target scene using a fixed camera or vehicle-mounted acquisition device. The target scene includes the target object.

[0104] In addition, the initial image can be further annotated with 3D targets to obtain the annotation information corresponding to the target objects, including target detection boxes, target categories, etc.

[0105] The S720 performs 3D reconstruction on multiple initial images to obtain a 3D model of the target scene.

[0106] In this step, multiple initial images are used to recover the three-dimensional geometric structure and surface texture information of the target scene through three-dimensional reconstruction technology, thereby obtaining a three-dimensional model of the target scene.

[0107] Optionally, a 3DGS model can be used to realize the three-dimensional reconstruction process.

[0108] S730, extract at least one target object from a 3D model.

[0109] By utilizing the annotation information corresponding to the target object, the target object is extracted from the 3D model, decoupled from the overall scene, and made into an editable entity with independent pose parameters.

[0110] The S740 edits the target scene by changing the pose of the vehicle and / or the pose of the target object, and then performs multi-view rendering on the edited target scene to generate multiple simulation images.

[0111] The virtual camera's pose refers to its position and orientation within a 3D scene, used to simulate changes in the camera's viewpoint during real-world data acquisition. The target object's pose refers to the position, rotation angle, and other parameters of the 3D target model extracted in the previous steps within the target scene.

[0112] By changing the vehicle's pose, the effect of a data acquisition device observing the same scene from different positions and angles can be simulated. By changing the pose of the target object, the movement of the target object itself (such as vehicle driving or turning), abnormal postures (such as rollover or stagnation) or changes in spatial layout can be simulated.

[0113] After the target scene is edited, it is rendered from multiple perspectives to generate simulated images from multiple perspectives.

[0114] Since the simulation images from multiple perspectives are all rendered from the same 3D scene model, the multiple simulation images have spatiotemporal consistency in terms of geometric structure, spatial relationships, and target morphology. Furthermore, through flexible editing of the vehicle pose and target object pose, diverse long-tail image data of the target scene can be generated efficiently.

[0115] However, in the above embodiments, when the pose of the vehicle or the target object is changed, the scene needs to be re-rendered. During this process, distortion problems such as geometric distortion, artifacts, and loss of detail are likely to occur. Therefore, it is also necessary to use a target restoration model to restore its realism.

[0116] The above embodiments introduced the synthesis of simulated image data and the specific implementation method of performing realistic restoration of simulated image data based on the target restoration model. The training process of the target restoration model will be described below.

[0117] Figure 8 The diagram shown is a flowchart illustrating a training method for a target repair model provided in an embodiment of this disclosure. Figure 9 The diagram shown is a schematic diagram of the structure of the target repair model provided in an embodiment of this disclosure during the training phase.

[0118] like Figure 8 As shown, the training process of the target repair model includes the following steps.

[0119] S810 acquires training data pairs, which include simulated sample images and real sample images.

[0120] Simulated sample images have non-realistic visual characteristics, while real sample images are highly realistic images whose content is consistent with that of simulated sample images.

[0121] Specifically, both simulated and real sample images include the target object. In the simulated sample images, the target object exhibits non-realistic visual features, including visual inconsistencies between the target object and the background environment, or the presence of artifacts or blurred textures. In the real sample images, the target object demonstrates high visual realism and scene plausibility.

[0122] S820, determine the simulation background image and simulation foreground mask image corresponding to the simulation sample image.

[0123] The simulation background image is the background portion isolated from the simulation sample image, excluding the target object. The simulation foreground mask image is a mask representation used to identify the foreground image in the simulation sample image.

[0124] S830 extracts features from the simulated sample image and the simulated foreground mask image through the coding layer of the target repair model to obtain sample image features, and samples the simulated background image to obtain sample scene condition features.

[0125] The encoding layer includes a first encoding module, a second encoding module, and a third encoding module.

[0126] The first and second encoding modules encode the simulated sample image and the simulated foreground mask image, respectively, and then fuse the encoding results to obtain the sample image features. The sample image features are a comprehensive feature representation of the non-real-time visual information of the simulated sample image, as well as the position and structural information of the target object.

[0127] The third encoding module samples the simulated background image to obtain sample scene condition features. These sample scene condition features are feature representations of the global contextual information of the background environment in the simulated sample image.

[0128] The specific implementation of this step can be found in the above embodiments, and will not be repeated here.

[0129] S840 uses the diffusion generation layer of the target restoration model as input, takes the sample image features as input, and takes the sample scene condition features as diffusion conditions to generate sample restoration image features.

[0130] During the training phase, the diffusion generation layer only includes a prediction module. Sample image features and sample scene condition features are input into the prediction module, which outputs sample restoration image features. These features consist of the overall features predicted by the prediction module, foreground image features, and foreground mask features.

[0131] The S850 encodes real sample images through a target encoding module to obtain real sample image features.

[0132] contrast Figure 3 , Figure 9 The target encoding module E4 exists only during the training phase, and during training, it adopts a frozen paradigm and does not update gradients with training. Optionally, the target encoding module uses a variational autoencoder.

[0133] The target encoding module encodes the real sample images to obtain the features of the real sample images.

[0134] S860 calculates the loss value based on the difference between the features of the sample-restored image and the features of the real sample image, and adjusts the encoding layer and diffusion generation layer of the target restoration model based on the loss value.

[0135] After obtaining the features of the sample inpainted image and the real sample image, the loss values ​​of both are calculated according to a preset loss function formula. Subsequently, the parameters of the encoding layer and the diffusion generation layer in the target inpainting model are optimized based on the loss values.

[0136] By iteratively executing the above steps, the parameters of the target restoration model are continuously adjusted, gradually reducing the difference between the features of the restored sample image and the features of the real sample image, thereby gradually improving the restoration capability of the input image.

[0137] This disclosure provides a training method for a target restoration model. At the feature level, the difference between the features of the restored sample image and the features of the real sample image is calculated to drive training, enabling the target restoration model to converge quickly and stably. Through an end-to-end training framework, the encoding layer and the diffusion generation layer can collaboratively learn the mapping relationship from low-fidelity images to high-fidelity images.

[0138] The following section introduces a specific implementation method for the model training process.

[0139] Figure 10 The diagram shown is a flowchart illustrating the steps of calculating the loss value based on the difference between the features of the restored sample image and the features of the real sample image, according to an embodiment of this disclosure. Figure 10 As shown in this embodiment, the step of calculating the loss value based on the difference between the features of the sample-restored image and the features of the real sample image includes the following steps.

[0140] S1010, respectively determine the overall prediction loss, foreground image prediction loss, and foreground mask prediction loss between the sample restored image features and the real sample image features.

[0141] During the training phase, a joint loss function is used to comprehensively and accurately guide the model optimization process through multi-dimensional and decoupled loss calculation.

[0142] Specifically, the joint loss function includes the overall prediction loss, the foreground image prediction loss, and the foreground mask prediction loss.

[0143] The overall prediction loss characterizes the difference between the overall feature portion of the sample inpainted image and the features of the real sample image. The overall prediction loss is used to evaluate the prediction accuracy of the target inpainting model in terms of global consistency and semantic coherence.

[0144] The foreground image prediction loss characterizes the difference between the foreground image features in the sample inpainted image and the corresponding foreground features in the ground truth sample image. The foreground image prediction loss is used to evaluate the accuracy of the target inpainting model in predicting the visual details of the inpainted target object.

[0145] The foreground mask prediction loss characterizes the difference between the foreground mask features in the sample restored image and the foreground mask features obtained from real sample image features or simulated foreground mask images. The foreground mask prediction loss is used to evaluate the accuracy of the target restoration model in predicting the structure and spatial location of the restored target object.

[0146] S1020: The overall prediction loss, the foreground image prediction loss, and the foreground mask prediction loss are weighted and summed to obtain the loss value.

[0147] By assigning different weights to each component, the training process can be flexibly controlled. For example, the joint loss function L can be expressed as: in, Foreground mask prediction loss, To predict the overall loss, For the foreground image, predict the loss. β , λ , δ These are the weights corresponding to the foreground mask prediction loss, the overall prediction loss, and the foreground image prediction loss, respectively.

[0148] The embodiments of this disclosure calculate the independent losses of the three dimensions of overall prediction, foreground image prediction, and foreground mask prediction respectively. The parameters of the target restoration model are precisely measured and optimized from different dimensions, avoiding the problem of target ambiguity caused by mixed loss. This enhances the controllability and targeting of the training process, thereby effectively driving the target restoration model to converge to a better performance state quickly and stably.

[0149] As described in the above embodiments, the encoding process for image data such as simulated sample images, real sample images, and simulated images can be implemented using a variational autoencoder. Correspondingly, the encoding results of sample-restored image features, real sample image features, and image features are all feature representations of the latent feature space. Therefore, to further improve the training effect, this disclosure provides a more refined scheme for calculating the loss value.

[0150] Figure 11 The diagram shown is a flowchart illustrating the steps of weighted summation of the overall prediction loss, foreground image prediction loss, and foreground mask prediction loss to obtain the loss value, according to an embodiment of this disclosure. Figure 11 As shown in the embodiments of this disclosure, the step of weighted summation of the overall prediction loss, the foreground image prediction loss, and the foreground mask prediction loss to obtain the loss value includes the following steps.

[0151] S1110, determine the divergence loss.

[0152] Divergence loss, also known as Kullback-Leibler (KL) divergence loss, is an asymmetric measure used to quantify the difference between two probability distributions. Specifically, divergence loss measures the difference between the latent variable probability distributions output by the diffusion generation layer (i.e., the prediction module) and the latent variable probability distributions output by the target encoding module, and can be expressed in the following form: in, For divergence loss, The probability distribution of latent variables output by the diffusion generation layer. Z represents the probability distribution of latent variables output by the target encoding module, where z represents the latent variable.

[0153] S1120, the divergence loss, overall prediction loss, foreground image prediction loss, and foreground mask prediction loss are weighted and summed to obtain the loss value.

[0154] Specifically, the joint loss function L can be expressed in the following form: in, These are the weights corresponding to the divergence loss. Overall prediction loss. It can be expressed in the following forms: in, Let i be the i-th feature point in the real sample image features. The i-th feature point in the image features is repaired for the sample, where N is the total number of feature points. Understandably, the calculation methods for the foreground image prediction loss and foreground mask prediction loss are similar to those for the overall prediction loss, and will not be elaborated upon here.

[0155] Optionally, the weight of the overall prediction loss is greater than the weight of the foreground image prediction loss and the weight of the foreground mask prediction loss, and the weights of both the foreground image prediction loss and the foreground mask prediction loss are greater than the weight of the divergence loss. For example, the weights can be set as follows: =0.1, λ=1.0, β=δ =0.5; or, it can also be set to: α =0.2, λ =1.0, β =0.5, δ =0.4, and this disclosure does not impose specific restrictions on the specific numerical setting of the weight.

[0156] The above weighting strategy is based on the priority of each optimization objective during the repair process.

[0157] The overall prediction loss directly relates to the global consistency and realism of the image content and should be given the highest weight. The foreground image prediction loss and foreground mask prediction loss are applied to the prediction of local visual details and structures, respectively, and are given high weights for fine-tuning. The divergence loss is mainly used to stabilize training and improve distribution quality, and is given a lower weight to avoid over-constraining and affecting the reconstruction of image content.

[0158] In this embodiment, by introducing divergence loss, the target restoration model focuses not only on the accuracy of the image features to be restored but also on the overall distribution characteristics of the features, thereby improving the accuracy and reliability of the target restoration model. Secondly, by weighting the multi-dimensional losses and configuring differentiated weights, a balance between training stability and generation quality is achieved.

[0159] The above embodiments describe the training process of the target repair model. The following describes the method for generating training data pairs.

[0160] Figure 12 The diagram shown is a schematic flowchart illustrating the generation process of training data pairs according to an embodiment of this disclosure. Figure 12 As shown in this embodiment of the disclosure, the process of generating training data pairs includes: S1210, Obtain the real sample image and the semantic annotation information corresponding to the real sample image.

[0161] Realistic sample images refer to highly realistic images actually captured by vehicle-mounted or fixed cameras. Optionally, realistic sample images are images actually captured in the target scene.

[0162] Semantic annotation information consists of labels that identify the category (such as vehicle, pedestrian, road, crushing station, background, etc.) of each pixel in a real sample image, corresponding to the real sample image. Semantic annotation information can be obtained manually or automatically using a semantic segmentation model.

[0163] S1220 generates a real foreground mask image corresponding to a real sample image based on semantic annotation information.

[0164] Based on semantic annotation information, pixel regions belonging to the preset foreground category are identified as foreground regions, and other regions are identified as background regions, and corresponding real foreground mask images are generated.

[0165] The preset foreground categories predefine the categories of target objects. For example, the preset foreground categories include dynamic target categories such as vehicles and pedestrians, as well as static target categories such as crushing plants and tunnels.

[0166] S1230, based on the real foreground mask image, segments the real sample image into a real foreground image and a real background image.

[0167] The real foreground mask image can distinguish between the foreground region and the background region in a real sample image. Therefore, based on the real foreground mask image, a real sample image can be segmented into the real foreground image corresponding to the foreground region and the real background image corresponding to the background region.

[0168] S1240 performs a non-realistic transformation on the real foreground image and merges the transformed real foreground image with the real background image to generate a simulation sample image.

[0169] Non-realistic transformation is used to increase the diversity of the foreground region. By applying non-realistic transformation to the real foreground image, the problem of poor foreground realism in the simulation image is simulated. Then, the real foreground image after non-realistic transformation is fused with the real background image to obtain a simulation sample image with low realism.

[0170] Real sample images and simulated sample images together form a training data pair, and the real foreground mask image can be used as the simulated foreground mask image corresponding to the simulated sample image.

[0171] Optionally, the non-realistic transformation processing includes at least one of brightness adjustment, contrast adjustment, blurring, and geometric transformation.

[0172] Brightness adjustment involves enhancing or reducing the brightness of a realistic foreground image to simulate potential disharmony between the foreground and background under varying lighting conditions. The brightness adjustment process can be represented as: in, The transformed image, The image before transformation. This is the brightness scaling factor. Optionally, The value range is [0.6, 1.4].

[0173] Contrast adjustment involves adjusting the contrast of a real foreground image to simulate the visual differences between the foreground and background in different scenes, thereby improving the diversity and robustness of the training data. The contrast adjustment process can be represented as: in, This is the brightness scaling factor. Optionally, The value range is [0.7, 1.3].

[0174] Blur processing applies a blurring operation to a real foreground image to simulate the blurring phenomenon that may occur during the synthesis of new perspectives. The blur processing process can be represented as follows: Wherein, σ is the fuzzy intensity control parameter. Optionally, the value range of σ is [0.5, 2.0].

[0175] The process of fusing the real foreground image after non-realistic transformation with the real background image to generate a simulated sample image C can be represented as follows: Where M is the simulated foreground mask image, This represents a non-realistic transformation applied to the real foreground image, where F is the real foreground image and B is the real background image. This indicates pixel-by-pixel multiplication.

[0176] The above text combined Figures 1 to 12 The method embodiments of this disclosure have been described in detail below, in conjunction with... Figure 13 The apparatus embodiments of this disclosure are described in detail below. It should be understood that the descriptions of the method embodiments correspond to the descriptions of the apparatus embodiments; therefore, any parts not described in detail can be referred to the foregoing method embodiments.

[0177] Figure 13 The diagram shown is a structural schematic of a simulation-based image generation apparatus provided in an embodiment of this disclosure. Figure 13 As shown, the simulation-based image generation apparatus 1300 of this disclosure includes: an acquisition module 1310, an extraction module 1320, and a generation module 1330.

[0178] The acquisition module 1310 is configured to acquire simulation image data of the target scene and input the simulation image data into the target repair model. The simulation image data includes the simulation image, as well as the background image and foreground mask image corresponding to the simulation image.

[0179] The extraction module 1320 is configured to extract features from the simulation image and the foreground mask image through the coding layer of the target repair model to obtain image features, and to extract features from the background image to obtain scene condition features.

[0180] The generation module 1330 is configured to take image features as input and scene condition features as diffusion conditions through the diffusion generation layer of the target restoration model to generate a real restored image corresponding to the simulation image.

[0181] Below, for reference Figure 14 To describe an electronic device according to embodiments of the present disclosure.

[0182] Figure 14 The diagram shown is a structural schematic of an electronic device provided according to an embodiment of this disclosure. Figure 14 As shown, the electronic device 1400 includes one or more processors 1410 and memory 1420.

[0183] The processor 1410 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1400 to perform desired functions.

[0184] The memory 1420 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1410 may execute the program instructions to implement the simulation-based image generation methods of the various embodiments of this disclosure described above, and / or other desired functions.

[0185] In one example, the electronic device 1400 may also include an input device 1430 and an output device 1440, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0186] The input device 1430 may include, for example, a keyboard, a mouse, etc.

[0187] The output device 1440 can output various information to the outside, including realistically restored images. The output device 1440 may include, for example, a display, a printer, a communication network, and remote output devices connected to it.

[0188] Of course, for the sake of simplicity, Figure 14 Only some of the components of the electronic device 1400 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 1400 may include any other suitable components depending on the specific application.

[0189] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the simulation-based image generation methods according to various embodiments of this disclosure described above.

[0190] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0191] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the simulation-based image generation methods according to various embodiments of this disclosure described above.

[0192] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0193] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0194] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0195] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.

[0196] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0197] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A simulation-based image generation method, characterized in that, include: Acquire simulation image data of the target scene and input the simulation image data into the target repair model. The simulation image data includes the simulation image, as well as the background image and foreground mask image corresponding to the simulation image. Through the coding layer of the target repair model, feature extraction is performed on the simulation image and the foreground mask image to obtain image features, and feature extraction is performed on the background image to obtain scene condition features; The target restoration model uses the diffusion generation layer to generate a real restored image corresponding to the simulated image, taking the image features as input and the scene condition features as diffusion conditions.

2. The method according to claim 1, characterized in that, The encoding layer includes a first encoding module and a second encoding module; The step involves extracting features from the simulated image and the foreground mask image through the encoding layer of the target restoration model to obtain image features, including: The simulation image is encoded by the first encoding module to obtain the first image sub-feature corresponding to the simulation image; The foreground mask image is encoded by the second encoding module to obtain the second image sub-feature corresponding to the foreground mask image; The first image sub-feature and the second image sub-feature are concatenated to obtain the image feature.

3. The method according to claim 1, characterized in that, The diffusion generation layer includes a prediction module and a decoding module; The step of generating a realistic restored image corresponding to the simulated image through the diffusion generation layer of the target restoration model, using the image features as input and the scene condition features as diffusion conditions, includes: The prediction module predicts the overall features, foreground image features, and foreground mask features corresponding to the real restored image based on the image features and the scene condition features. The overall features, the foreground image features, and the foreground mask features are fused to obtain the fused features; The fused features are decoded by the decoding module to obtain the real restored image.

4. The method according to claim 1, characterized in that, The target repair model is trained using the following method: Acquire training data pairs, which include simulated sample images and real sample images, wherein the simulated sample images have non-realistic visual features; Determine the simulation background image and simulation foreground mask image corresponding to the simulation sample image; The target repair model's encoding layer extracts features from the simulated sample image and the simulated foreground mask image to obtain sample image features, and samples the simulated background image to obtain sample scene condition features. The sample image features are generated by using the diffusion generation layer of the target restoration model as input and the sample scene condition features as diffusion conditions. The real sample image is encoded by the target encoding module to obtain the features of the real sample image; Based on the difference between the features of the sample-restored image and the features of the real sample image, a loss value is calculated, and the encoding layer and the diffusion generation layer of the target restoration model are adjusted according to the loss value.

5. The method according to claim 4, characterized in that, The step of calculating the loss value based on the difference between the features of the repaired sample image and the features of the real sample image includes: Determine the overall prediction loss, foreground image prediction loss, and foreground mask prediction loss between the features of the restored sample image and the features of the real sample image, respectively; The loss value is obtained by weighted summing of the overall prediction loss, the foreground image prediction loss, and the foreground mask prediction loss.

6. The method according to claim 5, characterized in that, Both the sample restored image features and the real sample image features are feature representations in the latent feature space. The step of weighted summing of the overall prediction loss, the foreground image prediction loss, and the foreground mask prediction loss to obtain the loss value includes: Determine the divergence loss, which is used to measure the difference between the latent variable probability distribution output by the diffusion generation layer and the latent variable probability distribution output by the target encoding module; The loss value is obtained by weighted summing of the divergence loss, the overall prediction loss, the foreground image prediction loss, and the foreground mask prediction loss. Preferably, the weight of the overall prediction loss is greater than the weight of the foreground image prediction loss and the weight of the foreground mask prediction loss, and the weights of the foreground image prediction loss and the foreground mask prediction loss are both greater than the weight of the divergence loss.

7. The method according to any one of claims 2 to 6, characterized in that, Also includes: Obtain the real sample image and the semantic annotation information corresponding to the real sample image; Based on the semantic annotation information, a real foreground mask image corresponding to the real sample image is generated; Based on the real foreground mask image, the real sample image is segmented into a real foreground image and a real background image; The real foreground image is subjected to a non-realistic transformation, and the real foreground image after the non-realistic transformation is fused with the real background image to generate the simulation sample image; Preferably, the non-realistic transformation processing includes at least one of brightness adjustment, contrast adjustment, blurring, and geometric transformation.

8. The method according to claim 1, characterized in that, The simulated image data includes multiple simulated images with spatiotemporal consistency; The acquisition of simulated image data in the target scene includes: Acquire the background image of the target scene; Extract the 3D model of the target object from the 3D model library, and perform multi-view rendering on the 3D model of the target object to obtain multiple foreground target images from different viewpoints and the foreground mask image corresponding to each of the multiple foreground target images; Using the foreground mask image corresponding to each foreground target image, the foreground target image is fused with the background image to generate a simulated image from the perspective of the foreground target image; Preferably, the target object includes dynamic targets and / or static targets; Preferably, the target scene includes a mining scene; Preferably, the target object includes at least one of a vehicle, a pedestrian, a crushing station, and a tunnel.

9. The method according to claim 1, characterized in that, The simulated image data includes multiple simulated images with spatiotemporal consistency; The acquisition of simulated image data in the target scene includes: Acquire multiple initial images from different perspectives in the target scene; The multiple initial images are reconstructed in three dimensions to obtain a three-dimensional model of the target scene; Extract at least one target object from the three-dimensional model; The target scene is edited by changing the pose of the vehicle and / or the pose of the target object, and the edited target scene is rendered from multiple perspectives to generate the multiple simulation images.

10. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1 to 9 by executing the executable instructions.