Model training method, synthetic data generation method and device

Through a new model training method, the parameters of the image generation model are updated using the loss function, which solves the problem of insufficient realism of synthetic images in the prior art, and achieves a higher realism and universal image generation effect.

CN119360159BActive Publication Date: 2025-05-16QINGKE LINGJING (ANHUI) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411918523.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-16
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The synthetic images generated by the existing style transfer image generation model are insufficiently realistic, which is difficult to meet the demand for high-reality images of autonomous driving simulation testing systems.

Method used

A model training method is proposed, by obtaining simulated images, image data and real images, inputting the initial image generation model, generating a synthetic image, and updating the model parameters based on the loss function to obtain the target image generation model to improve the reality of the synthetic image.

Benefits of technology

Through this method, the authenticity of the synthetic images generated by the image generation model and the universality of the model are significantly improved, and synthetic data closer to the real image can be generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360159B_ABST
    Figure CN119360159B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model training method, a synthetic data generation method and a device, which relate to the field of artificial intelligence technology, specifically to the technical fields of model training, image processing, etc., and can be applied to scenes such as image generation and simulation testing. The specific implementation scheme of the model training method includes: obtaining a first simulated image, image data corresponding to the first simulated image, and a first real image; inputting the first simulated image, the image data corresponding to the first simulated image, and the first real image into an initial image generation model, and generating a first synthetic image through the initial image generation model; determining a first loss function value based on the first synthetic image and the first real image, and a preset loss function; updating the model parameters of the initial image generation model based on the loss function value, and obtaining a target image generation model, and the loss function value includes the first loss function value. The present disclosure can improve the authenticity of the synthetic image generated by the image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as model training and image processing, and especially to model training methods, synthetic data generation methods and devices, which can be applied to scenarios such as image generation and simulation testing. Background Art

[0002] The autonomous driving simulation test system is an important tool in the development and testing of real autonomous driving systems, and can generate highly diverse simulation images. However, the simulation images generated by the autonomous driving simulation test system are quite different from real images in terms of realism.

[0003] Currently, the image generation model of style transfer can be used to transfer the features of real images to simulated images to obtain synthetic images with better realism.

[0004] However, the realism of synthetic images generated by current style transfer image generation models is still insufficient. Summary of the invention

[0005] The present disclosure provides a model training method, which can improve the authenticity of synthetic images obtained by an image generation model.

[0006] According to a first aspect of the present disclosure, a model training method is provided, which is applied to rendering high-fidelity synthetic data in a simulation environment based on a diffusion model, and the method comprises:

[0007] Acquire a first simulated image, image data corresponding to the first simulated image, and a first real image; input the first simulated image, image data corresponding to the first simulated image, and the first real image into an initial image generation model, and generate a first synthetic image through the initial image generation model; determine a first loss function value based on the first synthetic image and the first real image, and a preset loss function; update model parameters of the initial image generation model based on the loss function value to obtain a target image generation model, and the loss function value includes the first loss function value.

[0008] In some possible implementations, before inputting the first simulated image, the image data corresponding to the first simulated image, and the first real image into an initial image generation model and generating the first synthetic image through the initial image generation model, the method further includes:

[0009] From the first real image, determining a second real image that matches the first simulated image;

[0010] Inputting the first simulated image, image data corresponding to the first simulated image, and the first real image into an initial image generation model, and generating a first synthetic image through the initial image generation model, including:

[0011] The first simulated image, image data corresponding to the first simulated image, and the second real image are input into the initial image generation model, and the first synthetic image is generated by the initial image generation model.

[0012] In some possible implementations, determining, from the first real image, a second real image that matches the first simulated image includes:

[0013] Feature extraction is performed on the first simulated image and the first real image respectively to obtain a first feature corresponding to the first simulated image and a second feature corresponding to the first real image; a third feature is determined from the second feature according to the first feature and the second feature; and the first real image corresponding to the third feature is determined to be the second real image.

[0014] In some possible implementations, determining the third feature from the second feature according to the first feature and the second feature includes:

[0015] Determine the feature distance between the first feature and the second feature; and determine the third feature from the second feature through a clustering algorithm.

[0016] In some possible implementations, determining the third feature from the second feature according to the first feature and the second feature includes:

[0017] Determine a feature distance between the first feature and the second feature; and determine a third feature from the second feature according to the feature distance between the first feature and the second feature and a preset distance threshold.

[0018] In some possible implementations, the loss function value further includes a second loss function value, and the second loss function value is determined based on the first synthetic image and the first simulated image, and a preset loss function; updating the model parameters of the initial image generation model based on the loss function value to obtain the target semantic segmentation model includes:

[0019] Based on the first loss function value, the model parameters of the initial image generation model are updated to obtain a transition image generation model; based on the second loss function value, the model parameters of the transition image generation model are updated to obtain a target image generation model.

[0020] In some possible implementations, the image data corresponding to the first simulated image includes at least one of rendering pipeline intermediate data, depth information, semantic segmentation information, and normal information.

[0021] In some possible implementations, the initial image generation model includes a stable diffusion model.

[0022] The first aspect of the present disclosure has at least the following beneficial effects: it can improve the realism of the synthetic images generated by the image generation model and the versatility of the model.

[0023] According to a second aspect of the present disclosure, a method for generating synthetic data is provided, the method comprising:

[0024] Input the simulated image to be processed and the image data target image generation model, where the target image generation model is obtained based on the model training method of the first aspect; output the synthetic data through the target image generation model.

[0025] According to a third aspect of the present disclosure, a model training device is provided, which is used for rendering high-fidelity synthetic data in a simulation environment based on a diffusion model, and the device includes:

[0026] The acquisition module is used to acquire the first simulated image, the image data corresponding to the first simulated image and the first real image.

[0027] The generation module is used to input the first simulated image, the image data corresponding to the first simulated image and the first real image into the initial image generation model, and generate the first synthetic image through the initial image generation model.

[0028] A determination module is used to determine a first loss function value based on the first synthetic image, the first real image, and a preset loss function.

[0029] The training module updates the model parameters of the initial image generation model based on the loss function value to obtain the target image generation model, and the loss function value includes the first loss function value.

[0030] Optionally, the device further comprises:

[0031] The matching module is used to determine a second real image matching the first simulated image from the first real image before inputting the first simulated image, the image data corresponding to the first simulated image and the first real image into the initial image generation model to generate the first synthetic image through the initial image generation model. The generation module is specifically used to input the first simulated image, the image data corresponding to the first simulated image and the second real image into the initial image generation model to generate the first synthetic image through the initial image generation model.

[0032] Optionally, the matching module is specifically used to:

[0033] Feature extraction is performed on the first simulated image and the first real image respectively to obtain a first feature corresponding to the first simulated image and a second feature corresponding to the first real image; a third feature is determined from the second feature according to the first feature and the second feature; and the first real image corresponding to the third feature is determined to be the second real image.

[0034] Optionally, the matching module is specifically used to:

[0035] Determine the feature distance between the first feature and the second feature; and determine the third feature from the second feature through a clustering algorithm.

[0036] Optionally, the matching module is specifically used to:

[0037] Determine a feature distance between the first feature and the second feature; and determine a third feature from the second feature according to the feature distance between the first feature and the second feature and a preset distance threshold.

[0038] Optionally, the loss function value also includes a second loss function value, which is determined based on the first synthetic image and the first simulated image, and a preset loss function; the training module is specifically used to: update the model parameters of the initial image generation model based on the first loss function value to obtain a transition image generation model; update the model parameters of the transition image generation model based on the second loss function value to obtain a target image generation model.

[0039] Optionally, the image data corresponding to the first simulated image includes at least one of rendering pipeline intermediate data, depth information, semantic segmentation information, and normal information.

[0040] Optionally, the initial image generation model comprises a stable diffusion model.

[0041] According to a fourth aspect of the present disclosure, there is provided an image generating device, the device comprising:

[0042] The input module is used to input the simulated image and image data to be processed into the target image generation model, and the target image generation model is obtained based on the model training method of the first aspect. The output module is used to output the synthetic image through the target image generation model.

[0043] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method as described in the first aspect or the second aspect.

[0044] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method according to the first aspect or the second aspect.

[0045] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method according to the first aspect or the second aspect when executed by a processor.

[0046] The beneficial effects of the second to seventh aspects of the present disclosure can refer to the beneficial effects of the first aspect and will not be repeated here.

[0047] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0049] Figure 1 A flowchart of a model training method provided in an embodiment of the present disclosure;

[0050] Figure 2 Another flowchart of the model training method provided in the embodiment of the present disclosure;

[0051] Figure 3 A schematic diagram of another flow chart of the model training method provided in the embodiment of the present disclosure;

[0052] Figure 4 A schematic diagram of another flow chart of the model training method provided in the embodiment of the present disclosure;

[0053] Figure 5 A schematic diagram of another flow chart of the model training method provided in the embodiment of the present disclosure;

[0054] Figure 6 A schematic diagram of another flow chart of the model training method provided in the embodiment of the present disclosure;

[0055] Figure 7 A schematic diagram of the principle of the model training method provided in the embodiment of the present disclosure;

[0056] Figure 8 A schematic diagram of a flow chart of a synthetic data generation method provided in an embodiment of the present disclosure;

[0057] Fig. 9 A schematic diagram of a principle of a synthetic data generation method provided in an embodiment of the present disclosure;

[0058] Fig.10 A schematic diagram of the composition of a model training device provided in an embodiment of the present disclosure;

[0059] Fig.11 Another schematic diagram of the composition of the model training device provided in the embodiment of the present disclosure;

[0060] Fig.12 A schematic diagram of the composition of a synthetic data generating device provided in an embodiment of the present disclosure;

[0061] Fig.13 FIG. 1 is a schematic block diagram of an example electronic device 1300 that may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0062] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0063] It should be understood that in the embodiments of the present disclosure, the character " / " generally indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", etc. are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features.

[0064] The autonomous driving simulation test system is an important tool in the development and testing of real autonomous driving systems. With the help of the sensor model of the autonomous driving simulation test system, data with accurate annotation information can be generated. Benefiting from the configurability and programmability of the autonomous driving simulation test system, the autonomous driving simulation test system can generate highly diverse synthetic images, and the acquisition cost of synthetic images is much lower than that of real images. However, due to the limitations of current computational graphics rendering technology, the images generated by the autonomous driving simulation test system are still quite different from real images in terms of realism.

[0065] Based on this, the current application of synthetic images in the field of autonomous driving, in addition to some special and limited scene tasks, a considerable amount of research focuses on how to improve the realism of synthetic images and the performance of synthetic images in autonomous driving perception tasks. From a mathematical perspective, the realism of data can be divided into distribution-based realism and preference-based realism. One angle to evaluate the preference-based realism is visual realism, and the evaluation of human vision is an important reference for evaluating data authenticity.

[0066] Currently, the image generation model of style transfer can be used to transfer the features of real images to simulated images to obtain synthetic images with better realism.

[0067] However, the realism of synthetic images generated by current style transfer image generation models is still insufficient.

[0068] Against this background, the present disclosure provides a model training method capable of improving the realism of synthetic images generated by an image generation model.

[0069] Exemplarily, the model training method provided by the present disclosure can be applied to image generation and simulation test scenarios.

[0070] Exemplarily, the execution subject of the model training method provided in the embodiment of the present disclosure may be a terminal device, a computer, or a server, or other devices with data processing capabilities. The execution subject of the method is not limited here. In some embodiments, the execution subject of the model training method provided in the embodiment of the present disclosure may be a terminal device (such as a vehicle-mounted computer) on the main vehicle.

[0071] Optionally, the terminal device may be a mobile phone, or a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of the present disclosure do not limit the specific type of the terminal device.

[0072] In some embodiments, the server may be a single server, or a server cluster composed of multiple servers. In some implementations, the server cluster may also be a distributed cluster. The present disclosure does not limit the specific implementation of the server.

[0073] Figure 1 The following is a flow chart of the model training method provided in the embodiment of the present disclosure. Figure 1 As shown, the method may include:

[0074] S101. Acquire a first simulation image, image data corresponding to the first simulation image, and a first real image.

[0075] For example, the first simulation image may be acquired through an autonomous driving simulation test system, or may be acquired through other methods (such as an image generation engine), or may be acquired directly from a simulation image in an existing simulation image set, without limitation.

[0076] Exemplarily, the image data corresponding to the first simulation image may be intermediate state data of the rendering pipeline during the generation process of the first simulation image in the autonomous driving simulation test system, as well as any one or more of the depth information, semantic segmentation information, and normal information corresponding to the first simulation image, without limitation.

[0077] For example, the first real image can be obtained by using a data collection vehicle to collect a real environment using a collection device (such as a camera or a video camera), or it can be obtained by directly obtaining a real image from an existing open source real image set, and there is no limitation to this.

[0078] For example, the first real image may be a real image including different styles, such as sunny, cloudy, overcast, rainy and snowy weather, etc., and there is no limitation on the style of the first real image.

[0079] In some possible embodiments, the execution entity may obtain the first simulated image, image data corresponding to the first simulated image, and the first real image from other electronic devices or storage media through communication methods such as wireless or wired networks and wired connections.

[0080] For example, after the first simulation image, the image data corresponding to the first simulation image, and the first real image are acquired, S103 may be executed.

[0081] S103: Input the first simulated image, image data corresponding to the first simulated image, and the first real image into an initial image generation model, and generate a first synthetic image through the initial image generation model.

[0082] Exemplarily, the initial image generation model may be a generative adversarial network (GANs), a variational autoencoder (VAEs), a diffusion model (Diffusion Models), or other image generation models, without limitation.

[0083] Exemplarily, taking the stable diffusion model as the initial image generation model, the stable diffusion model can be adjusted by adding a multilayer perceptron (MLP) to the stable diffusion model, and the image data corresponding to the first simulated image of the input model is compressed using the multilayer perceptron, so that the image data corresponding to the first simulated image of the input is in the same dimension as the input image, thereby performing model training.

[0084] S104: Determine a first loss function value based on the first synthetic image, the first real image, and a preset loss function.

[0085] Exemplarily, the first loss function value can be determined by using the feature values ​​corresponding to the first synthetic image and the first real image and a preset loss function.

[0086] It should be noted that the preset loss function can use the loss function of the existing image generation model, which will not be described here; the loss function can also be redesigned without limitation.

[0087] S105. Update the model parameters of the initial image generation model based on the loss function value to obtain the target image generation model.

[0088] The loss function value includes a first loss function value.

[0089] Exemplarily, the model parameters of the initial image generation model can be updated by back propagation according to the first loss function value to obtain the target image generation model.

[0090] The disclosed embodiment acquires a first simulated image, image data corresponding to the first simulated image, and a first real image, inputs the first simulated image, the image data corresponding to the first simulated image, and the first real image into an initial image generation model, generates a first synthetic image through the initial image generation model, determines a first loss function value based on the first synthetic image and the first real image, and a preset loss function, updates the model parameters of the initial image generation model based on the loss function value, and obtains a target image generation model. The image data corresponding to the first simulated image can be used as one of the conditions for model training, so that the training of the initial image generation model can better link the intrinsic features of the simulated image with the real image, thereby making the synthetic image generated by the trained target image generation model more realistic, thereby improving the realistic degree of the synthetic image obtained by the image generation model.

[0091] At the same time, since the initial image generation model can be trained using first real images of different styles, the trained target image generation model can generate synthetic images of different styles, thereby improving the versatility of the image generation model.

[0092] Figure 2 Another flow chart of the model training method provided in the embodiment of the present disclosure. Figure 2 As shown, before inputting the first simulated image, the image data corresponding to the first simulated image, and the first real image into the initial image generation model and generating the first synthetic image through the initial image generation model, the method may further include:

[0093] S102: Determine, from the first real image, a second real image that matches the first simulated image.

[0094] Inputting the first simulated image, image data corresponding to the first simulated image, and the first real image into an initial image generation model, and generating a first synthetic image through the initial image generation model may include:

[0095] The first simulated image, image data corresponding to the first simulated image, and the second real image are input into the initial image generation model, and the first synthetic image is generated by the initial image generation model.

[0096] For example, after the first simulation image, the image data corresponding to the first simulation image, and the first real image are acquired, S102 may be executed.

[0097] Exemplarily, the similarity between the first simulated image and the first real image may be determined by a similarity calculation formula or a similarity calculation model, and then the second real image may be determined from the first real image according to the similarity and a preset similarity threshold.

[0098] For example, taking the first real image including image A, image B, and image C, and the first simulated image including image D as an example, the similarity calculation formula can be used to calculate that the similarity between image A and image D is 0.75, the similarity between image B and image D is 0.85, and the similarity between image C and image D is 0.9. The preset similarity threshold is 0.8, and image B and image C can be determined as the second real image.

[0099] It should be noted that the above example of similarity determination is only an exemplary description. In practical applications, other related technologies can also be used to implement similarity determination, which is not limited here.

[0100] Exemplarily, after determining a second real image matching the first simulation image from the first real image, S103 may be performed.

[0101] This embodiment determines a second real image that matches the first simulated image from the first real image, inputs the first simulated image, the image data corresponding to the first simulated image, and the second real image into an initial image generation model, and generates a first synthetic image through the initial image generation model. This can make the intrinsic features between the simulated image used for training and the real image more closely connected, so that the training of the initial image generation model can further connect the simulated image with the intrinsic features of the real image, thereby making the synthetic image generated by the trained target image generation model more realistic, further improving the authenticity of the synthetic image obtained by the image generation model.

[0102] At the same time, this embodiment does not need to build a simulation scene that is twin to the real image in the autonomous driving simulation test system, which reduces the difficulty of implementing style transfer.

[0103] Figure 3 Another flow chart of the model training method provided in the embodiment of the present disclosure. Figure 3As shown, determining a second real image matching the first simulated image from the first real image may include:

[0104] S301 , performing feature extraction on a first simulated image and a first real image respectively to obtain a first feature corresponding to the first simulated image and a second feature corresponding to the first real image.

[0105] Exemplarily, the first real image and the first simulated image can be input into the same convolutional neural network, and a layer of the network is selected as a feature extraction layer, so as to extract the corresponding first feature of the first simulated image and the corresponding second feature of the first real image.

[0106] It should be noted that the above example of feature extraction is only an exemplary description. In practical applications, feature extraction can also be achieved by other related technologies, which is not limited here.

[0107] S302. Determine a third feature from the second feature based on the first feature and the second feature.

[0108] Exemplarily, the third feature may be determined from the second feature by a feature distance between the first feature and the second feature.

[0109] For example, the second real image may be determined by a preset distance threshold or a clustering algorithm.

[0110] For example, taking the feature distance between the first feature A and the second feature B as 0.5, the feature distance between the first feature A and the second feature B as 0.8, the feature distance between the first feature A and the second feature C as 0.9, and the preset distance threshold as 0.7, the second feature B and the second feature C can be determined as the third feature.

[0111] S303: Determine that the first real image corresponding to the third feature is the second real image.

[0112] Exemplarily, the first real image corresponding to the third feature may be determined as the second real image according to the correspondence between the feature and the first real image.

[0113] For example, the correspondence between the feature and the first real image may be obtained by the execution subject establishing a mapping relationship table according to the correspondence between the feature and the first real image after executing S301.

[0114] Exemplarily, after determining that the first real image corresponding to the third feature is the second real image, S103 may be executed.

[0115] This embodiment extracts features from the first simulated image and the first real image respectively to obtain a first feature corresponding to the first simulated image and a second feature corresponding to the first real image. Based on the first feature and the second feature, a third feature is determined from the second feature, and the first real image corresponding to the third feature is determined to be the second real image. Based on the relationship between the first feature and the second feature, the second real image can be accurately determined.

[0116] Figure 4 Another flow chart of the model training method provided in the embodiment of the present disclosure. Figure 4 As shown, according to the first feature and the second feature, determining the third feature from the second feature may include:

[0117] S401: Determine a feature distance between a first feature and a second feature.

[0118] Exemplarily, the feature distance between the first feature and the second feature may be determined by a Euclidean distance calculation formula.

[0119] It should be noted that the above example of determining the feature distance is only an exemplary description. In practical applications, the feature distance can also be determined by using related technologies such as the Manhattan distance calculation formula and the Chebyshev calculation formula, which is not limited here.

[0120] S402: Determine a third feature from the second feature by using a clustering algorithm.

[0121] For example, clustering algorithm is a statistical analysis method for studying (sample or indicator) classification problems, and is also an important algorithm for data mining. Cluster analysis is composed of several patterns. Usually, a pattern is a vector of measurement or a point in multidimensional space. Clustering algorithm is based on similarity. Patterns in a cluster have more similarities than patterns in different clusters.

[0122] Exemplarily, a K-means clustering algorithm may be used to find, from the second features, a plurality of third features that are closest to the first feature in a feature distance dimension.

[0123] It should be noted that the above example of clustering algorithm is only an exemplary description. In practical applications, clustering calculation can also be achieved through related technologies such as Mean-Shift clustering algorithm, density-based noise application space clustering algorithm, etc., which is not limited here.

[0124] Exemplarily, after the third feature is determined from the second feature by using a clustering algorithm, S303 may be performed.

[0125] This embodiment determines the feature distance between the first feature and the second feature, and determines the third feature from the second feature through a clustering algorithm, so that the third feature can be quickly determined.

[0126] Figure 5 Another flow chart of the model training method provided in the embodiment of the present disclosure. Figure 5 As shown, according to the first feature and the second feature, determining the third feature from the second feature may include:

[0127] S501: Determine a feature distance between a first feature and a second feature.

[0128] Exemplarily, the feature distance between the first feature and the second feature may be determined by a Euclidean distance calculation formula.

[0129] It should be noted that the above example of determining the feature distance is only an exemplary description. In practical applications, the feature distance can also be determined by using related technologies such as the Manhattan distance calculation formula and the Chebyshev calculation formula, which is not limited here.

[0130] S502: Determine a third feature from the second feature according to a feature distance between the first feature and the second feature and a preset distance threshold.

[0131] For example, taking the feature distance between the first feature A and the second feature B as 0.6, the feature distance between the first feature A and the second feature B as 0.7, the feature distance between the first feature A and the second feature C as 0.8, and the preset distance threshold as 0.7, the second feature C can be determined as the third feature.

[0132] Exemplarily, after determining the third feature from the second feature based on the feature distance between the first feature and the second feature and a preset distance threshold, S303 may be executed.

[0133] This embodiment determines the feature distance between the first feature and the second feature, and determines the third feature from the second feature according to the feature distance between the first feature and the second feature and a preset distance threshold, so that the third feature can be accurately determined.

[0134] Figure 6 Another flow chart of the model training method provided by the embodiment of the present disclosure. The loss function value also includes a second loss function value, and the second loss function value is determined based on the first synthetic image and the first simulation image, and a preset loss function. Figure 6 As shown, updating the model parameters of the initial image generation model based on the loss function value to obtain the target semantic segmentation model may include:

[0135] S601. Update the model parameters of the initial image generation model based on the first loss function value to obtain a transition image generation model.

[0136] Exemplarily, the model parameters of the initial image generation model can be updated by back propagation according to the first loss function value to obtain the transition image generation model.

[0137] S602. Update the model parameters of the transition image generation model based on the second loss function value to obtain the target image generation model.

[0138] Exemplarily, the second loss function value may include any one or more of structural similarity loss, content loss, and semantic information.

[0139] Among them, the structural similarity loss is to maximize the structural similarity between the first simulated image and the first synthesized image during the training process, and the structural similarity is used to characterize the layout, content arrangement and other features of the two images; the content loss is to minimize the difference in image content between the first simulated image and the first synthesized image during the training process; the semantic loss is to minimize the loss of semantic information (such as categories and other dimensions) between the first simulated image and the first synthesized image during the training process. The role of these loss function values ​​is to improve the consistency of the content, layout and category between the synthesized image and the simulated image of the input model, under the premise that the synthesized image generated by the trained target image generation network is close to the real image.

[0140] Exemplarily, the model parameters of the transition image generation model can be updated by back propagation according to the second loss function value to obtain the target image generation model.

[0141] In some possible embodiments, the execution subject may execute S602 and then execute S601, and there is no restriction on the execution order of S601 and S602.

[0142] This embodiment obtains a transition image generation model by updating the model parameters of the initial image generation model based on the first loss function value, and obtains a target image generation model by updating the model parameters of the transition image generation model based on the second loss function value determined by the first synthetic image and the first simulated image and a preset loss function. This can improve the consistency between the synthetic image and the simulated image on the premise that the style of the synthetic image generated by the trained target image generation model is close to that of the real image.

[0143] The above embodiments introduce the model training method provided by the embodiments of the present disclosure. Figure 7 , through a specific example, the model training method is explained in more detail. Figure 7The following is a schematic diagram of the principle of the model training method provided in the embodiment of the present disclosure. The model training method may include the following steps 1 to 7:

[0144] Step 1: Acquire a first simulated image, image data corresponding to the first simulated image, and a first real image.

[0145] Step 2: Determine, from the first real image, a second real image that matches the first simulated image.

[0146] Step 3: input the first simulated image, the image data corresponding to the first simulated image, and the second real image into an initial image generation model, and generate a first synthetic image through the initial image generation model.

[0147] Step 4: Determine a first loss function value based on the first synthetic image, the second real image, and a preset loss function.

[0148] Step 5: Determine a second loss function value based on the first synthetic image, the first simulation image, and a preset loss function.

[0149] Step 6: Update the model parameters of the initial image generation model based on the first loss function value to obtain a transition image generation model.

[0150] Step 7: Update the model parameters of the transition image generation model based on the second loss function value to obtain the target image generation model.

[0151] The initial image generation model may include a stable diffusion model.

[0152] The technical effects of this embodiment can refer to the technical effects of the aforementioned embodiments, which will not be described in detail here.

[0153] The present disclosure also provides a method for generating synthetic data. Figure 8 Schematic diagram of the process of generating synthetic data provided by the present disclosure. Figure 8 As shown, the method may include:

[0154] S801, inputting the simulation image and image data to be processed into the target image generation model.

[0155] Among them, the target image generation model is obtained based on the model training method in the aforementioned embodiment.

[0156] Exemplarily, the image data may be any one or more of rendering pipeline intermediate data, depth information, semantic segmentation information, and normal information.

[0157] S802: Output synthetic data through the target image generation model.

[0158] For example, a simulated image to be processed and a set of image data can be directly input into the target image generation model, or multiple simulated images to be processed and multiple sets of image data can be first input into the target image generation model. Fig. 9 In the control network shown, Fig. 9 A schematic diagram of the principle of a synthetic data generation method provided in an embodiment of the present disclosure is provided, in which a control network is used to control the simulated image and image data input into a target image generation model to obtain a synthetic image. The control network can be used to improve the generation efficiency of the target image generation model in generating a synthetic image expected by the user.

[0159] The disclosed embodiment can obtain a more realistic synthetic image by inputting the simulated image and image data to be processed into a target image generation model and outputting a synthetic image through the target image generation model.

[0160] In an exemplary embodiment, the present disclosure also provides a model training device, which can be used to implement the model training method as described in the aforementioned embodiment. Fig.10 Schematic diagram of the composition of the model training device provided in the embodiment of the present disclosure. Fig.10 As shown, the device may include:

[0161] The acquisition module 1001 is used to acquire a first simulation image, image data corresponding to the first simulation image, and a first real image.

[0162] The generation module 1003 is used to input the first simulated image, the image data corresponding to the first simulated image and the first real image into the initial image generation model, and generate a first synthetic image through the initial image generation model.

[0163] The determination module 1004 is used to determine a first loss function value based on the first synthetic image, the first real image, and a preset loss function.

[0164] The training module 1005 updates the model parameters of the initial image generation model based on the loss function value to obtain the target image generation model, and the loss function value includes the first loss function value.

[0165] Fig.11 Another schematic diagram of the composition of the model training device provided in the embodiment of the present disclosure. Fig.11 As shown, the device also includes:

[0166] The matching module 1002 is used to determine a second real image that matches the first simulated image from the first real image before inputting the first simulated image, the image data corresponding to the first simulated image and the first real image into the initial image generation model to generate the first synthetic image through the initial image generation model.

[0167] The generation module 1003 is specifically used to: input the first simulation image, the image data corresponding to the first simulation image and the second real image into the initial image generation model, and generate the first synthetic image through the initial image generation model.

[0168] Optionally, the matching module 1002 is specifically configured to:

[0169] Feature extraction is performed on the first simulated image and the first real image respectively to obtain a first feature corresponding to the first simulated image and a second feature corresponding to the first real image; a third feature is determined from the second feature according to the first feature and the second feature; and the first real image corresponding to the third feature is determined to be the second real image.

[0170] Optionally, the matching module 1002 is specifically configured to:

[0171] Determine the feature distance between the first feature and the second feature; and determine the third feature from the second feature through a clustering algorithm.

[0172] Optionally, the matching module 1002 is specifically configured to:

[0173] Determine a feature distance between the first feature and the second feature; and determine a third feature from the second feature according to the feature distance between the first feature and the second feature and a preset distance threshold.

[0174] Optionally, the loss function value also includes a second loss function value, which is determined based on the first synthetic image and the first simulated image, and a preset loss function; the training module 1005 is specifically used to: update the model parameters of the initial image generation model based on the first loss function value to obtain a transition image generation model; update the model parameters of the transition image generation model based on the second loss function value to obtain a target image generation model.

[0175] Optionally, the image data corresponding to the first simulated image includes at least one of rendering pipeline intermediate data, depth information, semantic segmentation information, and normal information.

[0176] Optionally, the initial image generation model comprises a stable diffusion model.

[0177] The beneficial effects of the above-mentioned model training device can refer to the beneficial effects of the model training method in the aforementioned embodiment, which will not be repeated here.

[0178] In an exemplary embodiment, the present disclosure also provides an image generating device, which can be used to implement the synthetic data generating method as described in the above embodiment. Fig.12 Schematic diagram of the composition of the image generating device provided by the embodiment of the present disclosure. Fig.12 As shown, the device may include:

[0179] The input module 1201 is used to input the simulation image and image data to be processed into the target image generation model, and the target image generation model is obtained based on the model training method in the aforementioned embodiment.

[0180] The output module 1202 is used to output the synthesized data through the target image generation model.

[0181] The beneficial effects of the above-mentioned image generating device can refer to the beneficial effects of the synthetic data generating method described in the above-mentioned embodiment, which will not be repeated here.

[0182] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device. The electronic device may be a server, a computer or other device as described in the aforementioned embodiment, and can be used to implement the model training method or synthetic data generation method provided in the embodiment of the present disclosure.

[0183] In an exemplary embodiment, the electronic device may include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model training method or synthetic data generation method described in the above embodiments.

[0184] For example, Fig.13 1 is a schematic block diagram of an example electronic device 1300 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0185] like Fig.13 As shown, the electronic device 1300 may include a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit 1308 into a random access memory (RAM). In the RAM 1303, various programs and data required for the operation of the electronic device 1300 can also be stored. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface is also connected to the bus 1304.

[0186] Multiple components in the electronic device 1300 are connected to the I / O interface 1305, including: an input unit 1306, such as a keyboard, a mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a disk, an optical disk, etc.; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1309 allows the electronic device 1300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0187] The computing unit 1301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), any appropriate processors, controllers, microcontrollers, etc. The computing unit 1301 performs the various methods and processes described above, such as a model training method or a synthetic data generation method. For example, in some embodiments, the model training method or the synthetic data generation method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1308.

[0188] In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the model training method or synthetic data generation method described above may be performed.

[0189] Alternatively, in other embodiments, the computing unit 1301 may be configured to execute the model training method or the synthetic data generation method in any other suitable manner (eg, by means of firmware).

[0190] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium and a computer program product.

[0191] In an exemplary embodiment, the readable storage medium may be a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to the above embodiments.

[0192] In an exemplary embodiment, a computer program product includes a computer program, and when the computer program is executed by a processor, the method according to the above embodiments is implemented.

[0193] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor, which can receive data and instructions from a storage system, at least one input device, at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0194] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0195] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0196] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0197] The systems and techniques described herein may be implemented in computing systems that include back-end components (e.g., as a data server), computing systems that include middleware components (e.g., an application server), computing systems that include front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or computing systems that include any combination of such back-end components, middleware components, and front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0198] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0199] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0200] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A model training method for rendering high-fidelity synthetic data in a simulation environment based on a diffusion model, characterized in that: The method comprises: Acquire a first simulated image, image data corresponding to the first simulated image, and a first real image, and determine, from the first real image, a second real image that matches the first simulated image; Inputting the first simulated image, image data corresponding to the first simulated image, and a second real image into an initial image generation model, and generating a first synthetic image through the initial image generation model; Determine a first loss function value based on the first synthetic image, the second real image, and a preset loss function; Determining a second loss function value based on the first synthetic image, the first simulation image, and a preset loss function; Updating the model parameters of the initial image generation model based on the first loss function value to obtain a transition image generation model; The model parameters of the transition image generation model are updated based on the second loss function value to obtain a target image generation model.

2. The method according to claim 1, characterized in that The step of determining, from the first real image, a second real image matching the first simulated image comprises: Performing feature extraction on the first simulated image and the first real image respectively to obtain a first feature corresponding to the first simulated image and a second feature corresponding to the first real image; determining a third feature from the second feature according to a feature distance between the first feature and the second feature; Determine that the first real image corresponding to the third feature is the second real image.

3. The method according to claim 2, characterized in that The determining, based on the first feature and the second feature, a third feature from the second feature includes: determining a feature distance between the first feature and the second feature; The third feature is determined from the second feature by using a clustering algorithm.

4. The method according to claim 2, characterized in that: The determining, based on the first feature and the second feature, a third feature from the second feature includes: determining a feature distance between the first feature and the second feature; The third feature is determined from the second feature according to a feature distance between the first feature and the second feature and a preset distance threshold.

5. The method according to any one of claims 1 to 4, characterized in that: The image data corresponding to the first simulation image includes at least one of rendering pipeline intermediate state data, depth information, semantic segmentation information, and normal information.

6. The method according to claim 5, characterized in that The initial image generation model includes a stable diffusion model.

7. A method for generating synthetic data, characterized in that: The method comprises: Inputting the simulated image and image data to be processed into a target image generation model, wherein the target image generation model is obtained based on the model training method according to any one of claims 1 to 6; The synthetic data is output by the target image generation model.

8. A model training device for rendering high-fidelity synthetic data in a simulation environment based on a diffusion model, characterized in that: The device comprises: an acquisition module, configured to acquire a first simulated image, image data corresponding to the first simulated image, and a first real image, and determine, from the first real image, a second real image matching the first simulated image; A generating module, configured to input the first simulated image, image data corresponding to the first simulated image, and a second real image into an initial image generating model, and generate a first synthetic image through the initial image generating model; A determination module, configured to determine a first loss function value based on the first synthetic image, the second real image, and a preset loss function; Determine a second loss function value based on the first synthetic image, the first simulation image, and a preset loss function; A training module, which updates the model parameters of the initial image generation model based on the first loss function value to obtain a transition image generation model; The model parameters of the transition image generation model are updated based on the second loss function value to obtain a target image generation model.

Citation Information

Patent Citations

  • Image rendering model training method and device, image rendering method and related equipment

    CN116863056A

  • Image generation model training method and system, image generation method and system and electronic equipment

    CN117541883A