Model training method and device, image generation method and device, electronic equipment, computer readable storage medium and computer program product
By acquiring and predicting noise under multiple light sources and training the image generation model to follow the physical laws of illumination changes, the problem in existing technologies where light enhancement is difficult to maintain the inherent characteristics of the image is solved, thereby improving the model training efficiency and the quality of generated images.
Patent Information
- Application Number
- CN202510673074.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies make it difficult to enhance image light while maintaining the original inherent characteristics of the image, which affects the accuracy of target detection and path planning of autonomous vehicles.
By obtaining multiple sample data pairs, predicting the noise of the target object under different light sources, and determining the fusion noise and reference noise based on the noise, the image generation model is trained to follow the physical laws of illumination changes and generate new images.
The model training efficiency and accuracy are improved, and the quality of the generated images does not affect the original inherent properties of the image, meets the target lighting conditions and maintains the inherent characteristics of the scene.
Smart Images

Figure CN120672882A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to image generation technology, and in particular to a model training method, image generation method, device, electronic device, computer-readable storage medium and computer program product. Background Art
[0002] During operation, autonomous vehicles need to accurately perceive their surroundings under a variety of complex lighting conditions. However, at night, in low-light conditions, or in conditions with uneven lighting, the quality of images captured by sensors can significantly degrade. This degradation in image quality can impact the accuracy of key autonomous vehicle tasks such as target detection and path planning. Currently, common solutions to these problems include data augmentation, low-light image enhancement, and light transformation techniques based on generative adversarial networks. However, these methods share a common problem: it is difficult to enhance image lighting while preserving the original inherent characteristics of the image. Summary of the Invention
[0003] The embodiments of the present application provide a model training method, an image generation method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the training efficiency and accuracy of the model and improve the quality of images generated by the trained image generation model.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides a model training method, which includes:
[0006] Acquire a plurality of sample data pairs, wherein the sample data pairs include first sample data of a target object under a first light source and second sample data of the target object under a second light source;
[0007] For each of the sample data pairs, predicting a first noise of the target object under the first light source and a second noise of the target object under the second light source by using an image generation model to be trained;
[0008] Based on the first noise and the second noise, determining, by the image generation model to be trained, a fusion noise of the target object under the first light source and the second light source;
[0009] Determining reference noise of the target object under the first light source and the second light source, and determining a loss value corresponding to the sample data pair based on the fusion noise and the reference noise;
[0010] The image generation model to be trained is trained based on the loss value corresponding to each pair of sample data.
[0011] The present invention provides an image generation method, which includes:
[0012] When the image generation timing is reached, determining an image to be processed including the target object and second light source information corresponding to the image to be processed, where the second light source information includes at least one of a second light source parameter of the target light source, second description information, and a second light source image;
[0013] determining, based on the image to be processed and the second light source information, a target noise of the image to be processed under the target light source using a trained image generation model, wherein the trained image generation model is obtained using the above-mentioned model training method;
[0014] The target noise is superimposed on the image to be processed through the trained image generation model to obtain a target image of the target object under the target light source.
[0015] The present invention provides a model training device, comprising:
[0016] an acquisition module, configured to acquire a plurality of sample data pairs, wherein the sample data pairs include first sample data of a target object under a first light source and second sample data of the target object under a second light source;
[0017] a first determination module, configured to determine, for each pair of sample data, a first noise of the target object under the first light source and a second noise of the target object under the second light source by using a to-be-trained image generation model;
[0018] The first determination module is further configured to determine, based on the first noise and the second noise, a fusion noise of the target object under the first light source and the second light source by using the image generation model to be trained;
[0019] The first determining module is further configured to determine a reference noise of the target object under the first light source and the second light source, and determine a loss value corresponding to the sample data pair based on the fusion noise and the reference noise;
[0020] A training module is used to train the image generation model to be trained based on the loss value corresponding to each pair of sample data.
[0021] An embodiment of the present application provides an image generating device, comprising:
[0022] a second determining module, configured to determine, when an image generation timing is reached, an image to be processed including the target object and second light source information corresponding to the image to be processed, wherein the second light source information includes at least one of a second light source parameter of the target light source, second description information, and a second light source image;
[0023] The second determination module is configured to determine the target noise of the image to be processed under the target light source using a trained image generation model based on the image to be processed and the second light source information, wherein the trained image generation model is obtained using the above-mentioned model training method;
[0024] The superposition module is used to superimpose the target noise onto the image to be processed through the trained image generation model to obtain a generated image of the target object under the target light source.
[0025] An embodiment of the present application provides an electronic device, comprising:
[0026] a memory for storing computer executable instructions or computer programs;
[0027] The processor is used to implement the model training method or image generation method provided in the embodiment of the present application when executing the computer-executable instructions or computer program stored in the memory.
[0028] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which is used to implement the model training method or image generation method provided in the embodiment of the present application when executed by a processor.
[0029] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the model training method or image generation method provided in the embodiment of the present application is implemented.
[0030] The embodiments of the present application have the following beneficial effects:
[0031] In an embodiment of the present application, multiple sample data pairs are acquired, each comprising first sample data of a target object under a first light source and second sample data of the target object under a second light source. For each sample data pair, a first noise of the target object under the first light source and a second noise of the target object under the second light source are predicted using a trained image generation model. Based on the first and second noises, the trained image generation model determines a fused noise of the target object under the first and second light sources. In this way, the trained image generation model predicts the first and second noises of the target object under different light sources, and then determines the fused noise of the target object under multiple light sources based on the first and second noises. This ensures that the fused noise conforms to the physical laws of linear illumination variation, thereby improving the accuracy of determining the fused noise. Subsequently, a reference noise of the target object under the first and second light sources is determined, and based on the fused noise and the reference noise, a loss value corresponding to the sample data pair is determined. The trained image generation model is trained based on the loss value corresponding to each sample data pair. This ensures that the image generation model fully learns the physical consistency of illumination variation, thereby improving model training efficiency and accuracy. Furthermore, the trained image generation model can generate images in accordance with physical laws of illumination changes, ensuring that when the image is edited and a new image is generated, the inherent properties of the image are not affected, thereby improving the quality of the generated image. Therefore, the embodiments of the present application can improve the training efficiency and accuracy of the model, and improve the quality of images generated by the trained image generation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a structural diagram of a data processing system provided in an embodiment of the present application;
[0033] Figure 2A This is a schematic diagram of the server structure provided by an embodiment of the present application;
[0034] Figure 2B This is another server structure diagram provided in an embodiment of the present application;
[0035] Figure 3A This is a flow chart of the model training method provided in the embodiment of the present application;
[0036] Figure 3B Schematic diagram of the noise prediction method provided in the embodiment of the present application;
[0037] Figure 3C Schematic diagram of the noise fusion method provided in the embodiment of the present application;
[0038] Figure 3D This is a flow chart of constructing a sample data pair provided in an embodiment of the present application;
[0039] Figure 4 Schematic diagram of the process of generating an image according to an embodiment of the present invention;
[0040] Figure 5 This is another flowchart of the model training method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0042] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0043] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0044] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0045] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0046] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0047] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0048] 1) Noise: Different light source conditions can cause different types and degrees of noise in images. In the embodiments of the present application, noise refers to the overall or local changes in brightness, color, and shadows of the target object in the image under the corresponding light source;
[0049] 2) Light source parameters: physical quantities used to describe the characteristics of a light source, including luminous flux (unit: lumens, reflecting the total amount of light emitted), luminous intensity (unit: candela, reflecting the ability to emit light in a specific direction), illuminance (unit: lux, measuring the brightness of the illuminated surface), color temperature (expressing the color characteristics of the light source in absolute temperature), color rendering index (measures the ability to reproduce color), spectral power distribution (describing the distribution of light power at different wavelengths), and beam angle (indicating the range of light source illumination);
[0050] 3) Diffusion model: A probability-based generation model that generates new data samples by simulating a forward diffusion process, gradually adding noise to the original data until the data becomes random noise. It then learns a reverse process to gradually recover the original data from the random noise.
[0051] 4) U-Net model: A deep learning model for image segmentation. Its core structure resembles the letter "U," consisting of an encoder (a downsampling path that extracts image features) and a decoder (an upsampling path that restores the image size). Its key innovation is the skip connection, which directly connects feature maps from different encoder layers to the corresponding decoder layers, while preserving both detailed information (such as edges and textures) and high-level semantic information (such as object categories), achieving highly accurate pixel-level segmentation.
[0052] 5) Multilayer Perceptron (MLP): An artificial neural network model consisting of an input layer, hidden layers, and an output layer. It operates by performing calculations through forward propagation and optimizing the model by adjusting weights using a backpropagation algorithm.
[0053] In existing technologies, image enhancement is typically achieved through data augmentation, low-light image enhancement, and light transformation techniques based on generative adversarial networks. Data augmentation primarily augments data through rotation, flipping, and noise addition, which has limited effectiveness in improving lighting conditions. Furthermore, operations such as random cropping can easily alter object shape and position. While increasing data diversity, this can affect the autonomous driving system's accurate scene perception, introducing errors that hinder target detection and path planning. Low-light image enhancement, while increasing image brightness and contrast, can also easily lose detail, such as blurring dark textures. Furthermore, adjusting color to enhance visual quality can cause distortion. Details like color and texture are crucial for recognizing traffic signs and lane markings. Distortion can bias autonomous driving's perception of the environment and affect decision-making and control. Light transformation techniques based on generative adversarial networks are complex to train and prone to problems such as mode collapse and vanishing gradients. During light transformation, the generator struggles to generate images that meet target lighting conditions while maintaining the inherent characteristics of the scene, potentially introducing artifacts or altering the true appearance of objects. Consequently, a common problem with related art methods is the difficulty in enhancing image lighting while maintaining the original inherent characteristics.
[0054] The embodiments of the present application provide a model training method, an image generation method, an apparatus, a device, a computer-readable storage medium, and a computer program product, which can improve the training efficiency and accuracy of the model, and the trained image generation model can generate images in accordance with the physical laws of illumination changes, and when the image is illuminated and a new image is generated, the original inherent properties of the image are not affected, thereby achieving the effect of satisfying the target light conditions and maintaining the inherent characteristics of the scene. The following describes an exemplary application of the electronic device provided by the embodiment of the present application. The electronic device provided by the embodiment of the present application can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and car terminals, and can also be implemented as servers. The following will describe an exemplary application of the electronic device when it is implemented as a server.
[0055] See also Figure 1 , Figure 1 is a structural diagram of a data processing system provided in an embodiment of the present application, Figure 1 The system involves a database 100, a server 200, a network 300 and a terminal 400. The terminal 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two. The trained prediction model can be stored in the database 100. The database 100 can be independent of the server 200 or deployed on the server 200. Figure 1 In the figure, the database 100 is shown as being independent of the server 200 .
[0056] During model training, server 200 can obtain the image generation model to be trained and multiple sample data pairs from database 100 or terminal 400. Then, for each sample data pair, the image generation model to be trained is used to predict the first noise of the target object under the first light source and the second noise under the second light source. Based on the first noise and the second noise, the image generation model to be trained is used to determine the fused noise of the target object under the first and second light sources. A reference noise of the target object under the first and second light sources is determined, and based on the fused noise and the reference noise, a loss value corresponding to the sample data pair is determined. The image generation model to be trained is trained based on the loss value corresponding to each sample data pair. The trained image generation model is then stored in database 100.
[0057] Afterwards, when the image generation timing arrives, server 200 determines an image to be processed that includes the target object, as well as second light source information corresponding to the image to be processed. The second light source information includes at least one of the second light source parameters, second descriptive information, and second light source image of the target light source. A trained image generation model is then retrieved from database 100. Based on the image to be processed and the second light source information, the trained image generation model is used to determine the target noise of the image to be processed under the target light source. The trained image generation model is used to superimpose the target noise onto the image to be processed, thereby obtaining a target image of the target object under the target light source. Server 200 then pushes the target image to terminal 400, which displays the target image through display interface 410 of terminal 400.
[0058] Take server 200 for model training as an example, see Figure 2A , Figure 2A This is a schematic diagram of the server structure provided by the embodiment of the present application. Figure 2A The server 200-1 shown includes: at least one processor 210-1, a memory 230-1 and at least one network interface 220-1. The various components in the server 200-1 are coupled together via a bus system 240-1. It is understood that the bus system 240-1 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 240-1 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 240-1 is not described in detail. Figure 2A Various buses are labeled as bus system 240 - 1 .
[0059] Processor 210-1 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0060] The memory 230-1 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, a hard drive, an optical drive, etc. The memory 230-1 may optionally include one or more storage devices that are physically remote from the processor 210-1.
[0061] The memory 230-1 includes volatile memory or nonvolatile memory, or may include both volatile and nonvolatile memory. The nonvolatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 230-1 described in the embodiments of the present application is intended to include any suitable type of memory.
[0062] In some embodiments, the memory 230 - 1 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0063] Operating system 231-1, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0064] The network communication module 232-1 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220-1. Exemplary network interfaces 220-1 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).
[0065] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2A The model training device 233 stored in the memory 230-1 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 2331, a prediction module 2332, a first determination model 2333, and a training module 2334. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0066] Take the server 200 for image generation as an example, see Figure 2B , Figure 2B This is another server structure diagram provided in an embodiment of the present application. Figure 2B The server 200-2 shown includes: at least one processor 210-2, a memory 230-2, and at least one network interface 220-2. The various components in the server 200-2 are coupled together via a bus system 240-2. It is understood that the bus system 240-2 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 240-2 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 240-2 is not described in detail. Figure 2B In the figure, various buses are labeled as bus system 240-2. The detailed description of processor 210-2 and memory 230-2 is as above, which will not be repeated here.
[0067] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2B The image generation device 234 stored in the memory 230-2 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a second determination module 2341 and an overlay module 2342. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0068] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the model training method or image generation method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0069] The model training method and image generation method provided in the embodiments of the present application will be explained in combination with the exemplary application and implementation of the terminal provided in the embodiments of the present application.
[0070] Below, the model training method provided by the embodiment of the present application is described. As mentioned above, the electronic device that implements the model training method of the embodiment of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.
[0071] See also Figure 3A , Figure 3A This is a flow chart of the model training method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.
[0072] In step 101, a plurality of sample data pairs are obtained.
[0073] Here, the sample data pair includes first sample data of the target object under the first light source, and second sample data of the target object under the second light source. The target object refers to the object that the unmanned vehicle needs to pay attention to during driving. For example, the target object can be a vehicle, including various types of vehicles such as cars, trucks, buses, motorcycles, etc. that are driving or parked on the road. The unmanned vehicle needs to monitor the position, speed and driving direction of surrounding vehicles in real time to avoid collisions and make reasonable path planning. The target object can also be road facilities, such as traffic lights, traffic signs (such as speed limit signs and stop and yield signs), road markings (lane lines, zebra crossings, etc.), curbs, etc. These elements are very important for the unmanned vehicle to identify road rules and determine the driving path. The target object can also be an obstacle, such as fallen rocks, branches, scattered objects, etc. on the road that may affect the driving safety of the unmanned vehicle. The target object can also be a pedestrian. Pedestrians are the objects that the unmanned vehicle needs to pay attention to during driving. The unmanned vehicle needs to identify the position, movement (such as walking, standing, waving, etc.) and intention of pedestrians in order to make safe decisions, such as slowing down and stopping. The first light source and the second light source represent two different light source conditions. For example, the first light source represents a "clear daytime" and the second light source represents a "rainy night." The first sample data may include an image of the target object under the first light source and light source information describing the first light source. The second sample data may include an image of the target object under the second light source and light source information describing the second light source.
[0074] In step 102 , for each sample data pair, a model to be trained is used to generate a first noise of the target object under a first light source and a second noise of the target object under a second light source.
[0075] Here, the image generation model to be trained refers to an untrained image generation model, and the image generation model to be trained is a diffusion model. For each sample data pair, by inputting the first sample data in the sample data pair into the image generation model to be trained, the image generation model to be trained can predict the first noise of the target object under the first light source; by inputting the second sample data in the sample data pair into the image generation model to be trained, the image generation model to be trained can predict the second noise of the target object under the second light source. Different light source conditions will result in different types and degrees of noise. Noise refers to the overall or local changes in brightness, color, and shadow of the target object under the corresponding light source.
[0076] In some embodiments, the first sample data includes a first sample image and first light source information, the second sample data includes a second sample image and second light source information, and the image generation model to be trained includes a noise prediction module. Figure 3B , step 102 can be implemented through steps 1021 to 1023, including:
[0077] In step 1021, the first sample image and the second sample image are respectively encoded by the noise prediction module in the image generation model to be trained, and a first encoding result and a second encoding result are obtained correspondingly.
[0078] Here, the first sample image refers to an image of the target object under the first light source, and the second sample image refers to an image of the target object under the second light source. The noise prediction module is a U-Net model within the image generation model to be trained. The U-Net model is a deep learning model for image segmentation. Its core structure resembles the letter "U," consisting of an encoder and a decoder. The encoder encodes the first and second sample images, respectively, producing first and second encoding results.
[0079] In step 1022, prediction processing is performed on the first encoding result and the first light source information to obtain a first noise of the target object under the first light source.
[0080] Here, the image generation model to be trained can predict first noise of the target object under the first light source by performing prediction processing on the first encoding result and the first light source information. The first noise refers to changes in overall or local brightness, color, and shadow of the target object under the first light source.
[0081] In step 1023, prediction processing is performed on the second encoding result and the second light source information to obtain second noise of the target object under the second light source.
[0082] Here, the image generation model to be trained can predict the first noise of the target object under the first light source by performing prediction processing on the second encoding result and the second light source information. The second noise refers to the overall or local changes in brightness, color, and shadow of the target object under the second light source.
[0083] In an embodiment of the present application, the first sample data includes a first sample image and first light source information, the second sample data includes a second sample image and second light source information, and the image generation model to be trained includes a noise prediction module. The first sample image and the second sample image are encoded and processed respectively by the noise prediction module in the image generation model to be trained, and the first encoding result and the second encoding result are obtained correspondingly; the first encoding result and the first light source information are predicted and processed to obtain the first noise of the target object under the first light source; the second encoding result and the second light source information are predicted and processed to obtain the second noise of the target object under the second light source. In this way, by performing noise prediction on the sample image and the light source information at the same time, the accuracy of the noise prediction can be improved. Moreover, by performing noise prediction on the noise prediction module in the image generation model to be trained, the prediction efficiency can be improved, the data processing process can be simplified, and the model training efficiency can be improved.
[0084] In step 103 , based on the first noise and the second noise, the fusion noise of the target object under the first light source and the second light source is determined by using the image generation model to be trained.
[0085] Here, the first noise and the second noise are fused by the image generation model to be trained, thereby obtaining the fused noise of the target object under the first light source and the second light source.
[0086] In some embodiments, the image generation model to be trained includes a noise fusion module. Figure 3C , step 103 may be implemented through steps 1031 to 1033, including:
[0087] In step 1031, the first noise and the second noise are fused by a noise fusion module in the image generation model to be trained to obtain fused noise.
[0088] Here, the noise fusion module is a multilayer perceptron in the image generation model to be trained. The multilayer perceptron consists of an input layer, two hidden layers, and an output layer. The input layer receives the first and second noise data and formats them into a format suitable for input to the MLP. For example, it concatenates them into a single feature vector, the fused noise, which is then provided as input to the hidden layer.
[0089] In step 1032, feature extraction and nonlinear transformation are performed on the fused noise to obtain a hidden feature representation of the fused noise.
[0090] Here, the fused noise is passed to the first hidden layer. In the first hidden layer, each neuron receives the fused noise and computes it based on a set of weights and biases. Specifically, each neuron calculates the dot product of the fused noise with the neuron's weight vector, then adds the bias and applies a nonlinear transformation to the result through an activation function. In this way, the first hidden layer performs preliminary feature extraction and nonlinear transformation on the fused noise, obtaining the corresponding feature representation of the first hidden layer. The output of the first hidden layer serves as the input to the second hidden layer. The second hidden layer similarly processes the input, further extracting features and performing nonlinear transformation through weight and bias operations and the application of the activation function. This process is similar to the first hidden layer, but because its input is already processed by the first layer, it can learn more complex and abstract feature representations. These feature representations are related to the first and second noise inputs and the relationship between them. The second hidden layer outputs the hidden feature representation of the fused noise.
[0091] In step 1033, activation processing is performed on the hidden feature representation to obtain fused noise of the target object under the first light source and the second light source.
[0092] Here, the hidden feature representation is passed to the output layer. In the output layer, neurons calculate the hidden feature representation based on another set of weights and biases, thereby activating the hidden feature representation and obtaining the output value, which is the fused noise of the target object under the first and second light sources. The activation function of the output layer may vary depending on the nature of the problem. If the noise value is continuous, an activation function may not be used or a linear activation function may be used. If the noise value needs to meet a certain range (for example, between 0 and 1), an activation function such as Sigmoid or Tanh may be used.
[0093] In an embodiment of the present application, the image generation model to be trained includes a noise fusion module. The noise fusion module in the image generation model to be trained fuses the first noise and the second noise to obtain a fused noise. Feature extraction and nonlinear transformation are performed on the fused noise to obtain a hidden feature representation of the fused noise. The hidden feature representation is activated to obtain the fused noise of the target object under the first and second light sources. In this way, the model can learn the combined effects of noise under multiple light sources, thereby more realistically simulating complex lighting environments and improving the fidelity and diversity of generated images. This also enhances the model's robustness to lighting changes and its generalization capabilities, supporting various tasks in multi-light source scenarios and thus improving model performance.
[0094] In step 104 , the reference noise of the target object under the first light source and the second light source is determined, and based on the fused noise and the reference noise, the loss value corresponding to the sample data pair is determined.
[0095] Here, reference noise refers to the noise that the target object should normally experience under the first and second light sources. The loss value is used to measure the difference between the fused noise and the reference noise. The L2 norm can be used to construct the model's loss function. The difference between the two values is then calculated based on the loss function. This loss value serves as feedback to guide model parameter optimization, thereby improving image generation quality and model generalization capabilities.
[0096] In some embodiments, the reference noise of the target object under the first light source and the second light source may be determined by the following process, including:
[0097] A first noise of the target object under a first light source and a second noise of the target object under a second light source are obtained; the first noise and the second noise are linearly superimposed to obtain a reference noise.
[0098] Here, since the first noise of the target object under the first light source and the second noise under the second light source have already been predicted by the noise prediction module in the image generation model to be trained, the first noise and the second noise can be directly obtained. Then, according to the physical light transmission theory, the change in the appearance image of the target object under different light sources is linear, that is, the superposition of the effects of multiple light sources can be directly equal to the synthesis of the effects of the individual processing under these multiple light sources. In other words, the reference noise of the target object under the first light source and the second light source is equal to the linear superposition result of the first noise and the second noise. Therefore, the first noise and the second noise are linearly superimposed to obtain the reference noise.
[0099] In the embodiment of the present application, a first noise level of a target object under a first light source and a second noise level of the target object under a second light source are obtained; the first noise level and the second noise level are linearly superimposed to obtain a reference noise level. This allows for efficient and accurate determination of the reference noise level of the target object under the first and second light sources based on physical light transmission theory, improving the accuracy and efficiency of determining the reference noise level.
[0100] In some embodiments, the loss value corresponding to the sample data pair can be determined by the following process:
[0101] Determine the noise difference matrix between the fusion noise and the reference noise; obtain the mask matrix corresponding to the sample data pair, perform weighted processing on the noise difference matrix based on the mask matrix to obtain a weighted noise difference matrix; quantize the weighted noise difference matrix to obtain the loss value corresponding to the sample data pair.
[0102] Here, by substituting the fusion noise and the reference noise into the loss function, the loss value corresponding to the sample data pair is determined, see formula (1):
[0103]
[0104] Among them, L ltc Represents the loss value corresponding to the sample data pair. ∈ L1+L2 represents the reference noise under the first light source L1 and the second light source L2, φ(∈ L1 ,∈ L2 ) represents the fusion noise under the first light source L1 and the second light source L2, ∈ L1+L2 -φ(∈ L1 ,∈ L2 =(\begin{aligned}) represents the noise difference matrix between the fused noise and the reference noise. M is the mask matrix, representing the foreground mask, which ensures that the loss is applied only to the foreground portion of the image, i.e., the portion where the target object is located. The mask matrix can be pre-determined through techniques such as image preprocessing and feature extraction, object detection and localization, and semantic segmentation. The noise difference matrix is then weighted based on the mask matrix to obtain a weighted noise difference matrix. This weighted noise difference matrix is then quantized using the L2 norm to obtain the loss value corresponding to the sample data pair.
[0105] In the embodiment of the present application, a noise difference matrix between the fused noise and the reference noise is determined; a mask matrix corresponding to a sample data pair is obtained, and the noise difference matrix is weighted based on the mask matrix to obtain a weighted noise difference matrix; the weighted noise difference matrix is quantized to obtain the loss value corresponding to the sample data pair. In this way, the mask matrix enhances the flexibility of determining the model loss value, and the L2 norm is used to ensure the smoothness of the loss value and its robustness to outliers, comprehensively reflecting the difference between the model prediction and the true value, thereby improving the accuracy of the loss value determination.
[0106] In step 105 , the image generation model to be trained is trained based on the loss value corresponding to each sample data pair.
[0107] Here, based on the loss value corresponding to each sample data pair, the loss value is back-propagated to the image generation model to be trained. Based on the loss value, the parameters of the image generation model to be trained are adjusted using a gradient descent algorithm. The above steps are then repeated for the loss value corresponding to the next sample data pair until the training termination condition is met, resulting in the trained image generation model. The training termination condition can be when the loss value falls below a loss threshold or when the difference between the current loss value and the previous adjacent loss value is less than a preset difference.
[0108] In an embodiment of the present application, multiple sample data pairs are acquired, each comprising first sample data of a target object under a first light source and second sample data of the target object under a second light source. For each sample data pair, a first noise of the target object under the first light source and a second noise of the target object under the second light source are predicted using a trained image generation model. Based on the first and second noises, the trained image generation model determines a fused noise of the target object under the first and second light sources. In this way, the trained image generation model predicts the first and second noises of the target object under different light sources, and then determines the fused noise of the target object under multiple light sources based on the first and second noises. This ensures that the fused noise conforms to the physical laws of linear illumination variation, thereby improving the accuracy of determining the fused noise. Subsequently, a reference noise of the target object under the first and second light sources is determined, and based on the fused noise and the reference noise, a loss value corresponding to the sample data pair is determined. The trained image generation model is trained based on the loss value corresponding to each sample data pair. This ensures that the image generation model fully learns the physical consistency of illumination variation, thereby improving model training efficiency and accuracy. Furthermore, the trained image generation model can generate images in accordance with physical laws of illumination changes, ensuring that when the image is edited and a new image is generated, the inherent properties of the image are not affected, thereby improving the quality of the generated image. Therefore, the embodiments of the present application can improve the training efficiency and accuracy of the model, and improve the quality of images generated by the trained image generation model.
[0109] In some embodiments, see Figure 3D , multiple sample data pairs can be constructed through steps 201 to 203, including:
[0110] In step 201 , a plurality of sample images are acquired.
[0111] Here, a sample image is an image of the target object under a sample light source. A sample light source is a light source chosen as a representative or reference, and can have arbitrary spectral characteristics, intensity distribution, or other optical parameters. Sample images can be obtained from three main sources: sample images acquired using a single light source to simulate various lighting conditions; sample images generated using the Objaverse dataset through a custom image rendering pipeline; and high-quality real-world images acquired from the internet or databases based on specific application scenarios and requirements.
[0112] In step 202 , for each sample image, first light source information corresponding to the sample image is determined.
[0113] Here, the first light source information is used to describe the sample light source corresponding to the sample image, and may include light source parameters of the sample light source, description information of the light source, a reference image of the light source, and the like.
[0114] In some embodiments, the first light source information corresponding to the sample image may be determined by the following process:
[0115] Based on the image type of the sample image, first light source parameters of the sample light source corresponding to the sample image are obtained; first description information and a first light source image of the sample light source corresponding to the sample image are determined; and the first light source parameters, the first description information and the first light source image are determined as the first light source information.
[0116] Here, the image types of the sample images include synthetic types and real types. The synthetic type means that the sample image is an image generated by computer graphics technology, digital image processing software or artificial intelligence algorithms, and the real type means that the sample image is an image obtained by shooting and recording scenes and objects in the real world with optical equipment such as a camera. The first light source parameters include parameter information such as the intensity, color, position, and angle of the light source. The first descriptive information is text information used to describe the sample light source, such as a paragraph of text such as "hazy morning". The first light source image is a reference image used to simulate the sample light source. For example, the description of "a day with strong sun" through text is very vague, so the text description can be replaced by simulating an image of light of "a day with strong sun", making the description of the light source more accurate and intuitive.
[0117] In an embodiment of the present application, based on the image type of a sample image, first light source parameters of a sample light source corresponding to the sample image are obtained; first descriptive information and a first light source image corresponding to the sample image are determined; and the first light source parameters, first descriptive information, and first light source image are determined as first light source information. Thus, determining the first light source parameters, first descriptive information, and first light source image as first light source information allows for a comprehensive and accurate presentation of light source characteristics from multiple dimensions, including numerical quantification, textual qualitative description, and visual intuition, thereby increasing the diversity and accuracy of the first light source information.
[0118] In some embodiments, the first light source parameter of the sample light source corresponding to the sample image may be obtained through the following process:
[0119] When the image type is a synthetic type, the first light source parameter is determined from the synthetic information of the sample image; when the image type is a real type, a light source analysis is performed on the sample image to obtain the first light source parameter.
[0120] Here, synthesis information refers to the various information related to image generation and light source settings involved in the image synthesis process. When the image type is synthetic, the first light source parameters can be determined by reading the light source setting parameters in the 3D modeling or image synthesis software, the illumination simulation principles of the analysis algorithm, and the light source information provided by the reference synthetic material. When the image type is real, the first light source parameters can be determined by comprehensive analysis using analysis tools in image editing software and professional optical measurement equipment, combined with shadow characteristics of objects in the scene, shooting equipment parameters, and environmental information.
[0121] In the embodiments of the present application, when the image type is synthetic, the first light source parameters can be determined directly from the synthetic information of the sample image, thereby improving the efficiency of determining the first light source parameters. When the image type is real, the first light source parameters are obtained by performing light source analysis on the sample image, which can improve the accuracy of determining the first light source parameters. In this way, the first light source information can be determined specifically for different types of images, which can improve the overall data processing efficiency and accuracy.
[0122] In some embodiments, the first description information of the sample light source and the first light source image corresponding to the sample image may be determined by the following process:
[0123] Text information for describing the sample light source is obtained, and the text information is determined as first description information; a simulation image for simulating the sample light source is obtained, and the simulation image is determined as a first light source image.
[0124] Here, for the text information describing the sample light source, the relevant parameter description can be extracted from the software parameter setting instructions, algorithm documentation, and the description of the material itself of the synthetic image, or in the real image scene, the text description can be formed by using the optical equipment measurement data, image editing software analysis results, and environmental characteristics summary to form a text description, which is then determined as the first description information. For the simulated image used to simulate the sample light source, the synthetic image can be re-rendered and output in professional software based on its light source parameters. For the real image, different light source effects can be simulated through image processing software or an image with an approximate light source performance can be generated using a computer vision algorithm to serve as the first light source image. In this way, the light source parameters, properties, and other details can be accurately recorded through text, and the lighting effect can be intuitively presented with the help of simulated images.
[0125] In step 203, the sample image and the first light source information corresponding to the sample image are determined as sample data of the target object under the sample light source.
[0126] Here, for each sample image, the sample image and the first light source information corresponding to the sample image are determined as sample data of the target object under the sample light source.
[0127] In step 204 , any two different sample data are constructed into a sample data pair to obtain a plurality of sample data pairs.
[0128] Here, by traversing multiple sample data, a loop algorithm can be used to select two different sample data in turn for combination. Each time a combination is completed, a sample data pair is generated. This operation is continued until all possible sample data combinations are traversed, thereby obtaining multiple sample data pairs.
[0129] In an embodiment of the present application, multiple sample images are obtained, where the sample images are images of a target object under a sample light source. For each sample image, the first light source information corresponding to the sample image is determined. The sample image and the first light source information corresponding to the sample image are determined as sample data of the target object under the sample light source. Any two different sample data are constructed into a sample data pair to obtain multiple sample data pairs. In this way, by combining sample data under different light sources to form multiple sample data pairs, the multiple sample data pairs can cover more sample image and light source features, enriching the information in the sample data pairs, helping the model learn the mapping relationship between the appearance image of the target object and the light source, improving the generalization ability of the model, and thereby improving the efficiency and accuracy of model training.
[0130] The following describes the image generation method provided by the embodiments of the present application. As previously mentioned, the electronic device implementing the image generation method of the embodiments of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.
[0131] See also Figure 4 , Figure 4 This is a flow chart of the image generation method provided in the embodiment of the present application, which will be combined with Figure 4 The steps shown are explained.
[0132] In step 301, when the image generation timing is reached, an image to be processed including a target object and second light source information corresponding to the image to be processed are determined.
[0133] Here, the image generation timing is determined to have arrived when an image generation instruction is received, when an image to be processed and second light source information are received, or when a preset time has elapsed. The second light source information includes at least one of second light source parameters of the target light source, second descriptive information, and a second light source image. The second light source parameters include parameter information such as the intensity, color, position, and angle of the target light source; the first descriptive information is text information describing the target light source; and the first light source image is a simulated image used to simulate the target light source.
[0134] In step 302, based on the image to be processed and the second light source information, the target noise of the image to be processed under the target light source is determined by using the trained image generation model.
[0135] Here, the trained image generation model is obtained by the model training method provided in other embodiments of the present application. By inputting the image to be processed and the second light source information into the trained image generation model, the model can determine the target noise of the image to be processed under the target light source. Among them, the target light source corresponds to the second light source information, and refers to the specific lighting conditions defined by the input second light source information. Target noise refers to the overall or local changes in brightness, color, and shadow of the target object in the image to be processed under the target light source.
[0136] In step 303, the target noise is superimposed on the image to be processed through the trained image generation model to obtain a target image of the target object under the target light source.
[0137] Here, the trained image generation model uses the acquired knowledge through internal neural network analysis and transformation to predict the appearance of the target object under the target light source. This model then superimposes the target noise onto the processed image to produce the target image of the target object under the target light source. Furthermore, the generated image can be further post-processed to optimize the visual effect, such as brightness and color, as needed.
[0138] In an embodiment of the present application, when the timing of image generation is reached, an image to be processed including a target object and second light source information corresponding to the image to be processed are determined, the second light source information including at least one of the second light source parameters, second description information, and second light source image of the target light source; based on the image to be processed and the second light source information, the target noise of the image to be processed under the target light source is determined through a trained image generation model, the trained image generation model being obtained through the model training method provided by other embodiments; the target noise is superimposed on the image to be processed through the trained image generation model to obtain a target image of the target object under the target light source. In this way, the trained image generation model can generate images in accordance with the illumination changes of physical laws, and when the image is illuminated and a new image is generated, the original inherent properties of the image will not be affected, thereby achieving the effect of both meeting the target lighting conditions and maintaining the inherent characteristics of the scene, thereby improving the quality of the target image.
[0139] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0140] During operation, autonomous vehicles need to accurately perceive their surroundings under a variety of complex lighting conditions. However, at night, in low-light conditions, or in conditions with uneven lighting, the quality of images captured by sensors can significantly degrade. This degradation in image quality can impact the accuracy of key autonomous vehicle tasks such as target detection and path planning. Currently, common solutions to these problems include data augmentation, low-light image enhancement, and light transformation techniques based on generative adversarial networks. However, these methods share a common problem: it is difficult to enhance image lighting while preserving the original inherent characteristics of the image.
[0141] In order to solve the above problems, the embodiment of the present application proposes an image light enhancement method based on a diffusion model. Due to its randomness and the process of noise introduction, the diffusion model can easily cause the inherent properties of the image (such as albedo, material, details, etc.) to change when generating the image. Especially when dealing with complex lighting modification tasks, such as adding or modifying shadows, highlights, etc., the model needs to retain these inherent properties (such as color, texture, etc.), otherwise it will cause unacceptable image effects. The embodiment of the present application aims to solve the problem of maintaining the inherent image properties when the image light is changed. By introducing light transmission consistency constraints and multi-data source training strategies, the accuracy and stability problems of image lighting editing are effectively solved, while improving the generalization ability of the model, so that it can be effectively edited under different lighting conditions without affecting the details and inherent properties of the image.
[0142] This embodiment of the application proposes an image light enhancement method, mainly based on the regularization constraint of physical light transport consistency (LTC), for the task of light enhancement of unmanned vehicle data. The goal of this technical solution is to preserve the inherent properties of objects (such as albedo and material details) from being affected by lighting modifications, while ensuring the physical consistency of lighting changes. The following is the specific implementation process of this embodiment of the application:
[0143] Step 1: Sample image collection:
[0144] The data sources of the sample images used in the embodiments of this application mainly include the following three types:
[0145] (1) Real Light Stage data: 20,000 sample images captured using a single light source (One-Light-At-a-Time, OLAT), capable of simulating a variety of lighting conditions;
[0146] (2) 3D rendering data: Using the Objaverse dataset, approximately 4 million sample images were generated through a custom image rendering pipeline;
[0147] (3) Real images: high-quality real images obtained from the Internet or database according to specific application scenarios and requirements.
[0148] Step 2: Construct sample data pairs:
[0149] All sample images are processed uniformly into a standard format, and each sample image is processed according to {I L ,L,text,I d} data format to obtain the first light source information corresponding to each sample image. L represents the image of the target object under the sample light source, that is, the sample image; L represents the first light source parameter of the sample light source; text represents the first description information of the sample light source corresponding to the sample image; I d The first light source image representing the sample light source corresponding to the sample image illuminates the added image.
[0150] Then, each sample image and the first light source information corresponding to the sample image are determined as sample data of the target object under this sample light source. After that, any two different sample data are constructed into a sample data pair to obtain multiple sample data pairs. Therefore, the sample data pair includes the first sample data of the target object under the first light source and the second sample data of the target object under the second light source.
[0151] Step 3: Model training:
[0152] See also Figure 5 , Figure 5 This is another flow chart of the model training method provided by the embodiment of the present application. For each sample data pair, the encoder 51 in the image generation model to be trained generates the first sample image I of the first light source. d1 , the first light source image I of the first light source d1 , a second sample image I of the second light source d2 , the second light source image I of the second light source d2 Then, the encoding result and the first light source L1, the first description information text1 of the first light source, the second light source L2, and the first description information text2 of the second light source are input into the noise prediction model 52, and the first noise ∈ L1 , and the second noise ∈ of the target object under the second light source L2 L2 . After that, the first noise ∈ L1 and the second noise ∈ L2 The input noise fusion module 53 obtains the fused noise 54, and according to the first noise ∈ L1 and the second noise ∈ L2Determine reference noise 55. Finally, determine loss value 56 corresponding to the sample data pair based on fusion noise 54 and reference noise 55. Afterwards, train the image generation model to be trained using loss value 56 corresponding to each sample data pair until the training end condition is met, thereby obtaining a trained image generation model.
[0153] Here, we first select a diffusion model (such as Stable Diffusion) as the image generation model to be trained. Alternatively, we can use an existing powerful basic model, such as SDXL, as the initialization model. Then, for each sample data pair, we use the image generation model to be trained to predict the first noise of the target object under the first light source and the second noise under the second light source, as shown in formulas (2) and (3):
[0154] ∈ L1 =δ(ε(I L1 ),L1,text1,ε(I d1 ))(2)
[0155] ∈ L2 =δ(ε(I L2 ),L2,text2,ε(I d2 ))(3)
[0156] Among them, L1 is the first light source, I L1 refers to the first sample image of the target object under the first light source L1, text1 represents the first description information of the first light source, I d1 represents the first light source image of the first light source; L2 is the second light source, I L2 refers to the second sample image of the target object under the second light source L2, text2 represents the first description information of the second light source, I d1 represents the first light source image of the second light source. ε represents the image encoder in the image generation model to be trained, and δ is the noise prediction module in the image generation model to be trained. ∈ L1 Refers to the first noise of the target object under the first light source L1, ∈ L2 It refers to the second noise of the target object under the second light source L2.
[0157] Then, the first noise and the second noise are fused by the multilayer perceptron, i.e., the noise fusion module in the image generation model to be trained, to obtain the fused noise φ(∈ L1 ,∈ L2 ).
[0158] According to the physical theory of light transmission, the appearance of the target object under different light sources changes linearly, that is, the superposition of the effects of multiple light sources can be directly equal to the synthesis of the effects of the multiple light sources processed separately. In other words, if there are two different light sources L1 and L2, the image I of the object under these two light sources is L1 and I L2 The linear superposition of should be equal to the image under L1+L2 mixed illumination. On this basis, given a certain ambient light source L (such as the direction, intensity, color of the light), the image generated by it satisfies the linear mapping relationship with the light source L, so formula (4) can be obtained:
[0159] I L =T·L(4)
[0160] Among them, T is a linear transformation matrix that has nothing to do with the image resolution, I L It refers to the image of the target object under the light source L.
[0161] By introducing the consistency constraint, we can derive the linear relationship between the image and the light source under the light source combination condition, as shown in formula (5):
[0162] I L1+L2 =I L1 +I L2 (5)
[0163] Among them, L1 and L2 are two different light sources, I L1 Refers to the first sample image of the target object under the first light source L1, I L2 It refers to the second sample image of the target object under the second light source L2, I L1+L2 It refers to the image of the target object under the mixed illumination of the first light source L1 and the second light source L2.
[0164] Correspondingly, by transforming formula (5), the reference noise of the target object under the mixed illumination of the first light source L1 and the second light source L2 can be obtained, see formula (6):
[0165] ∈ L1+L2 =∈ L1 +∈ L2 (6)
[0166] Among them, ∈ L1 It refers to the first noise of the target object under the first light source L1, I L2 Refers to the second noise of the target object under the second light source L2, ∈ L1+L2 It refers to the reference noise of the target object under the mixed illumination of the first light source L1 and the second light source L2.
[0167] Afterwards, the L2 norm is used to construct the loss function of the model, see formula (1):
[0168]
[0169] Among them, L ltc Represents the loss value, M is the mask matrix, which represents the foreground mask and is used to ensure that the loss only acts on the foreground part of the image.
[0170] Through this regularization constraint, the model learns a mapping relationship that ensures that the generated image adheres to physically-based lighting variations and avoids destroying the inherent properties of the image. Subsequently, a backpropagation algorithm is used to adjust the model parameters based on the calculated loss until the training end condition is met, resulting in a trained image generation model.
[0171] The embodiment of the present application proposes a method for light enhancement of image data for unmanned vehicles, which ensures the physical consistency of lighting changes by regularizing constraints based on the consistency of physical light transmission and forcing the training process to maintain "linear combination of appearances under different light sources = appearance under mixed light sources". Among them, light transmission refers to the process of propagation and interaction of light when passing through the surface of an object or scene. The embodiment of the present application utilizes the characteristic that "the superposition of the effects of multiple light sources can be directly equal to the synthesis of the effects of these multiple light sources processed separately. Only the lighting part of the image will be modified, and other inherent properties will not be changed at all" to continuously change the light source during the training process of the diffusion model. Ensure that the lighting editing operations learned by the constraint model will not affect these key inherent properties in the image.
[0172] The following continues to describe the exemplary structure of the model training method device 233 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2A As shown, the software modules stored in the model training method device 233 of the memory 230-1 may include:
[0173] An acquisition module 2331 is configured to acquire a plurality of sample data pairs, wherein the sample data pairs include first sample data of a target object under a first light source and second sample data of the target object under a second light source;
[0174] A prediction module 2332 is configured to determine, for each sample data pair, a first noise of the target object under the first light source and a second noise of the target object under the second light source by using a to-be-trained image generation model;
[0175] A first determining module 2333 is configured to determine, based on the first noise and the second noise, a fusion noise of the target object under the first light source and the second light source by using the image generation model to be trained;
[0176] The first determining module 2333 is further configured to determine a reference noise of the target object under the first light source and the second light source, and determine a loss value corresponding to the sample data pair based on the fused noise and the reference noise;
[0177] The training module 2334 is used to train the image generation model to be trained based on the loss value corresponding to each pair of sample data.
[0178] In some embodiments, the prediction module 2332 is further used to encode the first sample image and the second sample image respectively through the noise prediction module in the image generation model to be trained, and obtain a first encoding result and a second encoding result respectively; perform prediction processing on the first encoding result and the first light source information to obtain a first noise of the target object under the first light source; and perform prediction processing on the second encoding result and the second light source information to obtain a second noise of the target object under the second light source.
[0179] In some embodiments, the first determination module 2333 is further used to fuse the first noise and the second noise through the noise fusion module in the image generation model to be trained to obtain fused noise; perform feature extraction and nonlinear transformation on the fused noise to obtain a hidden feature representation of the fused noise; and activate the hidden feature representation to obtain the fused noise of the target object under the first light source and the second light source.
[0180] In some embodiments, the first determination module 2333 is further configured to obtain a first noise of the target object under the first light source and a second noise of the target object under the second light source; and linearly superimpose the first noise and the second noise to obtain the reference noise.
[0181] In some embodiments, the first determination module 2333 is further used to determine the noise difference matrix between the fused noise and the reference noise; obtain the mask matrix corresponding to the sample data pair, perform weighted processing on the noise difference matrix based on the mask matrix to obtain a weighted noise difference matrix; and quantize the weighted noise difference matrix to obtain a loss value corresponding to the sample data pair.
[0182] In some embodiments, the acquisition module 2331 is also used to acquire multiple sample images, where the sample images are images of the target object under a sample light source; for each of the sample images, the first light source information corresponding to the sample image is determined; the sample image and the first light source information corresponding to the sample image are determined as sample data of the target object under the sample light source; and any two different sample data are constructed into the sample data pair to obtain multiple sample data pairs.
[0183] In some embodiments, the acquisition module 2331 is further used to acquire first light source parameters of the sample light source corresponding to the sample image based on the image type of the sample image; determine first descriptive information and first light source image of the sample light source corresponding to the sample image; and determine the first light source parameters, the first descriptive information and the first light source image as the first light source information.
[0184] In some embodiments, the acquisition module 2331 is further used to determine the first light source parameters from the synthesis information of the sample image when the image type is a synthetic type; and to perform light source analysis on the sample image to obtain the first light source parameters when the image type is a real type.
[0185] In some embodiments, the acquisition module 2331 is further used to acquire text information used to describe the sample light source, and determine the text information as the first description information; acquire a simulation image used to simulate the sample light source, and determine the simulation image as the first light source image.
[0186] The following continues to describe the exemplary structure of the image generation device 234 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2B As shown, the software modules stored in the image generation device 234 of the memory 230-2 may include:
[0187] A second determining module 2341 is configured to determine, when an image generation timing is reached, an image to be processed including the target object and second light source information corresponding to the image to be processed, where the second light source information includes at least one of a second light source parameter of the target light source, second description information, and a second light source image;
[0188] The second determination module 2341 is configured to determine the target noise of the image to be processed under the target light source using a trained image generation model based on the image to be processed and the second light source information, wherein the trained image generation model is obtained using the model training method provided in other embodiments of the present application;
[0189] The superposition module 2342 is configured to superimpose the target noise onto the image to be processed by using the trained image generation model to obtain a generated image of the target object under the target light source.
[0190] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the model training method or image generation method described in the present invention.
[0191] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the model training method or image generation method provided in the embodiment of the present application, for example, Figure 3A The model training method shown or Figure 4 The image generation method is shown.
[0192] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0193] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0194] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored in part of a file that stores other programs or data, e.g., in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0195] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0196] In summary, through the embodiments of the present application, multiple sample data pairs are obtained, each of which includes first sample data of a target object under a first light source and second sample data of the target object under a second light source. For each sample data pair, the image generation model to be trained predicts the first noise of the target object under the first light source and the second noise of the target object under the second light source. Based on the first noise and the second noise, the image generation model to be trained determines the fused noise of the target object under the first and second light sources. In this way, the image generation model to be trained predicts the first noise and the second noise of the target object under different light sources, and then determines the fused noise of the target object under multiple light sources based on the first noise and the second noise. This ensures that the fused noise and the first noise and the second noise conform to the physical laws of linear illumination variation, thereby improving the accuracy of determining the fused noise. Subsequently, the reference noise of the target object under the first and second light sources is determined, and based on the fused noise and the reference noise, the loss value corresponding to the sample data pair is determined. The image generation model to be trained is trained based on the loss value corresponding to each sample data pair. In this way, it is ensured that the image generation model fully learns the physical consistency of illumination variation, thereby improving the efficiency and accuracy of model training. Furthermore, the trained image generation model can generate images in accordance with physical laws of illumination changes, ensuring that when the image is edited and a new image is generated, the inherent properties of the image are not affected, thereby improving the quality of the generated image. Therefore, the embodiments of the present application can improve the training efficiency and accuracy of the model, and improve the quality of images generated by the trained image generation model.
[0197] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Acquire a plurality of sample data pairs, wherein the sample data pairs include first sample data of a target object under a first light source and second sample data of the target object under a second light source; For each of the sample data pairs, predicting a first noise of the target object under the first light source and a second noise of the target object under the second light source by using an image generation model to be trained; Based on the first noise and the second noise, determining, by the image generation model to be trained, a fusion noise of the target object under the first light source and the second light source; Determining reference noise of the target object under the first light source and the second light source, and determining a loss value corresponding to the sample data pair based on the fusion noise and the reference noise; The image generation model to be trained is trained based on the loss value corresponding to each pair of sample data.
2. The method according to claim 1, characterized in that The first sample data includes a first sample image and first light source information, the second sample data includes a second sample image and second light source information, and the image generation model to be trained includes a noise prediction module. The predicting, by using the image generation model to be trained, a first noise of the target object under the first light source and a second noise under the second light source, comprises: encoding the first sample image and the second sample image respectively by using the noise prediction module in the image generation model to be trained, and obtaining a first encoding result and a second encoding result respectively; performing prediction processing on the first encoding result and the first light source information to obtain a first noise of the target object under the first light source; Prediction processing is performed on the second encoding result and the second light source information to obtain second noise of the target object under the second light source.
3. The method according to claim 1, characterized in that The image generation model to be trained includes a noise fusion module, The determining, based on the first noise and the second noise, by using the image generation model to be trained, the fusion noise of the target object under the first light source and the second light source includes: fusing the first noise and the second noise using a noise fusion module in the image generation model to be trained to obtain fused noise; Performing feature extraction and nonlinear transformation processing on the fused noise to obtain a hidden feature representation of the fused noise; An activation process is performed on the hidden feature representation to obtain fused noise of the target object under the first light source and the second light source.
4. The method according to claim 1, wherein The determining the reference noise of the target object under the first light source and the second light source includes: Acquire a first noise of the target object under the first light source, and a second noise of the target object under the second light source; The first noise and the second noise are linearly superimposed to obtain the reference noise.
5. The method according to claim 1, wherein The determining, based on the fusion noise and the reference noise, a loss value corresponding to the sample data pair, includes: determining a noise difference matrix between the fused noise and the reference noise; Obtaining a mask matrix corresponding to the sample data pair, and performing weighted processing on the noise difference matrix based on the mask matrix to obtain a weighted noise difference matrix; The weighted noise difference matrix is quantized to obtain a loss value corresponding to the sample data pair.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Acquire a plurality of sample images, where the sample images are images of the target object under a sample light source; For each of the sample images, determining first light source information corresponding to the sample image; determining the sample image and the first light source information corresponding to the sample image as sample data of the target object under the sample light source; Any two different sample data are constructed into the sample data pair to obtain a plurality of the sample data pairs.
7. The method according to claim 6, characterized in that The determining the first light source information corresponding to the sample image includes: acquiring, based on the image type of the sample image, a first light source parameter of a sample light source corresponding to the sample image; Determining first description information of a sample light source and a first light source image corresponding to the sample image; The first light source parameters, the first description information, and the first light source image are determined as the first light source information.
8. The method according to claim 7, characterized in that The acquiring, based on the image type of the sample image, a first light source parameter of a sample light source corresponding to the sample image includes: When the image type is a composite type, determining the first light source parameter from composite information of the sample image; When the image type is a real type, light source analysis is performed on the sample image to obtain the first light source parameter.
9. The method according to claim 7, characterized in that The determining the first description information of the sample light source and the first light source image corresponding to the sample image includes: Acquire text information for describing the sample light source, and determine the text information as the first description information; A simulation image for simulating the sample light source is acquired, and the simulation image is determined as the first light source image.
10. An image generation method, characterized in that: The method comprises: When the image generation timing is reached, determining an image to be processed including the target object and second light source information corresponding to the image to be processed, where the second light source information includes at least one of a second light source parameter of the target light source, second description information, and a second light source image; determining, based on the image to be processed and the second light source information, a target noise of the image to be processed under the target light source using a trained image generation model, wherein the trained image generation model is obtained by the method according to any one of claims 1 to 9; The target noise is superimposed on the image to be processed through the trained image generation model to obtain a target image of the target object under the target light source.
11. A model training device, characterized in that: The device comprises: an acquisition module, configured to acquire a plurality of sample data pairs, wherein the sample data pairs include first sample data of a target object under a first light source, and second sample data of the target object under a second light source; a prediction module, configured to predict, for each pair of sample data, a first noise of the target object under the first light source and a second noise of the target object under the second light source by using a to-be-trained image generation model; a first determining module, configured to determine, based on the first noise and the second noise, a fusion noise of the target object under the first light source and the second light source by using the image generation model to be trained; The first determining module is further configured to determine a reference noise of the target object under the first light source and the second light source, and determine a loss value corresponding to the sample data pair based on the fusion noise and the reference noise; A training module is used to train the image generation model to be trained based on the loss value corresponding to each pair of sample data.
12. An image generating device, characterized in that: The device comprises: a second determining module, configured to determine, when an image generation timing is reached, an image to be processed including the target object and second light source information corresponding to the image to be processed, wherein the second light source information includes at least one of a second light source parameter of the target light source, second description information, and a second light source image; The second determination module is configured to determine the target noise of the image to be processed under the target light source by using a trained image generation model based on the image to be processed and the second light source information, wherein the trained image generation model is obtained by the method according to any one of claims 1 to 9; The superposition module is used to superimpose the target noise onto the image to be processed through the trained image generation model to obtain a generated image of the target object under the target light source.
13. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is used to implement the model training method described in any one of claims 1 to 9, or the image generation method described in claim 10 when executing the computer-executable instructions or computer program stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the model training method according to any one of claims 1 to 9 or the image generation method according to claim 10 is implemented.
15. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the model training method according to any one of claims 1 to 9 or the image generation method according to claim 10 is implemented.