Image generation method, apparatus, readable storage medium and program product

CN122597562APending Publication Date: 2026-08-18XIAMEN MEITUZHIJIA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610672449.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

理想的人像拍摄需要专业的布光,对人脸施加适量的光照以达到较好的拍照效果,而大多数实际拍摄场景下很难达到理想的布光环境,因此实际拍摄得到的图像中的人像通常存在面部阴暗、缺乏立体感、面部光线不均匀等问题

Benefits of technology

[0010]The aforementioned image generation method, apparatus, computer device, computer-readable storage medium, and computer program product acquire an image to be processed and a pair of reference images, wherein the reference image pair is used to reflect the differences in luminous efficacy between reference images under different lighting conditions, and encodes the reference image pair to obtain an encoded image pair; further, based on the luminous efficacy comparison information contained in the encoded image pair, the luminous efficacy feature information of the encoded image pair is determined, and based on the image to be processed and the luminous efficacy feature information, a target image after lighting processing is generated. Since the reference image pairs in this application are image pairs used to reflect the differences in lighting effects between reference images under different lighting conditions, the reference image pairs in this application essentially contain lighting effect contrast information. This allows for encoding of the reference image pairs to obtain encoded image pairs. The lighting effect feature information of the encoded image pairs, determined based on the lighting effect contrast information contained within them, is more accurate, significantly improving the accuracy of lighting effect decoupling. Consequently, the lighting effect in the target image generated after lighting processing, based on the image to be processed and more accurate lighting effect feature information, is also more accurate. This effectively distinguishes lighting effects from inherent object attributes (such as texture, color, and geometry), directly avoiding the problems of insufficient lighting effect expression and chaotic generation results caused by insufficient, ambiguous, or conflicting conditional information in diffusion model schemes based on text, normal maps, or single intensity values. In other words, the method provided in this application ensures that the transferred lighting effect is more faithful to the user's intent, reducing unexpected artifacts or style changes introduced by inaccurate lighting decoupling or conditional conflicts. This makes the lighting effect more controllable and predictable, improving the stability and reliability of editing. This technology achieves the technical effect of accurately transferring the lighting effects specified by the user while effectively improving the visual effect of the generated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597562A_ABST
    Figure CN122597562A_ABST
Patent Text Reader

Abstract

The application relates to an image generation method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a to-be-processed image and a reference image pair, the reference image pair is used for reflecting the light effect difference between reference images under different light conditions; encoding the reference image pair to obtain an encoded image pair; determining light effect feature information of the encoded image pair based on light effect contrast information contained in the encoded image pair; and generating a target image after light processing based on the to-be-processed image and the light effect feature information. The method can meet the light effect requirement, and effectively improve the quality and visual effect of the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image generation method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology

[0002] With the development of computer and internet technologies, users' demand for images is increasing, and these images are widely used in scenarios such as job hunting, exam registration, school enrollment, portrait photography, and food photography. When mobile users capture portraits and food in common everyday scenes, ambient light has the most direct impact on the appearance of the subject. It's difficult to obtain high-quality and atmospheric images in low-light environments. To improve the overall quality and atmosphere of an image, users usually need to prepare professional lighting equipment or use post-processing methods such as image editing apps. For example, more and more users are frequently using mobile devices to take portrait photos and then beautify the images to achieve their desired results. Ideal portrait photography requires professional lighting to apply appropriate light to the face for better results. However, in most actual shooting scenarios, it's difficult to achieve ideal lighting conditions. Therefore, the portraits obtained in actual shooting often suffer from problems such as dark faces, lack of three-dimensionality, and uneven facial lighting.

[0003] However, current image generation methods have significant limitations in handling image lighting tasks. These limitations mainly stem from fragmented algorithm design, unstable lighting effect modeling mechanisms, and inherent defects in the generation model, resulting in poor image quality and visual effects after lighting processing. For example, the traditional physically based rendering (PBR) approach decomposes the face reflectivity, normal map, and other parameters required for lighting into physical rendering. Since physical rendering relies on the decomposition of relevant components, this method suffers from component prediction accuracy issues, often resulting in artifacts and unrealistic phenomena, leading to poor lighting effects in the final generated image. Therefore, effectively improving the visual effects of generated images has become an urgent problem to be solved. Summary of the Invention

[0004] Based on this, this application provides an image generation method, apparatus, computer device, computer-readable storage medium, and computer program product. While satisfying the lighting effect, it also effectively improves the visual effect of the generated image. Specifically, it seamlessly integrates light effect feature information with the existing content of the image to be processed, generating a lit image with smooth light and shadow transitions, well-preserved structure, and rich detail. This significantly enhances the visual quality and realism of the target image, making the lit image appear as if it were actually taken under new lighting conditions. This not only improves user satisfaction but also allows for better application in professional scenarios with high image quality requirements.

[0005] On one hand, this application provides an image generation method, comprising: acquiring an image to be processed and a pair of reference images, wherein the pair of reference images is used to reflect the difference in light effect between reference images under different lighting conditions; encoding the pair of reference images to obtain an encoded pair of images; determining the light effect feature information of the encoded pair of images based on the light effect comparison information contained in the encoded pair of images; and generating a target image after lighting processing based on the image to be processed and the light effect feature information.

[0006] On one hand, this application also provides an image generation apparatus, comprising: an acquisition module for acquiring an image to be processed and a pair of reference images, the pair of reference images being used to reflect the differences in luminous efficacy between reference images under different lighting conditions; an encoding module for encoding the pair of reference images to obtain an encoded pair of images; a determination module for determining luminous efficacy feature information of the encoded pair of images based on luminous efficacy comparison information contained in the encoded pair of images; and a generation module for generating a target image after lighting processing based on the image to be processed and the luminous efficacy feature information.

[0007] On one hand, this application also provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: acquiring an image to be processed and a pair of reference images, the pair of reference images being used to reflect the differences in luminous efficacy between reference images under different lighting conditions; encoding the pair of reference images to obtain an encoded pair of images; determining luminous efficacy feature information of the encoded pair based on the luminous efficacy comparison information contained in the encoded pair of images; and generating a target image after lighting processing based on the image to be processed and the luminous efficacy feature information.

[0008] On the one hand, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, performs the following steps: acquiring an image to be processed and a pair of reference images, the pair of reference images being used to reflect the differences in luminous efficacy between reference images under different lighting conditions; encoding the pair of reference images to obtain an encoded pair of images; determining luminous efficacy feature information of the encoded pair based on the luminous efficacy comparison information contained in the encoded pair of images; and generating a target image after lighting processing based on the image to be processed and the luminous efficacy feature information.

[0009] On the one hand, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps: acquiring an image to be processed and a pair of reference images, the pair of reference images being used to reflect the differences in luminous efficacy between reference images under different lighting conditions; encoding the pair of reference images to obtain an encoded pair of images; determining luminous efficacy feature information of the encoded pair based on the luminous efficacy comparison information contained in the encoded pair of images; and generating a target image after lighting processing based on the image to be processed and the luminous efficacy feature information.

[0010] The aforementioned image generation method, apparatus, computer device, computer-readable storage medium, and computer program product acquire an image to be processed and a pair of reference images, wherein the reference image pair is used to reflect the differences in luminous efficacy between reference images under different lighting conditions, and encodes the reference image pair to obtain an encoded image pair; further, based on the luminous efficacy comparison information contained in the encoded image pair, the luminous efficacy feature information of the encoded image pair is determined, and based on the image to be processed and the luminous efficacy feature information, a target image after lighting processing is generated. Since the reference image pairs in this application are image pairs used to reflect the differences in lighting effects between reference images under different lighting conditions, the reference image pairs in this application essentially contain lighting effect contrast information. This allows for encoding of the reference image pairs to obtain encoded image pairs. The lighting effect feature information of the encoded image pairs, determined based on the lighting effect contrast information contained within them, is more accurate, significantly improving the accuracy of lighting effect decoupling. Consequently, the lighting effect in the target image generated after lighting processing, based on the image to be processed and more accurate lighting effect feature information, is also more accurate. This effectively distinguishes lighting effects from inherent object attributes (such as texture, color, and geometry), directly avoiding the problems of insufficient lighting effect expression and chaotic generation results caused by insufficient, ambiguous, or conflicting conditional information in diffusion model schemes based on text, normal maps, or single intensity values. In other words, the method provided in this application ensures that the transferred lighting effect is more faithful to the user's intent, reducing unexpected artifacts or style changes introduced by inaccurate lighting decoupling or conditional conflicts. This makes the lighting effect more controllable and predictable, improving the stability and reliability of editing. This technology achieves the technical effect of accurately transferring the lighting effects specified by the user while effectively improving the visual effect of the generated image. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an application environment diagram of the image generation method in one embodiment;

[0013] Figure 2 This is a flowchart illustrating an image generation method in one embodiment;

[0014] Figure 3 This is a schematic diagram of the overall process of an image generation method provided in one embodiment;

[0015] Figure 4This is a schematic diagram of a controlled generation diffusion model framework provided in one embodiment;

[0016] Figure 5 This is a structural block diagram of an image generation device in one embodiment;

[0017] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0018] To make the objectives, technical solutions, and beneficial effects of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] The image generation method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Server 104 can be a backend server for the image application. After terminal 102 acquires a to-be-processed image containing the user's portrait and a reference image pair, the terminal can send the acquired to-be-processed image and reference image pair to the backend server of the image application, i.e., server 104, so that server 104 encodes the reference image pair to obtain an encoded image pair. Based on the light effect contrast information contained in the encoded image pair, server 104 determines the light effect feature information of the encoded image pair. Furthermore, server 104 can generate a target image after lighting processing based on the to-be-processed image and the light effect feature information. Server 104 can return the generated target image after lighting processing to terminal 102 so that terminal 102 can display the target image after lighting processing.

[0020] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0021] In one exemplary embodiment, such as Figure 2As shown, an image generation method is provided, which can be applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps 202 to 208. Wherein:

[0022] Step 202: Obtain the image to be processed and the reference image pair. The reference image pair is used to reflect the difference in light effect between images under different lighting conditions.

[0023] The image to be processed refers to the image to be illuminated, that is, the image to be processed in this application is an image to be illuminated. For example, the image to be processed in this application can be an image containing a user's portrait or face. It is understood that the image to be processed in this application can be a real-time image captured by the user, or a historical image stored in a database. In some cases, for example, the image to be processed in this application can be a portrait image to be processed, that is, the image to be processed in this application needs to contain the user's face. It can be an image containing only the face area, or an image containing the user's half-body or full-body. For example, such as... Figure 3 The diagram shown is a schematic representation of the overall process of the image generation method provided in this application. The image to be processed can be an image input by the operator (the user currently using the terminal device). Figure 3 The target lighting image img_b shown can be a user image input by the user (i.e., an image containing the user's face or portrait).

[0024] A user's portrait refers to an image that at least includes the user's facial region. For example, in this application, a user's portrait can be a portrait that only includes the facial region, or it can be a portrait that includes the user's half-body or full-body; there are no specific limitations here.

[0025] A reference image pair refers to a pair of reference images, that is, a pair (or a set) of reference images. It can be understood that the reference image pair in this application contains two images. That is, the reference image pair in this application is used to reflect the difference in light effect between the two images contained in the image pair. For example, the reference image pair in this application can be composed of, for example... Figure 3 The image shown consists of an image img_a for showing the effect before lighting and an image img_A for showing the effect after lighting.

[0026] The difference in lighting effect refers to the difference in lighting effect between two images. For example, if a reference image pair contains images img_a and img_A, where image img_a is the original lighting effect display image (such as the lighting effect display image under natural light) and image img_A is the target lighting effect display image (such as the lighting effect display image using Rembrandt lighting), then the difference in lighting effect between image img_a and image img_A is the difference in lighting effect in this application.

[0027] Step 204: Encode the reference image pair to obtain the encoded image pair.

[0028] Here, encoding refers to encoding each of the two images contained in the reference image pair separately. For example, the terminal can use, for instance,... Figure 3 The visual encoder shown encodes the two images contained in the reference image pair separately to obtain the encoded image features of the two images contained in the reference image pair. The two encoded image features form the encoded image pair.

[0029] An encoded image pair refers to an image feature pair (i.e., two encoded image features) obtained by encoding a pair of reference images. For example, assuming that the reference image pair in this application includes a first reference image and a second reference image, after encoding the first reference image and the second reference image respectively, a first encoded feature of the first reference image and a second encoded feature of the second reference image are obtained. The first encoded feature and the second encoded feature are then combined to form the encoded image pair in this application.

[0030] Optionally / exemplarily, devices used by different users (operation objects) can interact with the image application (or image generation system). When a user (operation object) wants to generate an image that meets personalized needs (such as an ID photo under specific lighting effects), the user can open the image application (APP) on the terminal by triggering an operation, and enter the image generation page of the image application by selecting an operation. That is, the user can log in to the image application (such as an image enhancement application) by triggering an operation. Furthermore, on the generation page of the image application displayed on the terminal, the user can trigger an information editing operation, such as triggering an information input operation or an information selection operation, so that the terminal responds to the above-mentioned information input operation or information selection operation triggered by the user and obtains the image to be processed and the reference image pair input by the user for generating an ID photo that meets the lighting requirements. Furthermore, in response to the image generation request triggered by the operation object on the generation page of the image application (such as triggering a request to generate an ID photo that meets the lighting requirements based on the input image to be processed), the terminal automatically encodes the reference image pair to obtain an encoded image pair, for example, such as Figure 3 The processing flow shown allows the terminal to call a pre-trained visual encoder SigLip and encode the reference images contained in the reference image pair respectively through the visual encoder SigLip, thereby obtaining the light effect features of each reference image and combining the light effect features of different reference images into an encoded image pair.

[0031] It is understood that the method provided in this application can be implemented through interaction between the terminal and the backend server of the image application, or through interaction between the frontend and backend of the terminal. That is, the frontend of the terminal is used to display the image to be processed, the reference image pair, and the generated target image after lighting processing, while the backend of the terminal is equivalent to the backend server, used to encode the reference image pair and obtain the encoded image pair, etc.

[0032] Specifically, users can input an image containing their own image through interactive interface operations, causing the terminal to respond to the user's input and retrieve the corresponding image. That is, for example... Figure 3 As shown, after the terminal acquires the target lighting image img_b from the image to be processed, the user can also select a lighting effect that meets their personalized needs through interface interaction. The terminal then responds to the user-triggered lighting effect selection operation (e.g., selecting "Rembrandt lighting effect") by acquiring a lighting effect display image corresponding to the user-triggered operation "Rembrandt lighting effect." This display image includes a first lighting effect display image showing the original lighting effect and a second lighting effect display image showing the effect after processing with the target lighting effect, i.e., after processing with "Rembrandt lighting effect." The first and second lighting effect display images are combined into a reference image pair. Further, as... Figure 3 As shown, the terminal can invoke the pre-trained visual encoder SigLip and input the first and second light effect display images contained in the reference image pair in parallel, as shown in the figure. Figure 3 The visual encoder SigLip shown in the figure, after passing through... Figure 3 The visual encoder SigLip shown outputs encoded features containing light effect contrast information of the reference image pair after processing. These encoded features are the encoded image pair.

[0033] In an exemplary embodiment, obtaining the image to be processed and the reference image pair includes: the terminal responding to a triggered image selection operation to obtain the image to be processed corresponding to the image selection operation, wherein the image to be processed is the image to be illuminated; the terminal responding to a triggered light effect selection operation to obtain the light effect display image corresponding to the light effect selection operation; wherein the light effect display image includes a first light effect display image for displaying the original light effect and a second light effect display image for displaying the light effect after the target light effect processing, and the first light effect display image and the second light effect display image are combined into a reference image pair.

[0034] In an exemplary embodiment, the reference image pair includes a first reference image and a second reference image; the first reference image is an image used to display the original lighting effect, and the second reference image is an image used to display the image after processing for the target lighting effect; the method further includes: the terminal preprocesses the first reference image and the second reference image to obtain preprocessed first reference image and second reference image; encoding the reference image pair to obtain an encoded image pair, including: the terminal can perform feature extraction on the preprocessed first reference image and the second reference image through a visual encoder to obtain a first lighting effect feature of the first reference image and a second lighting effect feature of the second reference image; wherein, the first lighting effect feature and the second lighting effect feature are an encoded image pair.

[0035] In an exemplary embodiment, preprocessing the first reference image and the second reference image to obtain preprocessed first reference image and second reference image includes: the terminal adjusting the size of the first reference image and the second reference image to obtain adjusted first reference image and second reference image; wherein the adjusted first reference image and the second reference image meet the input size requirements of the visual encoder, and the adjusted first reference image and the second reference image are segmented to obtain first reference image block and second reference image block.

[0036] For example, let's take a portrait image as the image to be processed. Assume the image containing the user's portrait, input by the user, is "a real-time upper-body image captured by user A," and the user selects "Rembrandt lighting effect" to match their personalized lighting needs through interface interaction. Then, in response to the user's selection of the lighting effect (i.e., clicking "Rembrandt lighting effect"), the terminal retrieves the corresponding lighting effect display image, including, for example... Figure 3 The first light effect display image img_a and the second light effect display image img_A are shown in the figure, and the first light effect display image img_a and the second light effect display image img_A are combined into a reference image pair.

[0037] The first light effect display image, img_a, is used to display the original light effect (such as the lighting effect under natural light), while the second light effect display image, img_A, is used to display the light effect after the target light effect processing, that is, the "Rembrandt lighting effect" processing (i.e., the "Rembrandt lighting effect").

[0038] Furthermore, such as Figure 3 As shown, the terminal can invoke the pre-trained visual encoder SigLip and input the first light effect display image img_a and the second light effect display image img_A contained in the reference image pair in parallel as shown. Figure 3The visual encoder SigLip shown in the figure, after passing through... Figure 3 The visual encoder SigLip shown outputs a first encoded feature vector of the first light effect display image img_a and a second encoded feature vector of the second light effect display image img_A. The first encoded feature vector and the second encoded feature vector are the encoded image pair.

[0039] Step 206: Determine the light effect feature information of the encoded image pair based on the light effect contrast information contained in the encoded image pair.

[0040] The light effect contrast information refers to the light effect difference information contained in the encoded reference image pair. For example, the reference image pair contains images img_a and img_A, where image img_a is the original light effect display image (such as the light effect display image under natural light), and image img_A is the target light effect display image (such as the light effect display image using Rembrandt lighting). The difference in lighting effect between image img_a and image img_A is the light effect contrast information in this application.

[0041] Light effect feature information is used to characterize the characteristics of the lighting environment contained in the reference image pair. For example, the light effect feature information of the encoded image pair in this application is a feature used to characterize the differences in the lighting environment contained in the reference image pair, that is, a feature used to characterize the differences in the lighting effects contained in the reference image pair.

[0042] Step 208: Based on the image to be processed and the light effect feature information, generate the target image after lighting processing.

[0043] The target image refers to the generated image after lighting processing. In other words, the image to be processed in this application can be understood as the image to be lit. The target image is the image after lighting processing with a specific light effect selected by the user.

[0044] Specifically, the terminal acquires the image to be processed and a reference image pair, and encodes the reference image pair to obtain the encoded image pair. After obtaining the encoded image pair, the terminal can determine the light effect feature information of the encoded image pair based on the light effect contrast information contained in the encoded image pair. For example, the terminal can call a pre-trained lightweight light effect adapter LightAdapter and perform light effect feature decoupling processing on the encoded image pair through the lightweight light effect adapter LightAdapter to obtain the light effect feature information of the encoded image pair.

[0045] In an exemplary embodiment, the reference image pair includes a first reference image and a second reference image; the first reference image is used to display the original lighting effect, and the second reference image is used to display the image after processing for the target lighting effect; based on the lighting effect comparison information contained in the coded image pair, the lighting effect feature information of the coded image pair is determined, including: the terminal concatenates the first lighting effect feature of the first reference image with the second lighting effect feature of the second reference image to obtain a combined feature after concatenation; the combined feature is the lighting effect feature information of the coded image pair; or, the terminal concatenates the first lighting effect feature of the first reference image with the second lighting effect feature of the second reference image to obtain a combined feature after concatenation, and performs lighting effect decoupling processing on the combined feature to obtain a lighting effect condition vector reflecting the lighting effect change between the first lighting effect feature and the second lighting effect feature; the lighting effect condition vector is the lighting effect feature information of the coded image pair.

[0046] In an exemplary embodiment, the combined feature is an image sequence feature; the combined feature is subjected to light effect decoupling processing to obtain a light effect condition vector reflecting the light effect change between the first light effect feature and the second light effect feature, including: the terminal iteratively processes the image sequence feature through a light effect adapter to obtain a light effect condition vector reflecting the light effect change between the first light effect feature and the second light effect feature; wherein, the dimension of the light effect condition vector is the same as the dimension of the first light effect feature and the second light effect feature.

[0047] In an exemplary embodiment, the image sequence features are iteratively processed by a light effect adapter to obtain a light effect condition vector reflecting the light effect change between a first light effect feature and a second light effect feature. This includes: during the iterative processing of the image sequence features by the terminal using the light effect adapter, dividing the image sequence features into a first light effect feature sequence and a second light effect feature sequence; performing self-attention interaction processing on the first light effect feature sequence to obtain a consistency feature of the first light effect feature sequence; performing cross-attention processing on the first light effect feature sequence and the second light effect feature sequence to obtain a difference feature reflecting the difference between the first light effect feature sequence and the second light effect feature sequence; performing a feedforward operation on the consistency feature and the difference feature to obtain a feedforward-processed target feature sequence; and performing pooling processing on the target feature sequence to obtain a target light effect feature of a preset dimension. The target light effect feature is the light effect condition vector.

[0048] In an exemplary embodiment, cross-attention processing is performed on a first light effect feature sequence and a second light effect feature sequence to obtain a difference feature reflecting the difference between the first light effect feature sequence and the second light effect feature sequence. This includes: the terminal performs cross-attention processing on the first light effect feature sequence as query information and the second light effect feature sequence as query information through a light effect adapter to obtain a difference feature reflecting the difference between the first light effect feature sequence and the second light effect feature sequence.

[0049] In an exemplary embodiment, a self-attention interaction processing is performed on the first light effect feature sequence to obtain a consistency feature of the first light effect feature sequence. This includes: the terminal performs self-attention processing on the first light effect feature sequence as query information and the first light effect feature sequence as query information through a light effect adapter to obtain a consistency feature that reflects the light effect feature of the first light effect feature sequence itself.

[0050] Furthermore, the terminal can invoke a pre-trained image generation model, namely a diffusion model, and process the image to be processed and the lighting effect feature information through the image generation model to obtain the target image after lighting processing. For example, the terminal can use the image to be processed as... Figure 3 The image generation model shown is the input parameter of the SD 2.1 diffusion generation architecture, which also uses light effect feature information as... Figure 3 The image generation model shown, namely the SD 2.1 diffusion generation architecture, has its conditional signal input to, as... Figure 3 The image generation model shown is within the cross-attention layer of the SD 2.1 diffusion generation architecture, used to guide the image generation process. That is, as... Figure 3 The image generation model shown, namely the SD 2.1 diffusion generation architecture, while progressively removing noise and reconstructing the latent representation of the image, internally adjusts the lighting effects of the generated image based on the injected "light effect feature information" as a lighting effect condition, making it approximate the lighting changes shown in the reference image pair. Finally, the latent representation generated through multiple denoising steps within the model is sent back to the VAE decoder for decoding, restoring it to a pixel-space image, thus obtaining the final lit-processed image img_B, i.e., the target image. Through this processing method provided in this application, the changes in lighting effects can be accurately captured using the lighting effect contrast information of the image pair, and the generation capability of the diffusion model can be used to naturally transfer this lighting effect to the target image, ultimately obtaining a high-resolution result image that incorporates the target lighting effect. In other words, in the technical solution of this application, the progressive denoising generation process of the diffusion model helps to smoothly and harmoniously integrate the decoupled lighting effect into the structure and texture of the target image, ultimately generating a target image with natural lighting transitions, good detail preservation, and a realistic overall visual effect.

[0051] In an exemplary embodiment, the image to be processed contains a user's portrait; generating a target image after lighting processing based on the image to be processed and light effect feature information includes: the terminal generating the target image after lighting processing based on the image to be processed and light effect feature information through a diffusion model; or, the terminal generating the target image after lighting processing based on the image to be processed, light effect feature information, and an image vector reflecting the user's identity features through a diffusion model; wherein, the light effect feature information is the control condition of the diffusion model, the image to be processed is the input parameter of the diffusion model, and the image vector is extracted from the image to be processed.

[0052] In an exemplary embodiment, a target image after lighting processing is generated based on the image to be processed and light effect feature information using a diffusion model. This includes: the terminal encoding the image to be processed using an encoder in the diffusion model to obtain latent space features of the image to be processed; adding noise to the latent space features to obtain noisy latent space features; using the light effect feature information and the noisy latent space features as input parameters, iteratively denoising them in the U-Net structure of the diffusion model to obtain denoised latent space features; and decoding the denoised latent space features using a decoder in the diffusion model to obtain a pixel space image; the pixel space image is the target image after lighting processing.

[0053] In one exemplary embodiment, the U-Net structure includes at least an input layer and a cross-attention layer. The light effect feature information and the noisy latent space features are used as input parameters and iteratively denoised within the U-Net structure of the diffusion model to obtain the denoised latent space features. This includes: the terminal using the light effect feature information and the noisy latent space features as input parameters of the input layer, processing them through the input layer and the cross-attention layer to obtain the denoised latent space features; or, using the noisy latent space features as input parameters of the input layer and the light effect feature information as input parameters of the cross-attention layer, processing them through the input layer and the cross-attention layer to obtain the denoised latent space features.

[0054] In one exemplary embodiment, the luminous effect feature information and the noisy latent space features are used as input parameters and input into the U-Net structure in the diffusion model for iterative denoising to obtain the denoised latent space features. This includes: during the process of the terminal iteratively denoising the luminous effect feature information and the noisy latent space features through the U-Net structure in the diffusion model, the intermediate features of the noisy latent space features are used as query information, and the luminous effect feature information is used as the query information for iterative denoising to obtain the denoised latent space features.

[0055] In this embodiment, by acquiring the image to be processed and a pair of reference images, the pair of reference images is used to reflect the difference in light effect between reference images under different lighting conditions, and the pair of reference images is encoded to obtain a pair of encoded images; further, based on the light effect comparison information contained in the pair of encoded images, the light effect feature information of the pair of encoded images is determined, and based on the image to be processed and the light effect feature information, the target image after lighting processing is generated. Since the reference image pairs in this application are image pairs used to reflect the differences in lighting effects between reference images under different lighting conditions, the reference image pairs in this application essentially contain lighting effect contrast information. This allows for encoding of the reference image pairs to obtain encoded image pairs. The lighting effect feature information of the encoded image pairs, determined based on the lighting effect contrast information contained within them, is more accurate, significantly improving the accuracy of lighting effect decoupling. Consequently, the lighting effect in the target image generated after lighting processing, based on the image to be processed and more accurate lighting effect feature information, is also more accurate. This effectively distinguishes lighting effects from inherent object attributes (such as texture, color, and geometry), directly avoiding the problems of insufficient lighting effect expression and chaotic generation results caused by insufficient, ambiguous, or conflicting conditional information in diffusion model schemes based on text, normal maps, or single intensity values. In other words, the method provided in this application ensures that the transferred lighting effect is more faithful to the user's intent, reducing unexpected artifacts or style changes introduced by inaccurate lighting decoupling or conditional conflicts. This makes the lighting effect more controllable and predictable, improving the stability and reliability of editing. This technology achieves the technical effect of accurately transferring the lighting effects specified by the user while effectively improving the visual effect of the generated image.

[0056] In an exemplary embodiment, the step of obtaining the image to be processed and the reference image pair includes:

[0057] In response to a triggered image selection operation, obtain the image to be processed corresponding to the image selection operation;

[0058] In response to the triggered light effect selection operation, obtain the light effect display image corresponding to the light effect selection operation; wherein, the light effect display image includes a first light effect display image for displaying the original light effect and a second light effect display image for displaying the light effect after the target light effect processing;

[0059] Combine the first and second light effect display images into a reference image pair.

[0060] The image to be processed is the image to be illuminated. In this application, the image to be processed can be a portrait image to be processed.

[0061] The first and second light effect display images in this application are only used to distinguish different light effect display images. For example, the first light effect display image is used to display the original light effect (such as the original light effect is the lighting effect of natural light), and the second light effect display image is used to display the light effect after the target light effect is processed (such as the light effect display image of "Rembrandt lighting").

[0062] Specifically, let's take a portrait image as an example to illustrate this. Figure 3 As shown, users can input an image containing their own image to be processed through interactive interface operations. The terminal then responds to the user's input by retrieving the corresponding image. For example, a user can select a historical image from the database as the image to be processed, or they can upload a real-time captured image to the database and select it. In this case, the terminal responds to the user's image selection operation by retrieving the corresponding image to be processed.

[0063] Furthermore, users can also select lighting effects that meet their personalized needs through interface interaction. The terminal responds to the user's selection of the lighting effect (e.g., selecting "Rembrandt lighting effect") by obtaining a lighting effect display image corresponding to the user's selection of "Rembrandt lighting effect." This lighting effect display image includes, for example,... Figure 3 The first lighting effect display image (img_a) and the second lighting effect display image (img_A) are shown in the diagram, and are combined into a reference image pair. The first lighting effect display image (img_a) is used to display the original lighting effect (such as the lighting effect under natural light), and the second lighting effect display image (img_A) is used to display the lighting effect after processing with the target lighting effect, i.e., after processing with "Rembrandt lighting effect" (i.e., "Rembrandt lighting effect"). Therefore, the solution provided in this application embodiment has significant advantages in terms of application scenarios and flexibility. Users only need to provide the image to be processed and a pair of reference images (before and after lighting) that can display the desired lighting effect changes to achieve a smooth transition of lighting effects. This approach does not rely on specific 3D scene reconstruction, complex physical parameter settings, or predefined simple light source types (as limited by traditional physical a priori methods or simple parameter decoupling methods). It can transfer any observed complex lighting effects, whether it is subtle changes in natural light or dramatic artistic lighting, which greatly expands the application scope and creative freedom of image lighting processing, thereby effectively enriching and expanding the application scenarios and flexibility of the image lighting task processing provided in this application.

[0064] In an exemplary embodiment, the reference image pair includes a first reference image and a second reference image; the first reference image is used to display the original lighting effect, and the second reference image is used to display the image after processing for the target lighting effect; the method further includes:

[0065] The first reference image and the second reference image are preprocessed to obtain the preprocessed first reference image and the second reference image.

[0066] Encode the reference image pair to obtain the encoded image pair, including:

[0067] The first and second reference images are preprocessed by a visual encoder to extract features, thereby obtaining the first light effect feature of the first reference image and the second light effect feature of the second reference image.

[0068] Among them, the first light effect feature and the second light effect feature are coded image pairs.

[0069] The preprocessing in this application may include adjusting the size of the reference image, segmenting image blocks, etc., and no specific limitations are imposed here.

[0070] Encoded image pairs refer to the image features after encoding, that is, a pair of encoded image feature vectors.

[0071] A visual encoder is used to extract features that characterize the lighting environment of each image. The visual encoder in this application can be as follows: Figure 3 The SigLip visual encoder shown is used to extract features from each reference image in a pair of reference images to obtain image feature vectors that characterize the lighting environment (lighting effects) contained in different reference images.

[0072] Specifically, let's take a portrait image as an example to illustrate this. Figure 3As shown, after the terminal acquires the image to be processed and the reference image pair containing the user's portrait, the terminal can preprocess the first and second reference images contained in the reference image pair to obtain preprocessed first and second reference images. For example, the terminal can adjust the size of the first and second reference images respectively to obtain adjusted first and second reference images, and then segment the adjusted first and second reference images to obtain first and second reference image blocks; wherein, the adjusted first and second reference images both meet the input size requirements of the visual encoder. In this scheme, the input consists of two images in the reference image pair: the image before lighting (img_a) and the image after lighting (img_A). These images need to be preprocessed before being input into the model, such as resizing them to the size accepted by the model (e.g., 384x384 pixels) and segmenting them into image patches to adapt to its underlying Vision Transformer (ViT) architecture.

[0073] Furthermore, the terminal can input the preprocessed first reference image and the second reference image in parallel, such as... Figure 3 The SigLip visual encoder shown here, through, as Figure 3 The SigLip visual encoder shown extracts features from the preprocessed first and second reference images to obtain first lighting effect features of the first reference image and second lighting effect features of the second reference image. These first and second lighting effect features are then combined into an encoded image pair. For example, the terminal can use... Figure 3 The SigLip visual encoder shown extracts features from the preprocessed first and second reference images. The resulting first and second lighting effect features of the first and second reference images are both 768-dimensional feature vectors. These high-dimensional features aim to capture the rich semantic information and visual content of the input images. In other words, the output of the SigLip visual encoder is image feature vectors, also known as image embeddings.

[0074] In this embodiment, by directly extracting the differences in lighting effects from visual paradigms (image pairs), the inherent problem of "insufficient ability to express lighting effects" in other control methods is effectively avoided. This allows the subsequent generative model to more robustly capture and reproduce the target lighting effect, while also mitigating the risks of lighting decoupling difficulties and loss of details in the two-stage models used in traditional methods. In other words, the paired visual inputs provided by this application offer the model crucial comparative information, enabling it to more accurately identify and separate the target lighting effect (e.g., changes in light source direction, intensity, color, and shadows) purely caused by lighting variations by analyzing the differences in lighting effects between two images. This effectively eliminates interference information such as high-frequency textures, material properties, or other static ambient light inherent in the image content itself, thereby improving the accuracy of lighting effect transfer.

[0075] In one exemplary embodiment, the reference image pair includes a first reference image and a second reference image; the first reference image is an image used to display the original lighting effect, and the second reference image is an image used to display the image after processing for the target lighting effect; the step of determining the lighting effect feature information of the encoded image pair based on the lighting effect contrast information contained in the encoded image pair includes:

[0076] The first light effect feature of the first reference image and the second light effect feature of the second reference image are concatenated to obtain the combined feature after concatenation; the combined feature is the light effect feature information of the encoded image pair; or...

[0077] The first light effect feature of the first reference image and the second light effect feature of the second reference image are concatenated to obtain the concatenated combined feature; the combined feature is decoupled by light effect to obtain the light effect condition vector that reflects the light effect change between the first light effect feature and the second light effect feature; the light effect condition vector is the light effect feature information of the encoded image pair.

[0078] In this application, the light effect decoupling process can be implemented using a pre-trained light effect adapter. The purpose is to further process the combined features after splicing, aiming to more accurately decouple and extract pure "light effect condition" information that can characterize the illumination change from the first reference image img_a to the second reference image img_A. For example, the light effect adapter used in this application can be as follows: Figure 3 The LightAdapter module shown is a lightweight light effect adapter.

[0079] Specifically, let's take a portrait image as an example to illustrate this. Figure 3As shown, the terminal acquires a pair of images to be processed containing the user's image and a pair of reference images, and encodes the pair of reference images to obtain an encoded image pair. The encoded image pair contains a first light effect feature of the first reference image and a second light effect feature of the second reference image. The terminal can then splice the first light effect feature of the first reference image with the second light effect feature of the second reference image to obtain a combined feature after splicing. The combined feature is the light effect feature information of the extracted encoded image pair.

[0080] Alternatively, the terminal can concatenate the first lighting effect feature of the first reference image with the second lighting effect feature of the second reference image to obtain a concatenated combined feature. A lightweight lighting effect adapter is then used to decouple the combined feature (feature extraction) to obtain a lighting effect condition vector reflecting the lighting effect changes between the first and second lighting effect features. This lighting effect condition vector represents the extracted lighting effect feature information of the encoded image pair. This allows for the concatenation of lighting effect feature information extracted from the pre-lighting reference image (img_a) and the post-lighting reference image (img_A) to obtain a concatenated combined feature. Further in-depth processing and analysis of this concatenated combined feature achieves precise decoupling and generates a lighting effect condition vector. This provides more accurate lighting effect data for subsequently leveraging the superior image generation capabilities of the diffusion model to efficiently and realistically integrate the decoupled lighting effect into the target image. Compared to various lighting solutions provided in traditional methods, the method provided in this application demonstrates significant advantages in terms of flexibility, accuracy, and generation quality, bringing important application value. Specifically, the method provided in the embodiments of this application can more accurately decouple pure light effect information and eliminate interference from the inherent attributes of image content. It can effectively avoid the "texture-like" or unrealistic feeling caused by inaccurate geometry or material estimation in traditional physical rendering methods, and can also overcome the serious distortion caused by simple pixel adjustment methods, thereby improving the quality and visual effect of the generated image.

[0081] In one exemplary embodiment, the combined feature is an image sequence feature; the step of performing light effect decoupling processing on the combined feature to obtain a light effect condition vector reflecting the light effect change between the first light effect feature and the second light effect feature includes:

[0082] The image sequence features are iteratively processed by the light effect adapter to obtain a light effect condition vector that reflects the light effect change between the first light effect feature and the second light effect feature.

[0083] The dimension of the light effect condition vector is the same as the dimension of the first light effect feature and the second light effect feature.

[0084] Specifically, the terminal can be accessed through methods such as Figure 3The light effect adapter shown iteratively processes the combined features (i.e., image sequence features) after stitching. The output of the light effect adapter is a light effect condition vector that reflects the light effect change between the first and second light effect features. The dimension of the light effect condition vector is the same as the dimensions of the first and second light effect features. For example, the dimensions of the first and second light effect features and the light effect condition vector are all 768-dimensional, meaning that the first, second, and light effect features and the light effect condition vector are all 768-dimensional light effect feature vectors.

[0085] The core function of the LightAdapter in this application is to receive and process the combined feature information output from the SigLip visual encoder. Specifically, the LightAdapter receives the combined features (features) obtained by concatenating the lighting effect feature information extracted from the reference image before lighting (img_a) and the reference image after lighting (img_A). Through its internal structure (e.g., the designed five internal modules including self-attention and cross-attention), the LightAdapter performs in-depth processing and analysis on these combined features to achieve precise decoupling and generate lighting effect conditional vectors.

[0086] like Figure 3 As shown, the input to the lighting adapter in this application is the stitched embedding of the images before lighting (img_a) and after lighting (img_A) as input parameters. It can be understood that the use of stitched embedding (i.e., image patch embedding sequence) in this embodiment preserves more spatial information compared to using only a single global embedding vector for each image, thus providing richer spatial information.

[0087] The internal iterative refinement (module) process in the light effect adapter of this application includes: the core of the light effect adapter consists of 5 stacked modules, each of which performs self-attention, cross-attention and feedforward operations respectively. This iterative structure allows the model to gradually deepen its understanding of illumination changes.

[0088] The self-attention processing for H_A includes: within each module, the self-attention mechanism allows different parts (image patches) of the "lit" image representation (H_A) to interact with each other. Benefits: This helps the model integrate features related to the consistent effects of new lighting in the image. It can recognize patterns and structures within the target lighting state itself, potentially highlighting the effects of a globally applied effect or consistent shadow orientation.

[0089] The cross-attention (H_A query H_a) process involves a crucial step in extracting differences introduced by lighting. The current state (H_A) of the "lit" image is used as a query, while the complete representation (H_a) of the "unlit" image, providing both keys and values, is attended to. Benefits: This step explicitly forces the model to compare the "lit" state with the "unlit" state. By attending to H_a, features in H_A can identify which aspects have changed due to lighting and which are part of the underlying scene content (which should exist in both H_a and H_A). This mechanism directly facilitates the decoupling of lighting effects from the image content.

[0090] In this embodiment, by comparing images before and after lighting, the model can focus more on the changes brought about by the lighting itself, effectively distinguishing the lighting effect from the inherent attributes of the object (such as texture, color, and geometry). This overcomes the difficulty in completely separating lighting information from content or style in traditional methods that rely on a single reference image (including some two-stage deep learning methods). Simultaneously, the method provided in this application directly avoids the problems of insufficient lighting effect expression and chaotic generation results caused by insufficient, ambiguous, or conflicting conditional information in diffusion models based on text, normal maps, or single intensity values. In other words, the image pairs in this application provide a direct, rich, and unambiguous visual paradigm for defining lighting effect transformation, offering more precise and reliable control compared to indirect textual or geometric constraints. Its value lies in ensuring that the transferred lighting effect is more faithful to the user's intent, reducing unexpected artifacts or style changes introduced by inaccurate lighting decoupling or conditional conflicts. This makes the lighting effect more controllable and predictable, improving the stability and reliability of editing. This technology achieves the technical effect of accurately transferring the lighting effects specified by the user while effectively improving the visual effect of the generated image.

[0091] In one exemplary embodiment, the step of iteratively processing the image sequence features using a light effect adapter to obtain a light effect condition vector reflecting the light effect change between the first light effect feature and the second light effect feature includes:

[0092] During the iterative processing of the image sequence features through the light effect adapter, the image sequence features are divided into a first light effect feature sequence and a second light effect feature sequence.

[0093] The first light effect feature sequence is subjected to self-attention interaction processing to obtain the consistency features of the first light effect feature sequence;

[0094] Cross-attention processing is performed on the first and second light effect feature sequences to obtain difference features that reflect the differences between the first and second light effect feature sequences;

[0095] Feature transformation is performed on consistent and dissimilar features to obtain the target feature sequence;

[0096] The target feature sequence is pooled to obtain target light effect features of a preset dimension; the target light effect features are light effect condition vectors.

[0097] The first and second light effect feature sequences refer to the light effect feature sequences used to distinguish between different light effect features. For example, the first light effect feature sequence in this application can be represented as H_A, where H_A represents the image "after lighting". H_A can be understood as the sequence to be processed within the model. The second light effect feature sequence in this application can be represented as H_a, where H_a represents the image "before lighting". H_a can be understood as the context sequence within the model, used to provide the "before lighting" state in cross-attention processing.

[0098] Feature transformation can be achieved through the feedforward network module in the LightAdapter, for example, the feedforward network module includes a standard MLP module for further feature transformation.

[0099] Specifically, the terminal, through, such as Figure 3 The LightAdapter shown iteratively processes the combined features, i.e., the image sequence features. The LightAdapter automatically divides the image sequence features into a first light effect feature sequence H_A and a second light effect feature sequence H_a. For example, if the combined features input to the LightAdapter are: Z_concat = concat(embedding_a, embedding_A) # Shape (2N, D), the LightAdapter can initially project the combined features Z_concat onto the working dimension (e.g., if D is different from 768). If D != 768: proj_a = Linear(embedding_a) # Shape (N, 768), proj_A = Linear(embedding_A) # Shape (N, 768). Otherwise: proj_a = embedding_a, proj_A = embedding_A, and let H_a = proj_a, H_A = proj_A. Here, H_a is the context (Key, Value), and H_A is the sequence to be processed (Query).

[0100] Furthermore, the terminal performs self-attention interaction processing on the first light effect feature sequence H_A through the self-attention module inside the LightAdapter to obtain the consistency features of the first light effect feature sequence H_A. For example, in this application, the self-attention module inside the LightAdapter acts on H_A, allowing the tokens inside the "lit" image representation to interact with each other to extract features. The specific processing method inside the self-attention module can be as follows:

[0101] attn_output_self = MultiHeadSelfAttention(H_A)

[0102] H_A = LayerNorm(H_A + attn_output_self) Residual connection or layer normalization.

[0103] In one exemplary embodiment, the step of performing cross-attention processing on the first luminous effect feature sequence and the second luminous effect feature sequence to obtain difference features reflecting the difference between the first luminous effect feature sequence and the second luminous effect feature sequence includes:

[0104] Cross-attention processing is performed on the first light effect feature sequence as the query information and the second light effect feature sequence as the query information to obtain the difference features that reflect the difference between the first light effect feature sequence and the second light effect feature sequence.

[0105] Here, the query information refers to the query vector Q (Query). The query vector represents the target sequence for which information needs to be generated or updated (such as the current state of the decoder), and is used to "ask": "Which parts of the source sequence do I need to focus on?"

[0106] The information to be queried refers to the key vector K (Key) and the value vector V (Value). The key vector K (Key) represents the index of the source sequence (such as encoder output), used to match Q and calculate relevance. The value vector V (Value) represents the actual content of the source sequence, ultimately generated as a contextual representation through weighted summation.

[0107] Specifically, the terminal performs cross-attention processing on the first light effect feature sequence H_A and the second light effect feature sequence H_a through the cross-attention module (H_A query H_a) inside the LightAdapter, obtaining difference features that reflect the differences between the first light effect feature sequence H_A and the second light effect feature sequence H_a. For example, in this application, the cross-attention module (H_A query H_a) inside the LightAdapter allows features "after lighting" to focus on features "before lighting" to discover differences / correspondences. That is, the terminal performs cross-attention processing on the first light effect feature sequence H_A as query information, i.e., query vector Q, and the second light effect feature sequence as query information, i.e., key vector K and value vector V, through the cross-attention module (H_A query H_a) inside the LightAdapter, obtaining difference features that reflect the differences between the first light effect feature sequence and the second light effect feature sequence.

[0108] Simultaneously, the terminal performs cross-attention processing on the first light effect feature sequence H_A and the second light effect feature sequence H_a through the cross-attention module (H_A queries H_a) within the LightAdapter, obtaining difference features that reflect the differences between the first light effect feature sequence H_A and the second light effect feature sequence H_a. For example, the cross-attention module (H_A queries H_a) within the LightAdapter in this application allows features "after lighting" to focus on features "before lighting" to discover differences / correspondences. The specific processing method within the cross-attention module can be as follows:

[0109] attn_output_cross=MultiHeadCrossAttention(query=H_A,key=H_a, value=H_a)

[0110] H_A = LayerNorm(H_A + attn_output_cross), residual connection or layer normalization.

[0111] Simultaneously, the terminal performs feature transformation on the consistent and dissimilar features through the feedforward network module inside the LightAdapter to obtain the target feature sequence. For example, in this application, the feedforward network module inside the LightAdapter operates on H_A, and its internal standard MLP module is used for further feature transformation. The specific processing method inside the feedforward network module can be as follows:

[0112] ff_output = FeedForward(H_A)

[0113] H_A = LayerNorm(H_A + ff_output), residual connection or layer normalization.

[0114] Furthermore, the terminal performs pooling processing on the target feature sequence through the LightAdapter to obtain target light effect features of a preset dimension. That is, the terminal pools the processed sequence H_A (i.e., the target sequence) through the LightAdapter to obtain a single vector representation pooled_output = AveragePooling(H_A, dimension=sequence_dim) of shape (768). Alternative solution: If a special [CLS] token is added in advance, its representation is used. For example, the final output is a 768-dimensional light effect feature vector, Output = pooled_output of shape (768), and this 768-dimensional light effect feature vector is the light effect condition vector.

[0115] Component details:

[0116] MultiHeadSelfAttention: Standard multi-head self-attention, where Q, K, and V all come from the input sequence H_A.

[0117] MultiHeadCrossAttention: Standard multi-head cross-attention, where Q comes from H_A, while K and V come from H_a.

[0118] FeedForward typically consists of two linear layers with a non-linear activation function, such as GELU, sandwiched in between.

[0119] LayerNorm: Standard layer normalization.

[0120] AveragePooling: Performs average pooling on features along the sequence dimension (N).

[0121] In this embodiment, by comparing images before and after lighting, the model can focus more on the changes brought about by the lighting itself, effectively distinguishing the lighting effect from the inherent attributes of the object (such as texture, color, and geometry). This overcomes the difficulty in completely separating lighting information from content or style in traditional methods that rely on a single reference image (including some two-stage deep learning methods). Simultaneously, the method provided in this application directly avoids the problems of insufficient lighting effect expression and chaotic generation results caused by insufficient, ambiguous, or conflicting conditional information in diffusion models based on text, normal maps, or single intensity values. In other words, the image pairs in this application provide a direct, rich, and unambiguous visual paradigm for defining lighting effect transformation, offering more precise and reliable control compared to indirect textual or geometric constraints. Its value lies in ensuring that the transferred lighting effect is more faithful to the user's intent, reducing unexpected artifacts or style changes introduced by inaccurate lighting decoupling or conditional conflicts. This makes the lighting effect more controllable and predictable, improving the stability and reliability of editing. This technology achieves the technical effect of accurately transferring the lighting effects specified by the user while effectively improving the visual effect of the generated image.

[0122] In one exemplary embodiment, the step of performing cross-attention processing on the first luminous effect feature sequence and the second luminous effect feature sequence to obtain difference features reflecting the difference between the first luminous effect feature sequence and the second luminous effect feature sequence includes:

[0123] Cross-attention processing is performed on the first light effect feature sequence as the query information and the second light effect feature sequence as the query information to obtain the difference features that reflect the difference between the first light effect feature sequence and the second light effect feature sequence.

[0124] Here, the query information refers to the query vector Q (Query). The query vector represents the target sequence for which information needs to be generated or updated (such as the current state of the decoder), and is used to "ask": "Which parts of the source sequence do I need to focus on?"

[0125] The information to be queried refers to the key vector K (Key) and the value vector V (Value). The key vector K (Key) represents the index of the source sequence (such as encoder output), used to match Q and calculate relevance. The value vector V (Value) represents the actual content of the source sequence, ultimately generated as a contextual representation through weighted summation.

[0126] Specifically, the terminal performs cross-attention processing on the first light effect feature sequence H_A and the second light effect feature sequence H_a through the cross-attention module (H_A query H_a) within the LightAdapter, obtaining difference features that reflect the differences between the first light effect feature sequence H_A and the second light effect feature sequence H_a. For example, the cross-attention module (H_A query H_a) within the LightAdapter in this application allows features "after lighting" to focus on features "before lighting" to discover differences / correspondences. That is, the terminal performs cross-attention processing on the first light effect feature sequence H_A as query information (query vector Q), and the second light effect feature sequence as query information (key vector K and value vector V), obtaining difference features that reflect the differences between the first and second light effect feature sequences. This significantly improves the accuracy of light effect decoupling and the precision of control by using image pairs for guidance.

[0127] In one exemplary embodiment, the step of performing self-attention interaction processing on the first light effect feature sequence to obtain the consistency features of the first light effect feature sequence includes:

[0128] The first luminous effect feature sequence is used as the query information and the first luminous effect feature sequence is used as the query information to perform self-attention processing, so as to obtain a consistency feature that reflects the luminous effect feature of the first luminous effect feature sequence itself.

[0129] Specifically, the terminal performs self-attention interaction processing on the first light effect feature sequence H_A through the self-attention module inside the LightAdapter to obtain the consistency features of the first light effect feature sequence H_A. For example, in this application, the self-attention module inside the LightAdapter acts on H_A, allowing the tokens inside the "lit" image representation to interact with each other to extract features. That is, the terminal uses the self-attention module inside the LightAdapter to perform self-attention processing on the first light effect feature sequence H_A as query information (query vector Q), the first light effect feature sequence H_A as query information (key vector K), and the value vector V) to obtain the consistency features that reflect the light effect features of the first light effect feature sequence itself. This significantly improves the accuracy of light effect decoupling and the precision of control by using image pairs for guidance.

[0130] In an exemplary embodiment, the image to be processed contains a portrait of a user; the step of generating a target image after lighting processing based on the image to be processed and lighting effect feature information includes:

[0131] Using a diffusion model, a target image after lighting processing is generated based on the image to be processed and the light effect feature information; or...

[0132] Using a diffusion model, a target image after lighting processing is generated based on the image to be processed, light effect feature information, and image vectors that reflect the user's identity features.

[0133] Among them, the light effect feature information is the control condition of the diffusion model, the image to be processed is the input parameter of the diffusion model, and the image vector is extracted from the image to be processed. The image vector can also be used as the control condition of the diffusion model. For example, the terminal can extract the identity features of the image to be processed through the CLIP image encoder to obtain the image vector id_embed, which is used to represent the identity appearance features of the user, so as to maintain the texture and detail features of the image person (such as moles, scars, skin texture, hair strands and other detail features).

[0134] Specifically, let's take a portrait image as an example to illustrate this. Figure 4 The diagram shown illustrates the framework for the controlled generation diffusion model provided in this application. The terminal can invoke a pre-trained diffusion model and, as shown... Figure 4 The light effect feature information output by the LightAdapter shown is used as control information (light effect control conditions) and input into the U-Net of the diffusion model in parallel. At the same time, the image to be processed is used as input parameters, which are input into the encoder of the diffusion model for encoding and into the U-Net for weak texture reconstruction. Finally, the result image after lighting processing that meets the lighting effect requirements specified by the user and contains the user's portrait is generated, which is the target image.

[0135] Alternatively, the terminal can also... Figure 4The light effect feature information output by the LightAdapter shown is used as control information (light effect control conditions) and input into the U-Net of the diffusion model in parallel. At the same time, the extracted identity / appearance vector, i.e., the image vector, is injected into the cross-attention layer of the U-Net. The image to be processed is used as input parameter and input into the U-Net of the diffusion model for weak texture reconstruction. Finally, the result image after lighting processing that meets the lighting effect requirements specified by the user and contains the user's portrait is generated, i.e., the target image. The IP-Adapter (identity / appearance branch) can use the CLIP image encoder to extract identity features from the input image, obtaining the image vector id_embed. An "image key-value" path is added to the cross-attention module of the U-Net in the diffusion model, which is equivalent to the KV replacement / sponging of the IP-Adapter. A more optimized approach is to enable the image vector id_embed in the mid- and deep layers (such as block2–block3) to ensure that identity / texture influences more on high-frequency details rather than large geometry. As a result, by adopting the processing method of diffusion generation + light effect control + identity embedding, the diffusion model can perform weak texture reconstruction under more precise light effect control, and stabilize individual differences with identity embedding features (such as ArcFace / reference image adaptation). Compared with the traditional method, the "smoothing" of details and homogenization are significantly reduced, that is, user-specific identity details such as moles / scars / skin textures are better preserved, achieving stronger consistency of "looking like the person", improving the pass rate of ID card verification and reducing false rejections in online KYC.

[0136] In one exemplary embodiment, the step of generating a target image after lighting processing based on the image to be processed and light effect feature information using a diffusion model includes:

[0137] The latent space features of the image to be processed are obtained by encoding the image in the diffusion model using an encoder.

[0138] Noise is added to the latent space features to obtain the noisy latent space features;

[0139] The light effect feature information and the noisy latent space feature are used as input parameters and iteratively denoised in the U-Net structure of the diffusion model to obtain the denoised latent space feature.

[0140] The latent space features after denoising are decoded by the decoder in the diffusion model to obtain the pixel space image; the pixel space image is the target image after lighting processing.

[0141] Specifically, such as Figure 4The processing flow shown allows the terminal to encode the image to be processed using a variational autoencoder (VAE) in the diffusion model, obtaining the latent space features (shallow representation) of the image. This enables subsequent diffusion processes to be performed in the low-dimensional latent space. The latent space features are then denoised to obtain noisy latent space features. Furthermore, the terminal can use the light effect feature information and the noisy latent space features as input parameters, inputting them into the U-Net structure in the diffusion model for iterative denoising to obtain denoised latent space features. The denoised latent space features are then decoded by the decoder in the diffusion model to obtain a pixel space image. The pixel space image is the target image after lighting processing.

[0142] In one exemplary embodiment, the U-Net structure includes at least an input layer and a cross-attention layer; the step of iteratively denoising the U-Net structure in the diffusion model by taking the light effect feature information and the noisy latent space features as input parameters to obtain the denoised latent space features includes:

[0143] The terminal can use the light effect feature information and the noisy latent space features as input parameters of the input layer in the U-Net structure. After processing by the input layer, the cross-attention layer, and other layers, the denoised latent space features are obtained. Alternatively, the terminal can use the noisy latent space features as input parameters of the input layer in the U-Net structure and the light effect feature information as input parameters of the cross-attention layer in the U-Net structure (that is, inject the light effect feature information into the cross-attention layer of U-Net). After processing by the input layer, the cross-attention layer, and other layers, the denoised latent space features are obtained.

[0144] In one exemplary embodiment, the step of using light effect feature information and noisy latent space features as input parameters, and iteratively denoising the U-Net structure in the diffusion model to obtain the denoised latent space features includes:

[0145] In the process of iteratively denoising the light effect feature information and the noisy latent space features through the U-Net structure in the diffusion model, the intermediate features of the noisy latent space features are used as query information, i.e., query vector Q, and the light effect feature information is used as the query information, i.e. key vector K and value vector V, to perform iterative denoising and obtain the denoised latent space features.

[0146] The key control element in this embodiment is the injection of the "light effect embedding" (representing the target lighting effect) proposed by the light effect adapter into specific layers within U-Net (such as Cross-Attention Layers). U-Net uses its intermediate features as a query to "attend to" this "light effect embedding" (as a key and value). Through this mechanism, the lighting effect information guides the noise prediction direction of U-Net, causing it to evolve latent variables in a direction consistent with the target lighting effect while denoising. Finally, after a predetermined number of iterative denoising steps, U-Net outputs an approximately clean latent variable table z_0. This latent variable table z_0 not only reconstructs the content of the input image but also incorporates the target lighting effect. The clean latent variable representation z_0 output by U-Net serves as... Figure 4 The input parameters of the VAE decoder shown are as follows: the output of the decoder is the resulting image, which is the target image after applying the target lighting effect. The function of the VAE decoder in this application is to map latent variables from a low-dimensional latent space back to a high-dimensional pixel space, reconstructing the final user-visible image.

[0147] In this embodiment, image processing is transferred to the more computationally efficient latent space using a VAE encoder. The core U-Net diffusion model (similar to Stable Diffusion 2.1) is responsible for progressively denoising and generating images in the latent space. Unlike standard SD, its generation process is not guided by text embedding. Instead, a dedicated LightAdapter module extracts "light effect embedding" from reference image pairs and precisely controls it through a cross-attention mechanism. Finally, the VAE decoder obtains a high-resolution result image that incorporates the target light effect, significantly improving the visual quality and realism of the final output image. This makes the edited image look like it was actually taken under new lighting conditions, thereby improving the user experience and enabling the solution provided in this application to be better applied to professional scenarios with high image quality requirements.

[0148] This application also provides an application scenario in which the above-described image generation method is applied. The method provided in this application embodiment can be applied to various scenarios of personalized lighting image generation. The following uses a scenario where a user interacts with an image enhancement system as an example to illustrate the image generation method provided in this application embodiment.

[0149] Image lighting, also known as image relighting, is an important computer vision and image processing task. Its core objective is to modify the lighting effects in an image while preserving its original content (such as object identity, texture, and scene structure). This typically involves changing the position, intensity, color, number, or type of virtual light sources in the scene to simulate different lighting conditions, thereby enhancing the image's aesthetics, correcting lighting defects, or achieving a specific artistic style. For example, it can transform a daytime photograph into a nighttime effect, add dramatic Rembrandt lighting to a portrait, or eliminate facial shadows caused by backlighting. This technique usually takes a single 2D image as input and outputs an image with new lighting effects.

[0150] Image lighting tasks face numerous challenges. The most fundamental challenge lies in the fact that recovering 3D scene information (including geometry, material reflection properties, and the current lighting environment) from a single 2D image is a classic "ill-posed problem." This means that these intrinsic properties cannot be uniquely determined from a single image. Specific challenges include:

[0151] 1) Difficulty in geometry and material estimation: Accurately estimating the 3D shape of an object (e.g., through normal maps or depth maps) and surface materials (e.g., albedo, roughness) is crucial for physically based rendering, but accurately estimating this information from a single viewpoint image is very difficult, especially when dealing with complex textures, transparent or reflective objects. The accuracy of normal prediction directly affects lighting effects, and boundary areas are particularly prone to errors.

[0152] 2) Modeling complex lighting effects: Real-world lighting includes not only direct lighting but also complex global illumination effects, such as soft shadows, indirect lighting (light bounce), and subsurface scattering. These effects are difficult to model and render accurately.

[0153] 3) Difficulty in acquiring data: Supervised deep learning methods usually require a large amount of paired training data (images of the same scene under different lighting conditions), and acquiring this type of real-world data is costly and time-consuming.

[0154] 4) Realism and content consistency: The generated result needs to be visually realistic and conform to the laws of physics. At the same time, it must strictly maintain the content of the original image and avoid distorting the identity of objects or introducing irrelevant artifacts.

[0155] 5) Controllability and Automation: While some end-to-end deep learning methods may yield good results, they often lack intuitive control methods, making it difficult for users to precisely specify the desired lighting effects (such as light source position and intensity). Furthermore, fully automatic optimization of scene lighting without relying on user-specified guide lighting is also a challenge.

[0156] 6) Exposure issues: Increasing the brightness of the virtual light source or changing the direction of the light can easily lead to overexposure of the original highlight areas of the image, requiring an effective exposure control mechanism.

[0157] Portrait relighting is an important and specialized subtask in image lighting technology, focusing on modifying and optimizing the lighting effects on people, especially faces, in images. Its goal is to simulate different lighting conditions while preserving core information such as the subject's identity, facial features, and skin texture, in order to improve the aesthetic quality of portrait photos or correct lighting defects during shooting. For example, it can eliminate "half-lit faces" caused by backlighting or sidelighting, add soft butterfly lighting or artistic Rembrandt lighting to faces, simulate professional studio lighting effects, or simply provide natural fill light for portraits. This technology is particularly important for improving the shooting performance of mobile devices (such as smartphones) in complex lighting environments, aiming to provide users with a convenient, efficient, high-quality, and atmospheric out-of-camera shooting experience.

[0158] Portrait lighting presents far greater challenges than general scene lighting. First, the human face possesses an extremely complex 3D structure and detail, such as the fine shapes of the eyes, nose, and lips, as well as the subtle textures and pores of the skin, all of which are highly sensitive to lighting effects. Achieving realistic lighting requires a precise understanding and modeling of the face's geometry, which typically relies on accurate normal vector prediction. Incorrect geometric information (such as inaccurate normal map predictions or imprecise boundaries) can lead to uneven brightness or artifacts after lighting. Second, the optical properties of skin are highly complex, including diffuse reflection, specular reflection, and subsurface scattering. Accurately simulating the skin's appearance under different lighting conditions is crucial for realism. Furthermore, the complex structure of individual hair strands and accessories such as glasses and hats further complicate the modeling of light interactions. Finally, maintaining consistency in the subject's identity is paramount in portrait lighting; any lighting adjustments should not distort the subject's facial features, making them resemble someone else entirely.

[0159] Limitations of traditional solutions:

[0160] In the research and practice of image and portrait lighting, a variety of technical solutions have emerged, but each of them also faces significant bottlenecks.

[0161] The first type of approach is physics-guided lighting. This type of method attempts to simulate the physical processes of lighting in the real world. It typically requires reconstructing or estimating the geometry of objects in a virtual 3D scene (e.g., using normal maps or depth maps), then setting virtual light sources (usually simulated point lights or parallel light sources, etc.), and calculating the interaction between light and the object's surface based on optical models (such as Blinn-Phong, Lambertian, etc.), ultimately rendering an image with new lighting effects. However, this method has limitations: a) It is limited in the types of lighting effects, and can usually only simulate simple, parametric electric point light sources well, making it difficult to represent the rich light and shadow changes brought by natural light, ambient light, or complex regional light sources; b) It essentially adds new lighting effects to the original image (or the estimated surface reflectivity), so it cannot effectively remove the shadow areas that already exist in the original image unless combined with additional, often imperfect, shadow removal algorithms; c) Since accurately estimating the three-dimensional geometry and surface material (BRDF) from a single two-dimensional image is difficult in itself, estimation errors will cause the rendered lighting effects to blend unnaturally with the content of the original image, producing a noticeable "texture look" and lacking realism, especially when dealing with complex materials such as skin and hair.

[0162] The second type of approach attempts to decouple several control parameters from simple lighting conditions and directly adjust pixel values. This method tries to bypass complex 3D reconstruction and physically based rendering by analyzing the image (sometimes by comparing it with a standard lighting image) to extract simplified lighting descriptors, such as the overall light intensity coefficient, the direction of the main light source, and low-order spherical harmonics (SH) coefficients. Then, based on these parameters, the pixel values ​​of the corresponding areas in the original image are directly modified through a mapping function or rule to achieve the lighting effect. However, this simplification strategy brings serious problems: a) Its lighting modeling scheme is too simple, lacking a deep understanding of the physical processes of real lighting, and cannot capture the rich phenomena generated by the interaction of light with complex geometry and materials (such as soft shadows, specular changes, subsurface scattering, etc.), resulting in poor applicability and only being able to work in very limited scenarios; b) Similar to the first type of method, it can usually only handle very simple lighting effects, such as simulating a single electric point light source, and cannot cope with complex or natural lighting environments; c) Due to the lack of physical constraints and operation at the pixel level, this method is very prone to destroying the original structure, texture and color relationships of the image, resulting in severe distortion of the effect, producing unnatural colors, loss of details or artifacts.

[0163] The third approach employs a two-stage deep learning model, separating "lighting estimation" from "image rendering." This is a common paradigm for deep learning in lighting tasks. Typically, the first neural network estimates the lighting representation of the scene from the input image (or a specified reference lighting image), commonly using spherical harmonic (SH) coefficients. Then, the second neural network receives the original image and the estimated lighting parameters, responsible for rendering or synthesizing the final lit image. Despite some progress, this approach still faces challenges: a) Lighting decoupling is extremely difficult. Accurately separating the scene's intrinsic properties (such as reflectivity and geometry) from the external lighting information from a single image is a typical "ill-posed problem," and the estimated lighting parameters (such as SH coefficients) may be inaccurate or incomplete; b) While modifying the image's lighting and shadows according to the lighting parameters, the rendering network needs to accurately preserve the detailed features of the person (such as skin texture, hair strands, and facial contours). However, during complex network transformations, these high-frequency details are easily smoothed, blurred, or lost, affecting the realism of the portrait.

[0164] The fourth approach utilizes powerful diffusion models as generators. These methods treat lighting as a conditional image generation task. The model starts with random noise and progressively denoises it in multiple steps to generate the target image. The entire generation process is guided by a series of conditional information, such as text prompts describing the desired lighting effect, pre-estimated surface normal maps, user-specified global lighting intensity values, or other control signals. However, diffusion-based lighting also has inherent bottlenecks: a) Existing conditional control methods (text, normal maps, single intensity values, etc.) are often insufficient to accurately and unambiguously model complex, detailed, and spatially varied lighting effects. Text descriptions may be vague, normal maps may be inaccurate, and single intensity values ​​cannot express the spatial distribution and directionality of light, all of which limit the ability to control fine lighting effects; b) Diffusion models are highly sensitive to the consistency of guiding conditions. When there are conflicts or inconsistencies among multiple input control conditions (for example, the light source direction described in the text does not match the optimal lighting direction implied by the normal map), the model may have difficulty effectively integrating these contradictory information, leading to an unstable generation process and the final output image may contain messy, strange artifacts or unexpected results.

[0165] To improve user experience and address the problems of traditional methods, this application proposes a portrait lighting transfer scheme and system based on guided image pairs. Specifically, this application proposes a novel image and portrait lighting method based on a diffusion model architecture. Its core advantage lies in utilizing the lighting effect control information decoupled from the guided image pairs to achieve more accurate and natural lighting effect transfer of the target image. Compared to previous schemes relying on a single reference image or other control methods, this method introduces a pair of reference images containing the pre-lighting and post-lighting states. This paired visual input provides the model with crucial contrast information, enabling it to more accurately identify and separate the target lighting effect (e.g., changes in light source direction, intensity, color, shadows, etc.) purely caused by lighting changes by analyzing the differences in lighting effects between the two images. This effectively eliminates interference information such as high-frequency textures, material properties, or other static ambient light inherent in the image content itself. After accurately decoupling the target lighting effect, this scheme further utilizes the powerful image generation capabilities of the diffusion model to naturally integrate the learned lighting effect into the target image.

[0166] The strategy proposed in this application, based on guided image pairs and a diffusion model, exhibits several advantages and effectively overcomes many bottlenecks of traditional solutions:

[0167] First, this solution offers significant advantages in terms of application scenarios and flexibility. Users only need to provide the target image to be processed and a pair of reference images (before and after lighting) that can demonstrate the desired changes in lighting effects to achieve the transfer of lighting effects. This approach does not rely on specific 3D scene reconstruction, complex physical parameter settings, or predefined simple light source types (as limited by traditional physical a priori methods or simple parameter decoupling methods). Theoretically, it can transfer any observed complex lighting effects, whether it is subtle changes in natural light or dramatic artistic lighting, greatly expanding the scope of application and creative freedom.

[0168] Secondly, using image pairs as guiding conditions for lighting effects significantly improves the accuracy of lighting effect decoupling and the precision of control. Compared to traditional diffusion models that rely on text descriptions or normal maps, image pairs provide richer, more intuitive, and unambiguous lighting information. Text descriptions often fail to accurately convey the spatial distribution, intensity, and color details of complex lighting effects, easily leading to misunderstandings; while geometric information such as normal maps, although important, may be inaccurate in prediction and cannot fully represent lighting characteristics. When these control conditions are insufficient or contradictory, it can easily lead to chaotic generation results. This approach, by directly learning lighting effect differences from visual paradigms (image pairs), effectively avoids the inherent problem of "insufficient lighting effect expressiveness" in other control methods. This allows the model to more robustly capture and reproduce target lighting effects, while also mitigating the risks of lighting decoupling difficulties and loss of details in the previous two-stage models.

[0169] Finally, leveraging the powerful generative capabilities of the diffusion model architecture itself, the fusion of the transferred lighting effects with the target image is more natural and realistic. The diffusion model is renowned for its superior performance in high-fidelity image synthesis and maintaining content consistency. In this approach, using the diffusion model for final lighting effect rendering and fusion effectively avoids the "texture-like" or unrealistic feel caused by inaccurate geometry or material estimation in traditional physically based rendering methods, and also overcomes the severe distortion caused by simple pixel adjustment methods. The progressive denoising generation process of the diffusion model helps to smoothly and harmoniously integrate the decoupled lighting effects into the structure and texture of the target image, ultimately generating a target image with natural light and shadow transitions, good detail retention, and a realistic overall visual effect.

[0170] like Figure 3 The diagram shown illustrates the overall process of the image generation method provided in this application. This portrait lighting effect transfer scheme is based on a conditional diffusion model architecture. Its core idea is to use a pair of reference images (before and after lighting) to accurately extract the target lighting effect information, and then use these as a conditional guided diffusion model to relight the target image. The specific process is as follows: Figure 3 As shown: First, the system receives a pair of reference images as input: image img_a showing the original lighting effect and image img_A after processing for the target lighting effect. These two images are fed in parallel into a SigLip visual encoder (labeled 1), which extracts features characterizing the lighting environment from each image. Subsequently, the lighting effect features extracted from img_a and img_A are concatenated to form a combined feature vector containing contrasting information about lighting changes, i.e., "lighting effect feature information." This combined feature is passed to a lightweight adapter module, LightAdapter (labeled 2), which is specifically designed to further process the concatenated features, aiming to more accurately decouple and extract the pure "lighting effect condition" information representing the lighting changes from img_a to img_A.

[0171] Meanwhile, in the main generation pipeline, the target image `img_b`, which requires lighting processing, is fed into a standard Stable Diffusion 2.1 diffusion generation architecture (labeled 3). First, `img_b` is encoded into the latent space using a variational autoencoder (VAE) to obtain its latent representation. Then, following the forward pass of the diffusion model, noise is added to this latent representation. This noisy latent representation serves as the initial input and is fed into the core U-Net structure of the Stable Diffusion 2.1 model for iterative denoising. Crucially, at each step (or specific step) of denoising, the "lighting effect condition" information output from the LightAdapter is injected into the diffusion model (typically into the cross-attention layer of the diffusion model) to guide the model's generation process. While progressively removing noise and reconstructing the latent representation of the image, the model adjusts the lighting effects of the generated image according to this lighting effect condition, making it approximate the lighting changes shown in the reference image. Finally, the latent representation generated through multi-step denoising is sent back to the VAE decoder for decoding, restoring it to a pixel-space image, resulting in the final illuminated image img_B. In this way, the scheme accurately captures the lighting effect using the contrast information of image pairs, and naturally transfers this lighting effect to the target image by leveraging the generation capability of the diffusion model.

[0172] The innovations of this application include: 1. Designing and building an algorithmic framework for controlling the generation of such guiding images; 2. Selecting appropriate open-source algorithm modules, such as SigmaPu, or self-designed algorithm modules, such as LightAdapter, to address the characteristics of lighting issues.

[0173] like Figure 3 The algorithm scheme shown below has a framework diagram. The following is a detailed introduction to each component module:

[0174] 1.1 SigLip Vision Encoder

[0175] SigLip (Sigmoid Loss for Language Image Pre-Training) is a model and method proposed by Google researchers for pre-training visual-language images. Its core objective is similar to CLIP (Contrastive Language-Image Pre-Training), aiming to learn powerful visual representations, that is, encoding images into feature vectors (embeddings) that capture their rich semantic information. The main innovation of SigLip lies in the loss function used in its pre-training stage. Traditional CLIP models employ contrastive learning methods, typically using the InfoNCE loss function, which relies on distinguishing correct image-text pairs (positive samples) from a large number of incorrect pairs (negative samples) within a batch, and then normalizing using softmax. While effective, this approach is highly sensitive to batch size, requiring very large batches to contain a sufficient number of negative samples, and necessitates careful tuning of the temperature hyperparameter. SigLip serves as the feature extraction module in this scheme.

[0176] In this scheme, the input consists of two images from a reference image pair: the image before lighting (img_a) and the image after lighting (img_A). These images undergo preprocessing before being input into the model, such as resizing to a size acceptable to the model (e.g., 384x384 pixels) and segmenting them into image patches to fit the underlying VisionTransformer (ViT) architecture. The output of the SigLip visual encoder is image feature vectors, also known as image embeddings. These are high-dimensional vectors designed to capture the rich semantic information and visual content of the input images. The output is a 768-dimensional feature vector.

[0177] 1.2 LightAdapter (Light Effect Adapter)

[0178] The core function of LightAdapter is to receive and process combined feature information from the SigLip visual encoder. Specifically, it receives the concatenated result of concatenating the lighting effect feature information extracted from the unlit reference image (img_a) and the lit reference image (img_A). Through its internal structure (e.g., the design includes 5 modules containing self-attention and cross-attention), LightAdapter performs in-depth processing and analysis on this concatenated feature. Objective: To accurately decouple and generate lighting effect conditional vectors.

[0179] Input processing: The concatenated embeddings from the pre-lighting (img_a) and post-lighting (img_A) images are used as input. Using the image patch embedding sequence preserves more spatial information than using only a single global embedding vector for each image.

[0180] Iterative Refinement (5 Modules): The core of the adapter consists of 5 stacked modules, each performing self-attention, cross-attention, and feedforward operations. This iterative structure allows the model to progressively deepen its understanding of lighting changes.

[0181] The pseudocode for the specific network structure design is as follows:

[0182] Loop 5 times (i ranges from 0 to 4):

[0183] 1. Self-attention module (operating on H_A)

[0184] Allowing tokens within the "lit" image to interact and extract features.

[0185] attn_output_self = MultiHeadSelfAttention(H_A)

[0186] H_A = LayerNorm(H_A + attn_output_self) # Residual connection & layer normalization

[0187] 2. Cross-attention module (H_A queries H_a)

[0188] Allow features "after lighting" to focus on features "before lighting" in order to discover differences / correspondences.

[0189] attn_output_cross=MultiHeadCrossAttention(query=H_A,key=H_a, value=H_a)

[0190] H_A = LayerNorm(H_A + attn_output_cross) # Residual connection & layer normalization

[0191] 3. Feedforward network module (operating on H_A)

[0192] Standard MLP modules for further feature transformation

[0193] ff_output = FeedForward(H_A)

[0194] H_A = LayerNorm(H_A + ff_output) # Residual connection & layer normalization

[0195] Output processing:

[0196] The processed sequence H_A is pooled to obtain a single vector representation.

[0197] pooled_output=AveragePooling(H_A, dimension=sequence_dim )shape(768)

[0198] Alternative: If a special [CLS] token is pre-added, then its representation is used.

[0199] Final output: 768-dimensional light effect feature vector, Output = pooled_output # Shape (768)

[0200] LightAdapter is a module specifically designed for controlling light effects.

[0201] 1.3 Framework for Controlled Generation-Diffusion Model

[0202] This is a typical structure of a Conditional Latent Diffusion Model, whose design concept is very similar to Stable Diffusion 2.1, but it uses specific light effect information as a control condition. Its main components and workflow are as follows:

[0203] 1. Input Image Processing and Latent Space Coding (Encoder): The framework starts with the "Input Image".

[0204] The image is first fed into an encoder. In latent variable diffusion models like Stable Diffusion 2.1, this encoder is typically the encoding part of a variational autoencoder (VAE).

[0205] The VAE encoder compresses a high-resolution input image into a lower-dimensional but more information-dense latent space, generating a latent representation of the image. This is primarily to significantly reduce the computational complexity and memory requirements of the subsequent diffusion process, since the diffusion process takes place in a low-dimensional latent space rather than a high-dimensional pixel space.

[0206] 2. Conditional Information Generation (LightAdapter -> Light Effect Embedding):

[0207] Unlike the standard Stable Diffusion, which relies primarily on textual conditions, this framework uses a dedicated condition generation path.

[0208] A module called LightAdapter (whose input, based on previous discussions, is features from a pair of reference images) is responsible for processing and extracting information that represents the lighting effect of the target.

[0209] The output of LightAdapter is a "light effect embedding". This can be understood as a feature vector containing the required lighting transformation instructions (for example, a 768-dimensional light effect feature vector), which will serve as a control signal to guide the core generation process.

[0210] 3. Core diffusion and denoising process (U-Net diffusion process):

[0211] This is the core engine of the entire generative model. It receives the latent variable representation from the VAE encoder as its basis. According to the principle of the diffusion model, this latent variable will undergo a forward noise addition process (not explicitly shown in the figure, but implicit in the "diffusion process"), or it can be used directly as z_0 to start inverse denoising.

[0212] The core component is a U-Net architecture neural network, which is consistent with the noise predictor structure used in Stable Diffusion 2.1.

[0213] U-Net's task is to perform denoising iteratively across multiple time steps. At each step, it receives the current noisy latent variable z_t and the time step encoding t, and predicts the noise epsilon to be added to the clean latent variable.

[0214] The key control element is the "light effect embedding," which is injected into specific layers within U-Net (typically cross-attention layers). U-Net uses its intermediate features as a query to "attend to" this "light effect embedding" (as a key and value). Through this mechanism, the light effect information guides the direction of U-Net's noise prediction, enabling it to denoise while simultaneously evolving latent variables in a direction that aligns with the target light effect.

[0215] After a predetermined number of iterations of denoising, U-Net outputs an approximately clean latent variable representation z_0, which not only reconstructs the content of the input image but also incorporates the target lighting effect.

[0216] 4. Latent Space Decoding and Result Output (Decoder → Result Image):

[0217] The clean latent variable representation z_0 output by U-Net is fed into a decoder.

[0218] This decoder is typically the decoding part of the VAE, corresponding to the previous encoder. The function of the VAE decoder is to map latent variables from the low-dimensional latent space back to the high-dimensional pixel space, reconstructing the final user-visible image. The output of the decoder is the "result image," which is the final image after applying the target lighting effect.

[0219] This framework utilizes a Visual Image Array (VAE) to shift image processing to the more computationally efficient latent space. The core U-Net diffusion model (similar to Stable Diffusion 2.1) is responsible for progressively denoising and generating images in the latent space. Unlike standard SD, its generation process is not guided by text embeddings, but rather precisely controlled by a dedicated LightAdapter module that extracts "light effect embeddings" from reference image pairs through a cross-attention mechanism. Finally, a high-resolution result image incorporating the target light effect is obtained through a VAE decoder.

[0220] The beneficial effects of the technical solution in this application include:

[0221] This approach leverages a robust diffusion model-based generative architecture, innovatively employing guiding image pairs (reference images containing both pre- and post-lighting states) as core conditional information to achieve the transfer of lighting effects from the target image. This strategy, by allowing the model to directly learn and compare the visual differences before and after lighting changes, more accurately decouples pure lighting effect information and eliminates interference from inherent image content attributes. Subsequently, utilizing the diffusion model's superior image generation capabilities, the decoupled lighting effects are naturally and realistically integrated into the target image. Compared to previous lighting schemes, this method demonstrates significant advantages in flexibility, accuracy, and generation quality, bringing substantial application value.

[0222] First, this solution significantly enhances the flexibility and applicability of image lighting. Users only need to provide the target image to be processed and a pair of reference images (before and after lighting) showing the desired lighting effect changes to achieve lighting effect transfer. This completely eliminates the reliance on simulated 3D scenes and simple electric light source models based on traditional physical prior methods, and also transcends the limitations of simple parameter decoupling methods on lighting effect types. Theoretically, any lighting effect that can be captured through image pairs, whether it is complex natural lighting, subtle ambient light atmosphere, or artistic lighting or creative light painting effects from a professional studio, can be learned and transferred. This approach is also superior to solutions that rely on text descriptions, because image pairs can convey spatial details and light and shadow nuances that are difficult to express precisely in text. Its value lies in: greatly reducing the technical threshold for achieving complex and diverse lighting effects; users do not need professional lighting knowledge or 3D modeling skills to easily reproduce or create ideal lighting effects, significantly enhancing the freedom and convenience of image editing.

[0223] Secondly, using image pairs for guidance significantly improves the accuracy and precision of lighting effect decoupling. By comparing images before and after lighting, the model can focus more on the changes brought about by the lighting itself, effectively distinguishing lighting effects from the inherent properties of objects (such as texture, color, and geometry). This overcomes the difficulty in completely separating lighting information from content or style in schemes relying on single reference images (including some two-stage deep learning methods). Simultaneously, it directly avoids the problems of insufficient lighting effect expression and chaotic generation results caused by insufficient, ambiguous, or conflicting conditional information in diffusion models based on text, normal maps, or single intensity values. Image pairs provide a direct, rich, and unambiguous visual paradigm for defining lighting effect transformations, offering more precise and reliable control compared to indirect textual or geometric constraints. Its value lies in ensuring that the transferred lighting effects are more faithful to the user's intent and reducing unexpected artifacts or style changes introduced by inaccurate lighting decoupling or conditional conflicts. This makes lighting effects more controllable and predictable, improving the stability and reliability of editing.

[0224] Finally, leveraging the powerful generative capabilities of the diffusion model, this solution ensures that the transferred lighting effects blend seamlessly and with the target image in an extremely natural and high-quality manner. The diffusion model is recognized for its advantages in generating realistic images, maintaining content consistency, and preserving detailed textures. Using the diffusion model for final lighting effect synthesis effectively avoids the "texture-like" effect caused by inaccurate geometry and material estimations in physically based rendering methods, as well as the severe distortion and loss of detail resulting from simple pixel adjustments. The progressive generation process of the diffusion model helps to seamlessly integrate the learned lighting effect information with the existing content of the target image, generating a final image with smooth light and shadow transitions, well-preserved structure, and rich detail. Its value lies in significantly improving the visual quality and realism of the final output image, making the edited image look as if it were actually taken under new lighting conditions. This not only improves user satisfaction but also allows this technology to be better applied to professional scenarios with high image quality requirements.

[0225] In summary, this solution, by cleverly combining the precise light effect decoupling capability of guided image pairs with the natural generation capability of diffusion models, provides a more flexible, accurate, and high-quality solution for image and portrait lighting tasks, and is expected to unleash enormous potential in multiple fields such as mobile photography, content creation, and digital art.

[0226] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0227] Based on the same inventive concept, this application also provides an image generation apparatus for implementing the image generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more image generation apparatus embodiments provided below can be found in the limitations of the image generation method described above, and will not be repeated here.

[0228] In one exemplary embodiment, such as Figure 5 As shown, an image generation apparatus is provided, including: an acquisition module 502, an encoding module 504, a determination module 506, and a generation module 508, wherein:

[0229] The acquisition module 502 is used to acquire the image to be processed and a pair of reference images, wherein the pair of reference images is used to reflect the difference in light effect between reference images under different lighting conditions.

[0230] The encoding module 504 is used to encode the reference image pair to obtain an encoded image pair.

[0231] The determining module 506 is used to determine the light effect feature information of the encoded image pair based on the light effect contrast information contained in the encoded image pair.

[0232] The generation module 508 is used to generate a target image after lighting processing based on the image to be processed and the light effect feature information.

[0233] In one embodiment, the acquisition module is further configured to acquire the image to be processed corresponding to the triggered image selection operation in response to the triggered image selection operation; and to acquire the light effect display image corresponding to the triggered light effect selection operation in response to the triggered light effect selection operation; wherein the light effect display image includes a first light effect display image for displaying the original light effect effect and a second light effect display image for displaying the light effect after the target light effect processing; the device further includes: a combination module, configured to combine the first light effect display image and the second light effect display image into the reference image pair.

[0234] In one embodiment, the reference image pair includes a first reference image and a second reference image; the first reference image is used to display the original lighting effect, and the second reference image is used to display the image after processing for the target lighting effect; the apparatus further includes: a processing module, used to preprocess the first reference image and the second reference image to obtain preprocessed first reference image and second reference image; the encoding module is further used to extract features from the preprocessed first reference image and second reference image using a visual encoder to obtain a first lighting effect feature of the first reference image and a second lighting effect feature of the second reference image; wherein, the first lighting effect feature and the second lighting effect feature are the encoded image pair.

[0235] In one embodiment, the reference image pair includes a first reference image and a second reference image; the first reference image is used to display the original lighting effect, and the second reference image is used to display the image after processing for the target lighting effect; the device further includes: a processing module, used to concatenate a first lighting effect feature of the first reference image and a second lighting effect feature of the second reference image to obtain a concatenated combined feature; the combined feature is the lighting effect feature information of the coded image pair; or, concatenate the first lighting effect feature of the first reference image and the second lighting effect feature of the second reference image to obtain a concatenated combined feature; and perform lighting effect decoupling processing on the combined feature to obtain a lighting effect condition vector reflecting the lighting effect change between the first lighting effect feature and the second lighting effect feature; the lighting effect condition vector is the lighting effect feature information of the coded image pair.

[0236] In one embodiment, the combined feature is an image sequence feature; the processing module is further configured to iteratively process the image sequence feature through a light effect adapter to obtain a light effect condition vector that reflects the light effect change between the first light effect feature and the second light effect feature; wherein the dimension of the light effect condition vector is the same as the dimension of the first light effect feature and the second light effect feature.

[0237] In one embodiment, the apparatus further includes: a partitioning module, configured to partition the image sequence features into a first light effect feature sequence and a second light effect feature sequence during iterative processing of the image sequence features via a light effect adapter; the processing module is further configured to perform self-attention interaction processing on the first light effect feature sequence to obtain a consistency feature of the first light effect feature sequence; and to perform cross-attention processing on the first light effect feature sequence and the second light effect feature sequence to obtain a difference feature reflecting the difference between the first light effect feature sequence and the second light effect feature sequence; the apparatus further includes: a transformation module, configured to perform feature transformation on the consistency feature and the difference feature to obtain a target feature sequence; the processing module is further configured to perform pooling processing on the target feature sequence to obtain a target light effect feature of a preset dimension; the target light effect feature is the light effect condition vector.

[0238] In one embodiment, the image to be processed contains a portrait of the user; the generation module is further configured to generate a target image after lighting processing based on the image to be processed and the light effect feature information using a diffusion model; or, to generate a target image after lighting processing based on the image to be processed, the light effect feature information, and an image vector reflecting the user's identity features using the diffusion model; wherein the light effect feature information is the control condition of the diffusion model, the image to be processed is the input parameter of the diffusion model, and the image vector is extracted from the image to be processed.

[0239] Each module in the aforementioned image generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0240] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an image generation method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0241] Those skilled in the art will understand that Figure 6The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0242] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0243] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0244] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0245] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0246] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0247] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0248] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An image generation method, characterized in that, The method includes: Obtain a pair of images to be processed and reference images, wherein the pair of reference images is used to reflect the differences in light effect between reference images under different lighting conditions; The reference image pair is encoded to obtain an encoded image pair; Based on the light effect contrast information contained in the encoded image pair, the light effect feature information of the encoded image pair is determined; Based on the image to be processed and the light effect feature information, a target image after lighting processing is generated.

2. The method according to claim 1, characterized in that, The acquisition of the image to be processed and the reference image pair includes: In response to a triggered image selection operation, the image to be processed corresponding to the image selection operation is obtained; In response to a triggered light effect selection operation, a light effect display image corresponding to the light effect selection operation is obtained; wherein, the light effect display image includes a first light effect display image for displaying the original light effect and a second light effect display image for displaying the light effect after processing by the target light effect. The first light effect display image and the second light effect display image are combined to form the reference image pair.

3. The method according to claim 1, characterized in that, The reference image pair includes a first reference image and a second reference image; the first reference image is used to display the original lighting effect, and the second reference image is used to display the image after the target lighting effect has been processed. The method further includes: The first reference image and the second reference image are preprocessed to obtain the preprocessed first reference image and the second reference image; The process of encoding the reference image pair to obtain an encoded image pair includes: The first light effect feature of the first reference image and the second reference image are obtained by extracting features from the preprocessed first reference image and the second light effect feature of the second reference image through a visual encoder. Wherein, the first light effect feature and the second light effect feature are the encoded image pair.

4. The method according to claim 1, characterized in that, The reference image pair includes a first reference image and a second reference image; the first reference image is used to display the original lighting effect, and the second reference image is used to display the image after the target lighting effect has been processed. The step of determining the light effect feature information of the encoded image pair based on the light effect contrast information contained in the encoded image pair includes: The first light effect feature of the first reference image and the second light effect feature of the second reference image are concatenated to obtain a combined feature; the combined feature is the light effect feature information of the coded image pair; or... The first light effect feature of the first reference image and the second light effect feature of the second reference image are concatenated to obtain a combined feature after concatenation. The combined feature is then subjected to light effect decoupling to obtain a light effect condition vector that reflects the light effect change between the first light effect feature and the second light effect feature. The light effect condition vector is the light effect feature information of the coded image pair.

5. The method according to claim 4, characterized in that, The combined features are image sequence features; The process of decoupling the combined features to obtain a light effect condition vector reflecting the light effect change between the first and second light effect features includes: The image sequence features are iteratively processed by the light effect adapter to obtain a light effect condition vector that reflects the light effect change between the first light effect feature and the second light effect feature. The dimension of the light effect condition vector is the same as the dimension of the first light effect feature and the second light effect feature.

6. The method according to claim 5, characterized in that, The iterative processing of the image sequence features through a light effect adapter to obtain a light effect condition vector reflecting the light effect change between the first light effect feature and the second light effect feature includes: During the iterative processing of the image sequence features through the light effect adapter, the image sequence features are divided into a first light effect feature sequence and a second light effect feature sequence. The first light effect feature sequence is subjected to self-attention interaction processing to obtain the consistency features of the first light effect feature sequence; Cross-attention processing is performed on the first luminous effect feature sequence and the second luminous effect feature sequence to obtain difference features that reflect the differences between the first luminous effect feature sequence and the second luminous effect feature sequence; The consistency features and the difference features are transformed to obtain the target feature sequence; The target feature sequence is pooled to obtain target light effect features of a preset dimension; the target light effect features are the light effect condition vector.

7. The method according to claim 1, characterized in that, The image to be processed contains a portrait of the user; The step of generating the target image after lighting processing based on the image to be processed and the light effect feature information includes: Using a diffusion model, a target image after lighting processing is generated based on the image to be processed and the light effect feature information; or... Using the diffusion model, a target image after lighting processing is generated based on the image to be processed, the light effect feature information, and the image vector reflecting the user's identity features; Wherein, the light effect feature information is the control condition of the diffusion model, the image to be processed is the input parameter of the diffusion model, and the image vector is extracted from the image to be processed.

8. An image generation apparatus, characterized in that, The device includes: An acquisition module is used to acquire a pair of images to be processed and a pair of reference images, wherein the pair of reference images is used to reflect the difference in light effect between reference images under different lighting conditions; The encoding module is used to encode the reference image pair to obtain an encoded image pair; The determining module is used to determine the light effect feature information of the encoded image pair based on the light effect contrast information contained in the encoded image pair; The generation module is used to generate a target image after lighting processing based on the image to be processed and the light effect feature information.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.