Diffusion model-based intrinsic image decomposition method and system
By using the Diffusion Transformer module based on the diffusion model in the eigenimage decomposition, combining the diffusion feedforward neural network and the diffusion self-attention mechanism, the problem of difficulty in effectively capturing high-frequency and low-frequency feature information in the prior art is successfully solved, and high-quality intrinsic image and light image decomposition is achieved.
Patent Information
- Application Number
- CN202510043150.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively capture the high-frequency and low-frequency characteristic information of the image during the intrinsic image decomposition process, resulting in poor decomposition effect.
The Diffusion Transformer module based on the diffusion model is adopted, combining the diffusion feedforward neural network and the diffusion self-attention mechanism, and the identification and lighting images are generated through the Y-shaped architecture of the multi-layer Diffusion Transformer, and enhanced training is carried out through MSE loss comparison and SSIM loss comparison to constrain high-frequency and low-frequency feature information.
High-quality intrinsic image and light image decomposition is achieved, which reduces MSE loss by nearly 70% and DSSIM loss by nearly 30% compared with traditional CNN frameworks, significantly improving the robustness of the decomposition effect.
Smart Images

Figure CN119942302A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an intrinsic image decomposition method and system based on a diffusion model. Background Art
[0002] As a very important part of image processing, intrinsic images usually refer to images that retain the physical properties of objects in the image after removing the illumination component. These physical properties can help other image processing to learn features or classification recognition more efficiently. For example, many of the illumination enhancement, target detection and image segmentation use intrinsic images as input data. In order to obtain the intrinsic image, it is necessary to choose a suitable method to decompose it from the original image. According to the Retinex theory, an image can be represented as the process of incident light reflecting from the object itself to the human retina.
[0003] Therefore, the image can be represented as the product of the illumination image and the reflection image. According to the non-Lamborghini model, an image is not only the product of the illumination image and the reflection image, but also needs to add the specular reflection image. The product of the first two can be represented as a diffuse reflection image. However, under the Lambert model, an image can be represented as the product of the illumination image and the reflection image. The process of obtaining the illumination image and the reflection image from the input image is called intrinsic image decomposition. The process of intrinsic image decomposition itself is also an image generation problem. As a pathological end-to-end generation algorithm study, it also needs an excellent image generation framework as a foundation. In recent years, the development of diffusion models has achieved major breakthroughs in efficiency and performance in the field of artificial intelligence generation, such as the generation of images, videos, and voices.
[0004] These generation problems have achieved great breakthroughs in function and processing speed with the development of diffusion models. Based on this, the present invention proposes an intrinsic image decomposition method and system based on diffusion model. Summary of the invention
[0005] The present invention aims at solving the problems in the prior art, and the technical solution adopted in this application is as follows: In a first aspect, the present invention provides an intrinsic image decomposition method based on a diffusion model, comprising: Get the image and input the image into a Diffusion Transformer module that combines a diffusion feedforward neural network and a diffusion self-attention mechanism to obtain the intrinsic image and lighting images ; The intrinsic image and lighting images Respectively with the label intrinsic image MSE loss comparison and SSIM loss comparison are performed to obtain the output image of the basic framework.
[0006] Furthermore, the Diffusion Transformer module includes two parts. The first part combines the input image and the noise image as input data, and the input data is converted into feature representation through an embedding layer; for each diffusion step t, t is converted into a fixed-dimensional vector through an embedding function; the fixed-dimensional vector is combined with the input feature to guide the model to generate data at each time step; each time the input feature F passes through the time step T and is calculated by the Diffusion Transformer network, the process can be expressed as: (Formula 1) Represents 1×1 pointwise convolution.
[0007] The process of calculating feature F from time step T under the diffusion self-attention mechanism can be expressed as:
[0008] Let Q, K, V be the query, key, and value matrices respectively. Q, K, V can be expressed as Equation 3:
[0009] in and They represent 1×1 point convolution and 3×3 depth convolution respectively. The calculation process of the self-attention mechanism A can be expressed as formula 4: α is a learnable scaling parameter used to control the magnitude of the dot product of K and Q before applying the softmax function; Split represents the segmentation operation, and LN(F) represents layer normalization; the second part combines these modules into a Y network architecture, namely an encoder and two decoders, to obtain the output intrinsic image of the multi-layer Diffusion Transformer module and lighting images , and compare the MSE loss and SSIM loss with the label image to obtain the output image of the basic framework.
[0010] Further, the steps for MSE loss comparison include: In order to learn the basic framework of Diffusion Transformer, the output image is illuminated Lighting image with label The low-frequency feature information difference and The average pooling downsampling is performed three times respectively, and an MSE loss comparison is performed in each downsampling process, so that the low-frequency feature information at multiple scales can be learned; the average pooling process can be regarded as the process of retaining the low-frequency feature information of the image, and the multi-scale loss comparison can more evenly realize the low-frequency information constraint of the illumination image, so as to be closer to the label image.
[0011] Furthermore, the SSIM loss comparison steps include: In order to learn the high-frequency information difference between the intrinsic output image of the Diffusion Transformer basic framework and the label intrinsic image, the output image (as a feature map) and label image (as feature map) Through a core operator with eight fixed directions, convolution is performed to obtain the high-frequency edge information features of each image for perceptual loss comparison, thereby constraining the post-processing of the image output by the basic Diffusion Transformer framework. This constraint will help the edge information features of the generated intrinsic image to be closer to the label image; let the operator in the i-th direction be , the process of high-frequency information convolution MSE constraint can be expressed as shown in Equation 5: (Formula 5) It can be expressed as taking the image of the basic Diffusion Transformer framework as the feature map and using convolution kernels in eight directions to perform convolution operations to extract edge features. The difference between the two formulas is the difference in perceptual information loss between the generated image and the label image. Finally, the mean square error loss function is used for optimization iteration.
[0012] In a second aspect, the present invention provides an intrinsic image decomposition system based on a diffusion model, comprising: The Diffusion Transformer module is configured to: obtain an image, input the image into a Diffusion Transformer module that combines a diffusion feedforward neural network and a diffusion self-attention mechanism to obtain the intrinsic image and lighting images ; The information constraint module is configured to: transform the intrinsic image and lighting images Respectively with the label intrinsic image MSE loss comparison and SSIM loss comparison are performed to obtain the output image of the basic framework.
[0013] In a third aspect, the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the intrinsic image decomposition method based on the diffusion model described in the first aspect.
[0014] In a fourth aspect, the present invention provides an electronic device comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to execute an intrinsic image decomposition method based on a diffusion model as described in the first aspect.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are: The present invention provides an intrinsic image decomposition method based on a diffusion model, comprising: acquiring an image, inputting the image into a Diffusion Transformer module combining a diffusion feedforward neural network and a diffusion self-attention mechanism to obtain an intrinsic image and lighting images ; The intrinsic image and lighting images Respectively with the label intrinsic image The output image of the basic framework is obtained by comparing the MSE loss and the SSIM loss. The method has a clear process and is easy to implement. It only needs to input a diffuse reflection image to obtain the corresponding intrinsic image and illumination image. This method uses the diffusion model to gradually denoise and generate images. The generation process is stable and the image quality is high. The Transformer architecture has excellent performance in capturing long-distance dependencies and complex patterns to achieve high-quality generation and enhance learning through different image frequency feature constraints. According to the results on the standard test set, it can be seen that this method has high robustness. The method of intrinsic image decomposition was tested on the public ShapeNet database. The loss function obtained by the intrinsic image decomposition method based on the diffusion model proposed in the present invention is greatly reduced. The MSE loss results of the intrinsic image and illumination image of this method are reduced by nearly 70% compared with the training method of the CNN framework, and the DSSIM loss results are also reduced by nearly 30%. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0017] Figure 1The process of the intrinsic image decomposition method based on the diffusion model of the present invention; Figure 2 It is a composition architecture diagram of the modules of the present invention. DETAILED DESCRIPTION
[0018] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0019] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways than those described herein. Therefore, the present invention is not limited to the specific embodiments of the following disclosure.
[0020] Embodiment 1, as Figure 1-Figure 2 As shown, the intrinsic image decomposition method based on the diffusion model proposed in the present invention has the following specific steps: In order to improve the accuracy of intrinsic image decomposition, the present invention provides an intrinsic image decomposition method based on a diffusion model. The method is based on a Diffusion Transformer module that combines a diffusion feedforward neural network and a diffusion self-attention mechanism, and then adds enhanced training modules for post-processing of intrinsic images and illumination images respectively for further constraints. The method generates images by using the diffusion model itself in a step-by-step denoising manner, starting from a noise image and gradually approaching the target image. This progressive generation method allows the model to gradually improve the image quality in each step, and finally generates a very high-quality image, and its multi-step generation process can smoothly generate details and textures in the image, avoiding artifacts and unnatural details in some other generation models (such as GAN), and it is relatively easy to adjust hyperparameters and optimize the model architecture. On this basis, the step-by-step denoising process of the diffusion model and the powerful representation ability of the Transformer are combined, and the Diffusion Transformer achieves higher quality and rich diversity image generation. The Y-shaped architecture of the multi-layer Diffusion Transformer uses a common encoder to gradually extract high-level features of the image, while two identical decoders gradually reconstruct the high-resolution output image. The jump connection between the encoder and the decoder directly transfers the feature map of the corresponding layer to the decoder part, retaining high-resolution information so that the decoder has more details when reconstructing the image. Two images are output through the basic framework. Compared with the illumination image, the intrinsic image emphasizes high-frequency edge information more, and the edge information can better represent the edge features and detail information of the image. The low-frequency feature information of the illumination image can better perceive the large-area changes and smooth information of the region. The edge feature extraction method of the fixed-direction operator convolution and the average pooling low-frequency extraction are used to process the enhanced training and learning of the intrinsic image and the illumination image respectively. The present invention has a simple structure and can effectively decompose after the image to be decomposed is input.
[0021] The technical solution adopted by the present invention to solve the technical problem is: a method for decomposing an intrinsic image based on a diffusion model is proposed. Based on the development of the diffusion model, the method selects the Diffusion Transformer model framework to achieve high-quality generation and constrains high-frequency and low-frequency feature information through edge information extraction and average pooling multi-scale comparison to achieve enhanced training, thereby decomposing the input image into two images. Its features include the following steps: Step 1: Use a Y-shaped multi-layer Diffusion Transformer module to reconstruct the input data. This module consists of two parts: the first part is to combine the input image and the noise image into input data The input data is first converted into feature representation through an embedding layer.
[0022] The basic framework of Diffusion Transformer is generated. Diffusion Transformer is an architecture that combines the diffusion model and the Transformer network for generation tasks. The core idea of this architecture is to use the diffusion process to gradually generate data, while using the Transformer's self-attention mechanism to capture long-range dependencies in the data. Put the input data into the Diffusion Transformer module. For each diffusion step t, t is converted into a vector of fixed dimension through an embedding function. This vector is combined with the input features to guide the model to generate data at each time step. Each layer of the Diffusion Transformer module includes a diffusion self-attention mechanism and a diffusion feedforward neural network structure as shown in Equations 1 and 2. The process of each input data being put into the Diffusion Transformer calculation can be expressed as: (Formula 1) (Formula 2) in represents a nonlinear activation function, represents layer normalization, DFN represents the diffusion forward propagation network architecture that puts time step T into the input feature F for calculation. DFSA represents the process of putting time step T into the input feature F for diffusion self-attention mechanism calculation. Q, K, V are query, key and value matrices respectively. and They represent 1×1 point convolution and 3×3 depth convolution respectively. Q, K, and V are obtained as in Formula 3 and the process of calculating the three is expressed as Formula 4: Let Q, K, V be the query, key, and value matrices respectively. Q, K, V represent: (Formula 3) in and They represent 1×1 point convolution and 3×3 depth convolution respectively, and the calculation process of the self-attention mechanism A is expressed as: (Formula 4) α is a learnable scaling parameter that controls the magnitude of the dot product of K and Q before applying the softmax function. Split represents the split operation, LN(F) represents layer normalization, and A represents self-attention calculation. Predict the noise component of the current time step t and output the prediction result through the linear layer. The learning and reconstruction process is implemented using a multi-layer Diffusion Transformer architecture. The second part is to combine these modules into a Y network architecture, namely an encoder and two decoders. The entire framework performs three downsampling and a single module output to achieve encoding and two identical three upsampling decodings and perform jump connections respectively to reconstruct and output the intrinsic image and lighting images , and compare the MSE loss and SSIM loss with the label image to obtain the output image of the basic framework.
[0023] Step 2: High-frequency feature information constraints, in order to learn the intrinsic image output by the Diffusion Transformer basic framework Intrinsic image with label The high-frequency feature information difference of and By convolving eight core operators with fixed directions to obtain their respective high-frequency edge information features and then comparing the losses, constraints are imposed on the intrinsic image output by the DiffusionTransformer framework for post-processing. This constraint will help the edge information features of the generated intrinsic image to be closer to the label image. Let the direction operator be , the output image (as a feature map) is , the label image (as a feature map) is , the process of image high-frequency information processing can be expressed as, and the calculation process is expressed by formula 5: ; (Formula 5) in It can be expressed as taking the image of the basic Diffusion Transformer framework as the feature map and using convolution kernels in eight directions to perform convolution operations to extract edge features. The difference between the two formulas is the difference in perceptual information loss between the generated image and the label image. Finally, the mean square error loss function is used for optimization iteration.
[0024] Step 3: Low-frequency feature information constraints, in order to learn the basic framework of Diffusion Transformer to illuminate the output image Lighting image with label The low-frequency feature information difference of the illumination image is enhanced by learning training. The average pooling module can help extract the low-frequency information of the image. and Three average pooling modules are downsampled respectively to extract multi-scale low-frequency features. Each downsampling process is compared with the MSE loss. The illumination image output in step 1 is and label light image Multi-scale low-frequency features are extracted separately and then the MSE loss is compared. This allows learning of low-frequency feature information at multiple scales. The average pooling process can be regarded as a process of retaining the low-frequency feature information of the image. Multi-scale loss comparison can more evenly implement the low-frequency information constraints of the illumination image, so as to be closer to the label image.
[0025] Embodiment 2, the present invention provides an intrinsic image decomposition system based on a diffusion model, comprising: The Diffusion Transformer module is configured to: obtain an image, input the image into a Diffusion Transformer module that combines a diffusion feedforward neural network and a diffusion self-attention mechanism to obtain the intrinsic image and lighting images ; The information constraint module is configured to: transform the intrinsic image and lighting images Respectively with the label intrinsic image MSE loss comparison and SSIM loss comparison are performed to obtain the output image of the basic framework.
[0026] Embodiment 3, the present invention provides a computer-readable storage medium, the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute an intrinsic image decomposition method based on a diffusion model described in Embodiment 1.
[0027] Embodiment 4, the present invention provides an electronic device, including a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to execute an intrinsic image decomposition method based on a diffusion model described in Embodiment 1.
[0028] The electronic device may include: a processor, a memory and a communication unit. These components communicate via one or more buses. Those skilled in the art will appreciate that the structure of the electronic device does not limit the embodiments of the present invention, and it may be a bus structure or a star structure, or a combination of certain components, or different component arrangements.
[0029] The communication unit is used to establish a communication channel so that the electronic device can communicate with other devices, receive user data sent by other devices or send user data to other devices.
[0030] The processor is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It runs or executes software programs and / or modules stored in the memory, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor can only include a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.
[0031] The memory is used to store the execution instructions of the processor. The memory can be implemented by any type of volatile or non-volatile storage device or a combination of them, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0032] When the execution instructions in the memory are executed by the processor, the electronic device is enabled to execute part or all of the steps of Embodiment 1.
[0033] The above description is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. A diffusion model-based intrinsic image decomposition method, characterized in that: include: Get the image and input the image into a Diffusion Transformer module that combines a diffusion feedforward neural network and a diffusion self-attention mechanism to obtain the intrinsic image and lighting images ; The intrinsic image and lighting images Respectively with the label intrinsic image Label Light Image Perform MSE loss comparison and SSIM loss comparison to obtain the output image of the basic framework; The Diffusion Transformer module includes two parts. The first part combines the input image and the noise image as input data, and the input data is converted into feature representation through an embedding layer. For each diffusion step t, t is converted into a fixed-dimensional vector through an embedding function. The fixed-dimensional vector is combined with the input feature to guide the model to generate data at each time step. The process of each input data being put into the Diffusion Transformer calculation can be expressed as: Let Q, K, V be the query, key, and value matrices respectively. Q, K, V represent: in and They represent 1×1 point convolution and 3×3 depth convolution respectively, and the calculation process of the self-attention mechanism A is expressed as: α is a learnable scaling parameter used to control the magnitude of the dot product of K and Q before applying the softmax function; Split represents the split operation, LN(F) represents layer normalization, and A represents the calculation process of the self-attention mechanism; the second part combines these modules into a Y network architecture, that is, an encoder and two decoders, to obtain the output intrinsic image of the multi-layer DiffusionTransformer module and lighting images , and compare the MSE loss and SSIM loss with the label image to obtain the output image of the basic framework.
2. The intrinsic image decomposition method based on the diffusion model as claimed in claim 1, characterized in that: include: The steps for comparing MSE loss include: generating illumination images to learn the basic framework of Diffusion Transformer Image with label The perceptual difference of low-frequency feature information will and The average pooling downsampling is performed three times respectively, and an MSE loss comparison is performed in each downsampling process, so that perceptual learning can be performed on the low-frequency feature information at multiple scales; the average pooling process can be regarded as the process of retaining the low-frequency feature information of the image, and the multi-scale loss comparison can more evenly realize the low-frequency information constraint of the illumination image, so as to be closer to the label image.
3. The intrinsic image decomposition method based on the diffusion model as claimed in claim 1, characterized in that: include: The SSIM loss comparison steps include: In order to learn the high-frequency information difference between the intrinsic output image of the Diffusion Transformer basic framework and the label intrinsic image, the output image and label image By convolving eight core operators in fixed directions, the high-frequency edge information features of each image are compared for loss, so as to impose constraints on the post-processing of the image output by the basic Diffusion Transformer framework; this constraint will help the edge information features of the generated intrinsic image to be closer to the label image; let the operator in the i-th direction be , the process of high-frequency information convolution MSE constraint is expressed as: ; In order to use the image of the basic Diffusion Transformer framework as a feature map, convolution operations are performed using convolution kernels in eight directions to extract edge features. The difference between the two formulas is the difference in perceptual information loss between the generated image and the label image. Finally, the mean square error loss function is used for optimization iteration.
4. An intrinsic image decomposition system based on a diffusion model, characterized in that: include: The Diffusion Transformer module is configured to: obtain an image, input the image into a Diffusion Transformer module that combines a diffusion feedforward neural network and a diffusion self-attention mechanism to obtain the intrinsic image and lighting images ; The information constraint module is configured to: transform the intrinsic image and lighting images Respectively with the label intrinsic image MSE loss comparison and SSIM loss comparison are performed to obtain the output image of the basic framework.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is run, the device where the computer-readable storage medium is located is controlled to execute the intrinsic image decomposition method based on the diffusion model as described in any one of claims 1 to 3.
6. An electronic device, characterized in that: It comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to execute an intrinsic image decomposition method based on a diffusion model as described in any one of claims 1-3.