Dynamic prior diffusion intrinsic image decomposition method

Through the dynamic prior diffusion method, combined with prior feature extraction and U-Net network, the problem of poor decomposition effect in complex environments in the existing technology is solved, and efficient image decomposition effect is achieved, which is suitable for diverse lighting conditions and complex backgrounds.

CN120707877APending Publication Date: 2025-09-26NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510780362.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing intrinsic image decomposition methods have poor decomposition effects in complex environments, making it difficult to meet the needs of real-time and resource-constrained scenarios. In addition, it is difficult to accurately decompose high-quality reflection images and illumination images under diverse lighting conditions and complex backgrounds.

Method used

The dynamic prior diffusion method is adopted to extract prior implicit features and conditional vectors through the first original image feature prior network. The diffusion model and U-Net network are combined for decomposition. The dynamic gated feedforward network and multi-head attention network are used to aggregate contextual information and capture global spatial information to optimize the image decomposition process.

Benefits of technology

The model's parameter efficiency and computational efficiency are improved, image details are effectively captured, and it is suitable for diverse lighting conditions and complex backgrounds. The controllability of decomposition effect and computational complexity is enhanced, estimation errors are reduced, and the accuracy and stability of image decomposition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707877A_ABST
    Figure CN120707877A_ABST
Patent Text Reader

Abstract

The invention discloses an intrinsic image decomposition method based on dynamic prior diffusion, and particularly relates to the field of image decomposition. Comprising the following steps: performing prior feature extraction on an original image to obtain prior implicit features; performing feature extraction on the original image to obtain a condition vector; inputting the priori implicit features and the condition vectors into a diffusion model, converting the priori implicit features into noise in the diffusion process, and removing the noise in combination with the condition vectors in the inverse diffusion process to obtain diffusion implicit features; and inputting the original image and the diffusion implicit features into a U-Net network, and guiding the original image to decompose by using the diffusion implicit features to obtain an illumination image or a reflection image. Based on the method, the intrinsic image decomposition effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image decomposition, and in particular to an intrinsic image decomposition method based on dynamic prior diffusion. Background Art

[0002] Intrinsic image decomposition extracts intrinsic reflectance features (such as color and texture) from the target, separating the effects of illumination variations on its appearance and providing a data foundation for object recognition under varying lighting conditions. In recent years, deep learning networks based on codec architectures have been widely used in intrinsic image analysis, but they are prone to detail loss and inaccurate reconstruction. Diffusion models overcome the limited latent space representation capabilities of codec networks by progressively denoising and meticulously modeling the image distribution. These models can also improve the quality of generated images through infilling and restoration. With the continuous advancement of deep learning technology, intrinsic image decomposition methods have significantly improved. By leveraging large-scale image data for training, these methods are able to extract more complex and comprehensive features. However, their significant computational and memory requirements pose challenges for practical applications. This results in inefficiency when processing large datasets, making it difficult to meet real-time and resource-constrained requirements. Furthermore, existing models often struggle to accurately decompose high-quality reflectance and illumination maps when faced with diverse lighting conditions, complex background environments, and different types of objects. Summary of the Invention

[0003] The main purpose of this application is to provide a dynamic prior diffusion intrinsic image decomposition method, which aims to solve the problem of poor decomposition effect in complex environmental backgrounds in existing decomposition methods.

[0004] To achieve the above-mentioned objectives, the present application provides a dynamic prior diffusion intrinsic image decomposition method, comprising: extracting prior features from the original image to obtain prior implicit features, the prior implicit features including space, color and texture; extracting features from the original image to obtain a conditional vector; the conditional vector is a labeled abstract representation; inputting the prior implicit features and the conditional vector into a diffusion model, converting the prior implicit features into noise during the diffusion process, and combining the conditional vector to remove the noise in the inverse diffusion process to obtain a diffusion implicit feature; inputting the original image and the diffusion implicit features into a U-Net network, using the diffusion implicit features to guide the decomposition of the original image to obtain an illumination map or a reflection map; wherein the encoder and decoder of the U-Net network are both Transformer modules, the Transformer module includes a dynamic gated feedforward network and a dynamic multi-head attention network, and the processing process of the Transformer module includes: aggregating contextual information of different scales on the diffusion implicit features and the original image through the dynamic gated feedforward network to obtain a first output feature; aggregating global spatial information on the first output feature and the diffusion implicit features through the dynamic multi-head attention network to obtain a second output feature.

[0005] Optionally, the prior implicit features are obtained through the first original image feature prior network, and the conditional vector is obtained through the second original image feature prior network; the first original image feature prior network includes a downsampling layer, a convolutional neural module, a residual module, a first convolutional layer, a pooling layer, a first linear layer and a first activation function layer; the prior feature extraction process includes: performing pixel reordering operation on the input image through the downsampling layer to obtain downsampling features; performing feature extraction on the downsampling features through the convolutional neural module to obtain a first output feature; performing residual processing on the first output feature through the residual module to obtain a second output feature; extracting features from the second output feature through the first convolution layer, and reducing the feature map dimension through the pooling layer to obtain low-dimensional features; mapping the low-dimensional features to the output space through the first linear layer, and increasing the nonlinear expression capability through the first activation function layer to obtain the prior implicit features.

[0006] Optionally, the first original image feature prior network, the second original image feature prior network, the diffusion model and the U-Net network are obtained through pre-training, and the training process includes: obtaining a training sample set, the training sample set including the original image, the illumination map and the reflection map; splicing the decomposition image and the original image to obtain a spliced ​​image; wherein the decomposition image is an illumination map or a reflection map; inputting the spliced ​​features into the first original image feature prior network to obtain prior latent features; inputting the prior latent features and the original image into the U-Net network to complete the first stage of training; inputting the original image into the second original image feature prior network to obtain a conditional vector; inputting the prior latent features and the conditional vector into the diffusion model to obtain a diffusion latent feature; inputting the diffusion latent features and the original image into the U-Net network obtained by the first stage of training to complete the second stage of training.

[0007] Optionally, the context information of different scales of the diffusion implicit feature and the original image are aggregated through a dynamic gated feedforward network to obtain a first output feature, including: linearly transforming the diffusion implicit feature through two second linear layers respectively to obtain a first linear feature and a second linear feature; performing layer normalization on the original image through a third linear layer to obtain a first normalized image; performing element-by-element multiplication of the first linear feature and the first normalized image to obtain a first fusion feature; adding the first fusion feature and the second linear feature to obtain a second fusion feature; aggregating the second fusion feature through a first convolution module and a second convolution module respectively to obtain a first linear fusion feature and a second linear fusion feature; processing the first linear fusion feature through a fourth activation function layer to obtain a nonlinear fusion feature; performing element-by-element multiplication of the nonlinear fusion feature and the second linear fusion feature, and transforming through a fifth convolution layer to obtain a transformed feature; performing a residual connection on the transformed feature and the original image to obtain the first output feature.

[0008] Optionally, a dynamic multi-head attention network is used to aggregate global spatial information of the first output feature and the diffuse latent feature to obtain a second output feature, including: linearly transforming the first output feature through a fourth linear layer to obtain a second normalized image; linearly transforming the diffuse latent feature through two fifth linear layers respectively to obtain a third linear feature and a fourth linear feature; performing element-by-element multiplication of the third linear feature and the second normalized image to obtain a third fused feature; adding the third fused feature and the fourth linear feature to obtain a fourth fused feature; performing a deconvolution operation on the fourth fused feature through three parallel sixth convolutional layers respectively to obtain a query, a key and a value; determining the similarity between the query and the key to generate a transposed attention map; multiplying the transposed attention map with the value to obtain a weighted feature, convolving the weighted feature through the convolution layer, and then adding it to the first output feature through a residual connection to obtain the second output feature.

[0009] Optionally, the prior implicit features are converted into noise during the forward diffusion process, and the noise is removed in combination with the conditional vector during the reverse diffusion process to obtain the diffusion implicit features, including: during the diffusion process, the prior implicit features are converted into pure noise; the estimated noise of each time step is determined, and the image representation of the current time step is updated using the estimated noise and the conditional vector; when the time step t=1, the image representation of the current time step is pure noise; during the reverse diffusion process, the image representation of the current time step is denoised to obtain the image representation of the previous time step; the image representation of the previous time step is used as the image representation of the current time step, and the estimated noise of the current time step is re-determined based on the conditional vector and the image representation of the current time step, and the next iteration is performed until the total number of iterations is reached to obtain the diffusion implicit features.

[0010] Optionally, the second original image feature prior network has the same structure as the first original image feature prior network.

[0011] To achieve the above-mentioned purpose, the present application also provides an intrinsic image decomposition device with dynamic prior diffusion, comprising: a first original image feature prior network, used to extract prior features of the original image to obtain prior implicit features; a second original image feature prior network, used to extract features of the original image to obtain a conditional vector; a diffusion model, used to convert the prior implicit features into noise during the diffusion process, and combine the conditional vector to remove the noise during the inverse diffusion process to obtain the diffusion implicit features; a U-Net network, used to use the diffusion implicit features to guide the decomposition of the original image to obtain an illumination map or a reflection map; wherein, the intrinsic image decomposition device is obtained through pre-training, and the loss functions in the training process of the intrinsic image decomposition device include Charbonnier loss, edge loss and perceptual loss.

[0012] Compared with the prior art, the present invention has the following advantages: The present invention discloses a dynamic prior diffusion intrinsic image decomposition method, which extracts prior implicit features through a first original image feature prior network, and recovers diffuse implicit features from the prior implicit features through a diffusion model. The diffuse implicit features are used as multi-scale prior guidance in the U-Net network module, which improves the parameter efficiency, computational efficiency and simplicity of data representation of the model, can effectively capture image details and enhance the intrinsic image decomposition effect. The first original image feature prior network first performs a downsampling operation and then performs prior feature extraction, which can extract implicit representations containing key information while reducing computing resource consumption. The diffuse implicit features and the original image are aggregated with contextual information of different scales through a dynamic gated feedforward network to expand the receptive field while keeping the computational complexity controllable. The method is suitable for image decomposition under diverse lighting conditions, complex background environments and different types of objects. The dynamic multi-head attention mechanism is used to effectively aggregate global spatial information, capture the intrinsic connection of the data, and enhance the model's feature encoding ability. The intrinsic image decomposition starts from a specific time step, runs all denoising iterations, and sends the results to the Transformer module for joint optimization, thereby reducing estimation errors and improving performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a flow chart of the dynamic prior diffusion intrinsic image decomposition method of this application; Figure 2 Schematic diagram of the refinement process of the intrinsic image decomposition method with dynamic prior diffusion in this application; Figure 3 This is a schematic diagram of the structure of the first original image feature prior network in the intrinsic image decomposition method of dynamic prior diffusion in this application; Figure 4 This is a schematic diagram of the structure of the dynamic feedforward network in the intrinsic image decomposition method with dynamic prior diffusion in this application; Figure 5 This is a schematic diagram of the structure of the dynamic multi-head attention network in the intrinsic image decomposition method with dynamic prior diffusion in this application; Figure 6 This is the result diagram of Example 1; Figure 7 This is the reflection map comparison result of the CG-Intrinsic dataset in Example 1; Figure 8 This is a comparison chart of the effects of the reflection map, illumination map, and label map generated by different image decomposition methods in Example 1 on the MIT dataset; Figure 9 This is a comparison chart of the effects of reflection maps, illumination maps, and label maps generated by different image decomposition methods in Example 1 on the ShapeNet dataset.

[0014] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0015] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0016] The first embodiment of the present invention provides a dynamic prior diffusion intrinsic image decomposition method, such as Figure 1 As shown, the specific steps include: Step S1: Input the original image into the first original image feature prior network to extract the prior features and obtain the prior implicit features. Among them, the prior implicit features include space, color and texture, which are compressed representations of the intrinsic visual attributes of the image; Specifically, the first original image feature prior network includes a downsampling layer, a convolutional neural module, a residual module, a first convolutional layer, a pooling layer, a first linear layer and a first activation function layer; Figure 2 As shown, the prior feature extraction process of step S10 includes: Step S11, performing a pixel unshuffle operation (PixelUnshuffle) on the input image through the downsampling layer to obtain downsampling features; Step S12, performing feature extraction on the downsampled features through the convolutional neural module to obtain a first output feature; Specifically, the convolutional neural module includes a second convolutional layer, a second activation function layer, a third convolutional layer, a third activation function layer, and a fourth convolutional layer; the feature extraction process of the convolutional neural module in step S12 includes: Step S121, performing local feature extraction on the downsampled features through a second convolutional layer, and converting the local features into first nonlinear features through a second activation function layer; In step S122, feature extraction, nonlinear transformation, and feature extraction are performed on the first nonlinear feature through the third convolutional layer, the third activation function layer, and the fourth convolutional layer, respectively, to obtain a first output feature. The convolutional neural module first extracts local features through each convolutional layer and then uses the LReLU activation function to solve the "dead neuron" problem.

[0017] Step S13, performing residual processing on the first output feature through the residual module to obtain a second output feature; Specifically, the residual module includes multiple residual blocks; each residual block contains multiple convolutional layers and LRelu activation functions, and the first nonlinear features and the second nonlinear features are input into the residual layer for residual processing to obtain the second output feature, which can effectively alleviate the gradient disappearance problem and improve learning ability and expressiveness.

[0018] Step S14: extract features from the second output features through the first convolutional layer, and reduce the dimension of the feature map through the pooling layer to obtain low-dimensional features; this can reduce the amount of calculation and the number of parameters.

[0019] Step S15: Map the low-dimensional features to the output space through the first linear layer, and increase the nonlinear expression capability through the first activation function layer to obtain the prior implicit features.

[0020] Step S2: inputting the original image into a pre-trained second original image feature prior network to extract features and obtain a conditional vector, which is a labeled abstract representation; The second original image feature prior network (S2) has the same structure as the first original image feature prior network, but the first convolutional layer has different input dimensions to accommodate feature extraction from the original image, and the output is different. The conditional vector D provides the diffusion model with important information about how to recover the image from noise, and its expression is:

[0021] The pixel reordering operation adjusts the image resolution during this process, increasing it by rearranging pixels to meet the network input size requirements. During the inverse diffusion process, due to the presence of pre-trained implicit prior information, the second original image feature prior extraction network can use fewer iterations and a smaller model size to obtain better estimates than traditional diffusion models.

[0022] In step S3, the prior implicit features and conditional vectors are input into the diffusion model. During the diffusion process, the prior implicit features are converted into noise, and during the inverse diffusion process, the noise is removed by combining with the conditional vector to obtain the diffusion implicit features. The prior implicit features are used to drive the diffusion process, and the conditional vector is a task-related control signal and is only injected as an external constraint in the denoising stage. This separation design achieves a balance between generation quality and control flexibility.

[0023] Specifically, in step S31, during the diffusion process, the prior implicit feature is transformed into pure noise. , whose probability distribution is defined as:

[0024] in, represents the total number of iterations, represents the time step embedding, ; In step S32, the denoising network determines the estimated noise at the current time step and updates the image representation of the current time step using the estimated noise, the conditional vector, and the image representation of the current time step. The specific update formula is as follows: when time step t = 1, the image representation of the current time step is the pure noise obtained in step S31. When t is greater than 1, the image representation of the current time step is obtained by the following formula:

[0025] Where, is the time step The estimated noise, is the image representation of the current time step, is the index of the time step.

[0026] Step S33, in the reverse diffusion process, denoise the image representation of the current time step obtained by step S32 to obtain the image representation of the previous time step, which is expressed as follows:

[0027] in, is a random noise term, usually sampled from a standard normal distribution. In the inverse diffusion process, this noise term is used to simulate the noise added in the original diffusion process.

[0028] In step S34, the image representation of the previous time step is used as the image representation of the current time step, and the process returns to step S33 for the next iteration until the total number of iterations is reached to obtain the diffusion implicit feature.

[0029] In this example, S2 is used to extract a conditional vector from the original image, a key step in image restoration. Next, the denoising network estimates the noise at each time step and updates the image representation, gradually restoring the original image quality. After a certain number of iterations, the final estimated implicit prior information is obtained. This approach more effectively utilizes denoising information, improving the accuracy and stability of image restoration.

[0030] In step S4, the original image and the diffusion latent features are input into a U-Net network, and the diffusion latent features are used to guide the decomposition of the original image to obtain an illumination map or a reflectance map. The encoder and decoder of the pre-trained U-Net network are both Transformer modules, which include a dynamic gated feedforward network and a dynamic multi-head attention network. The processing method step S4 specifically includes: Step S41, aggregating context information of different scales on the second input feature and the original image through the dynamic gated feedforward network to obtain a first output feature; when decomposing, the second input feature is a diffusion latent feature; Specifically, such as Figure 3 As shown, in step S411, the second input feature is linearly transformed by two second linear layers to obtain a first linear feature and a second linear feature; the original image is layer-normalized by a third linear layer to obtain a first normalized image; the first linear feature and the first normalized image are element-wise multiplied to obtain a first fused feature; the first fused feature and the second linear feature are added to obtain a second fused feature, and the calculation formula is as follows:

[0031] in, and They are the input and output feature maps, and in this step F is the original image, is the second fusion feature, the modulation matrix is a learnable parameter. This step enables the multi-attention Transformer to dynamically adjust the feature map according to the implicit prior information to better restore the details and texture of the image.

[0032] Step S412, respectively aggregating the second fusion feature through the first convolution module and the second convolution module to obtain a first linear fusion feature and a second linear fusion feature, wherein the first convolution module and the second convolution module both include Convolutional layers and Deep convolution layer; the first linear fusion feature is processed by the fourth activation function layer to obtain a nonlinear fusion feature. The fourth activation function layer is a GELU activation function, which helps to introduce nonlinear connections and enables the network to learn more complex features; the nonlinear fusion feature and the second linear fusion feature are element-wise multiplied and a The fifth convolutional layer performs transformation to obtain transformation features; the transformation features are residually connected with the original image to obtain the first output feature, which is expressed as follows:

[0033] in, is the first output feature, and are learnable weight matrices, which are used to perform linear transformation on the input features. It is called element-wise product and represents the multiplication of corresponding elements of two matrices or vectors of the same dimension. Is a nonlinear activation function used to introduce nonlinear characteristics. The residual connection adds the input features to the transformed features, helping to alleviate the vanishing gradient problem in deep networks. This structure is used in deep learning to improve model performance and training stability.

[0034] In this embodiment, the dynamic gated feedforward network uses a gating mechanism that allows the network to dynamically adjust its behavior based on the input data, thereby improving the model's adaptability and performance. This design is particularly suitable for tasks that require dynamic adjustment of feature representations based on context.

[0035] Step S42: Aggregate the global spatial information of the first output feature and the second input feature through the dynamic multi-head attention network to obtain the second output feature.

[0036] Specifically, such as Figure 4 As shown, in step S421, the first output feature is linearly transformed by the fourth linear layer to obtain a second normalized image; the second input feature is linearly transformed by two fifth linear layers to obtain a third linear feature and a fourth linear feature; the third linear feature is element-wise multiplied by the second normalized image to obtain a third fused feature; the third fused feature is summed with the fourth linear feature to obtain a fourth fused feature, and the calculation formula is as follows:

[0037] Among them, in this step F is the first output feature, It is the fourth fusion feature.

[0038] Step S422, respectively performing a deconvolution operation on the fourth fused feature through three parallel sixth convolutional layers to obtain a query, a key, and a value; determining the similarity between the query and the key, generating a transposed attention map, which reflects the correlation between different regions in the feature map; multiplying the transposed attention map with V to obtain a weighted feature, convolving the weighted feature through a convolutional layer, and adding it to the first output feature through a residual connection to obtain a second output feature , the expression is as follows:

[0039] Where, is a learnable scaling parameter, is a learnable matrix, Function that converts the output of a neural network into a probability distribution.

[0040] In this embodiment, a dynamic multi-head attention mechanism is used to effectively aggregate global spatial information, capture the intrinsic connections of the data, and enhance the model's ability to encode features.

[0041] Exemplarily, the U-Net network in this embodiment has a three-layer structure.

[0042] The device used in the intrinsic image decomposition method of the present invention includes a first original image feature prior network, a second original image feature prior network, a diffusion model and a U-Net network. The first original image feature prior network, the second original image feature prior network, the diffusion model and the U-Net network are obtained through pre-training. The training process includes two stages, such as Figure 5 As shown, the details are as follows.

[0043] Stage 1: Training of the first original image feature prior network and U-Net network, see Figure 5 Middle stage 1; A training sample set is obtained, the training sample set including an original image, an illumination map, and a reflectance map; the decomposition map is spliced ​​with the original image to obtain a spliced ​​image, wherein the decomposition map is an illumination map or a reflectance map; the spliced ​​features are input into a first original image feature prior network to obtain a priori latent features; since the input of the first original image feature prior network during training is a spliced ​​image of the original image and the illumination map, the priori latent features contain mixed information of space, color, and texture, and are a compressed representation of the intrinsic visual attributes of the image, directly participating in the model backbone training and driving the diffusion process; the priori latent features and the original image are input into a U-Net network (Transformer module) to complete the first stage of training; in this stage, the image generation capability is optimized through a dynamic gated feedforward network and a dynamic multi-head attention network; Specifically, in the first stage, the training process of the first original image feature prior network includes: Step 10: Obtain a training sample set, which includes the original image, its illumination map, and its reflectance map. To expand the training set, specifically, for each image sample, randomly extract M small image blocks of size N×N. All small image blocks are horizontally flipped, generating a corresponding flipped image block for each image block. This ultimately generates 2M small image blocks for each image sample (including M original image blocks and M flipped image blocks). The small image blocks obtained from all image processing are aggregated to obtain the processed training sample set.

[0044] Step 20: splicing the decomposition image with the original image to obtain a spliced ​​image; wherein the decomposition image is an illumination image and a reflection image; during the training process, the original image and the decomposition image are both expanded; In step 30, the splicing features are input into the first original image feature prior network (S1) to extract prior features and obtain prior implicit features. The expression of the prior implicit feature Z is:

[0045] in, is a reflectance map or a lighting map, is the input original image, It means to stitch two images together. is a downsampling operation, and Represents the first original image feature prior network. In this way, the first original image feature prior network can extract implicit representations containing key information while reducing computing resource consumption.

[0046] It is worth noting that in this embodiment, the first original image feature prior network is pre-trained using the original image and the corresponding illumination map (or the original image and the corresponding reflection map), and the extracted features are converted into the latent space. Therefore, during training, the input of the first original image feature prior network is the spliced ​​image of the original image and the corresponding illumination map. When in use, the input of the first original image feature prior network is the original image. Correspondingly, when the input is the spliced ​​image of the original image and the corresponding illumination map, the output of the first original image feature prior network is the prior implicit features related to the illumination map, and the final decomposition result is the illumination map; when the input is the spliced ​​image of the original image and the corresponding reflection map, the output is the prior implicit features related to the reflection map, and the decomposition result is the reflection map.

[0047] Phase II, diffusion model and U-Net network (trained in Phase I), see Figure 5 middle stage2; The original image is fed into a second original image feature prior network to generate a conditional vector. The prior latent features and conditional vector are then fed into a diffusion model, where they are converted into pure noise through the forward process of the diffusion model (the first half of stage 2). The noise is then removed layer by layer in a denoising process, combining them with the conditional vector D to ultimately generate the true implicit prior features. The diffused latent features and the original image are then fed into the U-Net network trained in the first stage (the second half of stage 2), completing the second stage of training. In this stage, the prior latent features enter the dynamically modulated parameters of the U-Net network, guiding the intrinsic image decomposition process. This two-stage training process first learns a compact prior representation of the image that captures important image features. This allows the network to focus on accurately estimating this prior representation from the original image during the diffusion model training phase, rather than learning the entire image generation process from scratch.

[0048] The second embodiment of the present invention provides an intrinsic image decomposition device with dynamic prior diffusion, comprising: a first original image feature prior network, used to extract prior features of the original image to obtain prior implicit features; a second original image feature prior network, used to extract features of the original image to obtain a conditional vector; a diffusion model, used to convert the prior implicit features into noise during the diffusion process, and to remove the noise by combining the conditional vector during the inverse diffusion process to obtain a diffusion implicit feature; a U-Net network, used to use the diffusion implicit feature to guide the decomposition of the original image to obtain an illumination map or a reflection map.

[0049] The entire intrinsic image decomposition device is pre-trained. The loss functions in the training process of the intrinsic image decomposition device include Charbonnier loss, edge loss and perceptual loss, as follows.

[0050]

[0051] This is achieved using the Charbonnier loss. Compared to the standard L1 loss function, the Charbonnier loss function adds a small constant ε for smoothing, effectively avoiding the non-differentiability of the L1 loss at zero, thereby ensuring the stability of the optimization process. The Charbonnier loss exhibits greater numerical stability, which is particularly important in applications such as image reconstruction and super-resolution.

[0052]

[0053] Edge detection methods like the Sobel operator enhance the model's ability to capture image edges, ensuring the edges of the generated image are clear and consistent with the target image. By integrating edge detection algorithms into the loss function, the generated image maintains a high degree of consistency in edge detail with the original image.

[0054]

[0055] Perceptual loss is an advanced loss function that measures the difference between the generated image and the target image by simulating the human visual system's evaluation of image perceptual quality. The core of this loss function is that it not only focuses on the differences at the pixel level, but also considers the perceptual characteristics of the image at a deeper level, such as texture, color, and overall structure. In implementation, perceptual loss usually uses pre-trained deep learning models, such as the VGG network, to extract feature representations of the image. By comparing the differences between the generated image and the target image at these feature levels, perceptual loss can more accurately reflect the human eye's perception of image quality. The key to this method is that it can capture subtle differences that may not be noticeable at the pixel level but are very important in visual perception; .

[0056] Example 1 Experiments were conducted on the ShapeNet, MPI-Sintel, and CG-Intrinsic datasets, and the algorithm performance was comprehensively evaluated from two dimensions: visual display and quantitative indicators.

[0057] The convolutional neural module of the first original image feature prior network contains five residual blocks with convolution kernels of 512, 256, 128, 64, and 32, respectively. This allows for early capture of rich details and gradual feature abstraction, ultimately resulting in a comprehensive feature representation. By optimizing residual connections and linear transformations, the network effectively extracts image features while maintaining its lightweight, laying the foundation for subsequent tasks.

[0058] Training and testing were performed using the NVIDIA A100 GPU on the Google Colab platform. This high-performance graphics processing unit significantly accelerates the training of deep learning models. To ensure the stability and reproducibility of experimental results, a uniform random seed strategy was adopted to ensure the same initial conditions for each experiment, making the results more reliable.

[0059] In network training, the Adam optimizer was selected, which combines the advantages of AdaGrad and RMSProp and can adaptively adjust the learning rate. The following parameters were set for it: the first-order moment estimation coefficient ( ) is set to 0.1, the second-order moment estimation coefficient ( ) was set to 0.99, and the weight decay parameter was set to 1e-4. After multiple experiments, we ultimately chose 1000 epochs as the training cycle. This epoch length ensures the model has sufficient time to learn the data characteristics while avoiding unnecessary overfitting. The initial learning rate was set to 0.0001, which helps avoid excessively large update steps in the early stages of training, thereby ensuring more stable convergence.

[0060] The number of samples processed per batch is uniformly set to 16. Input images are resized to a resolution of 256×256×3. This size is based on a balance between memory usage and training efficiency. A smaller batch size may increase noise during training, while a larger batch size may increase memory consumption and even lead to training instability. The value of 16 strikes a good balance between the two.

[0061] Experimental results and analysis (1) Combination of different modules Table 1 Experimental results of different module combinations

[0062] Tested on the CG-Intrinsic dataset, Table 1 shows that the decomposition device (stages 1 and 2) of this application outperforms either stage 1 (the original image feature prior network and the U-Net network) or stage 2 (the diffusion model and the U-Net network) alone in predicting both reflectance and illumination maps. This demonstrates that the combination of the U-Net network and the diffusion model can improve the performance of image decomposition tasks. Under the DSSIM evaluation criteria, stage 1 and stage 2 significantly outperformed each of the individual processes. Figure 7 For CG-Intrinsic reflectance image visualization, using only stage one will affect edge details, and using only stage two will affect color restoration. The complete network structure confirms that the image processed by these two stages maintains a high degree of structural consistency with the original image, thereby providing better image quality in practical applications.

[0063] (2) Loss function analysis Table 2 Experimental results of different loss functions

[0064] Table 2 shows the experimental results under different loss functions. The experimental data clearly demonstrates that by fusing Charbonnier loss, perceptual loss, and Sobel loss, the overall performance of the image processing algorithm can be significantly enhanced, especially in reducing the mean square error (MSE) and optimizing the structural similarity index (DSSIM).

[0065] (3) Performance on the MIT dataset Table 3 Experimental results of MIT dataset

[0066] In terms of local mean square error (LMSE), it is proved that this method has more advantages in local processing. Figure 8In this paper, a visual comparative analysis of the effects of reflectance maps, illumination maps and label maps produced by different image restoration methods was conducted. The results revealed that traditional algorithms, such as SIRFS, can achieve a good brightness level when reconstructing reflectance maps, but are not fine enough in capturing the details of illumination maps. At the same time, deep learning methods such as USI3D have shortcomings in reconstructing the details of reflectance maps, and the accuracy of colors needs to be improved. In contrast, the serial intrinsic image decomposition method proposed in this invention has shown obvious advantages in image restoration. This method not only achieved remarkable results in improving the contrast of illumination maps, but also demonstrated excellent performance in reconstructing the details of reflectance maps, and the training time is shorter.

[0067] (4) Performance on the ShapeNet dataset Table 4 Experimental results of ShapeNet dataset

[0068] As shown in Table 4, the dynamic prior diffusion intrinsic image decomposition method has an advantage on the ShapeNet dataset, showing the lowest error and the highest similarity in terms of mean square error, local mean square error and structural similarity. Figure 9 The visual comparison results of different methods are presented, from which we can observe that traditional techniques such as DI have shortcomings in color restoration, specifically manifested in overly bright colors and inaccurate edge details. Deep learning-based methods, while showing improvements in overall performance, still have some deviations in color rendering in local areas. This may be due to the deep learning model's insufficient capture of local features during training, or the model's lack of sensitivity to certain types of image features. In contrast, the method of this invention demonstrates superior results in color restoration and edge accuracy.

[0069] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A dynamic prior diffusion intrinsic image decomposition method, characterized in that: include: Extracting prior features from the original image to obtain prior implicit features; the prior implicit features include spatial, color, and texture information; Extracting features from the original image to obtain a conditional vector; the conditional vector is a labeled abstract representation; Inputting the prior implicit feature and the conditional vector into a diffusion model, converting the prior implicit feature into noise during the diffusion process, and removing the noise by combining the conditional vector in the inverse diffusion process to obtain a diffusion implicit feature; Inputting the original image and the diffusion latent features into a U-Net network, and using the diffusion latent features to guide the decomposition of the original image to obtain an illumination map or a reflection map; The encoder and decoder of the U-Net network are both Transformer modules, which include a dynamic gated feedforward network and a dynamic multi-head attention network. The processing of the Transformer module includes: The dynamic gated feedforward network is used to aggregate contextual information of different scales on the diffusion implicit features and the original image to obtain a first output feature; the dynamic multi-head attention network is used to aggregate global spatial information on the first output feature and the diffusion implicit features to obtain a second output feature.

2. The dynamic prior diffusion intrinsic image decomposition method according to claim 1, characterized in that: The prior implicit feature is obtained through a first original image feature prior network, and the conditional vector is obtained through a second original image feature prior network; the first original image feature prior network includes a downsampling layer, a convolutional neural module, a residual module, a first convolutional layer, a pooling layer, a first linear layer, and a first activation function layer; the prior feature extraction process includes: Performing pixel reordering on the input image through the downsampling layer to obtain downsampled features; Performing feature extraction on the downsampled features through the convolutional neural module to obtain a first output feature; Performing residual processing on the first output feature by the residual module to obtain a second output feature; Extract features from the second output features through the first convolutional layer, and reduce the dimension of the feature map through the pooling layer to obtain low-dimensional features; The low-dimensional features are mapped to the output space through the first linear layer, and the nonlinear expression capability is increased through the first activation function layer to obtain the prior implicit features.

3. The dynamic prior diffusion intrinsic image decomposition method according to claim 2, characterized in that: The first original image feature prior network, the second original image feature prior network, the diffusion model, and the U-Net network are obtained through pre-training, and the training process includes: Acquire a training sample set, wherein the training sample set includes an original image, an illumination map, and a reflection map; Splicing the decomposition image with the original image to obtain a spliced ​​image; wherein the decomposition image is an illumination image or a reflection image; Inputting the splicing features into a first original image feature prior network to obtain a priori latent features; Input the prior implicit features and the original image into the U-Net network to complete the first stage of training; Inputting the original image into a second original image feature prior network to obtain a conditional vector; Inputting the priori latent features and the conditional vector into a diffusion model to obtain a diffusion latent feature; The diffusion latent features and the original image are input into the U-Net network obtained by the first stage training to complete the second stage training.

4. The dynamic prior diffusion intrinsic image decomposition method according to claim 1, characterized in that: The aggregating context information of different scales on the diffusion implicit features and the original image through the dynamic gated feedforward network to obtain the first output feature includes: Performing linear transformation on the diffusion implicit feature through two second linear layers respectively to obtain a first linear feature and a second linear feature; Performing layer normalization on the original image through a third linear layer to obtain a first normalized image; Performing element-by-element multiplication on the first linear feature and the first normalized image to obtain a first fused feature; Adding the first fusion feature and the second linear feature to obtain a second fusion feature; Aggregating the second fused features through a first convolution module and a second convolution module respectively to obtain a first linear fused feature and a second linear fused feature; Processing the first linear fusion feature through a fourth activation function layer to obtain a nonlinear fusion feature; Performing element-wise product on the nonlinear fusion feature and the second linear fusion feature, and transforming the feature through a fifth convolutional layer to obtain a transformed feature; Perform a residual connection between the transformed feature and the original image to obtain a first output feature.

5. The dynamic prior diffusion intrinsic image decomposition method according to claim 4, characterized in that: The aggregating global spatial information of the first output feature and the diffusion implicit feature through the dynamic multi-head attention network to obtain the second output feature includes: Performing a linear transformation on the first output feature through a fourth linear layer to obtain a second normalized image; Performing linear transformation on the diffusion implicit feature through two fifth linear layers respectively to obtain a third linear feature and a fourth linear feature; Performing element-by-element multiplication on the third linear feature and the second normalized image to obtain a third fused feature; Adding the third fusion feature and the fourth linear feature to obtain a fourth fusion feature; Performing a deconvolution operation on the fourth fused feature through three parallel sixth convolutional layers to obtain a query, a key, and a value; Determine the similarity between the query and the key and generate a transposed attention graph; The transposed attention map is multiplied by the value to obtain a weighted feature, the weighted feature is convolved by the convolution layer, and then added to the first output feature through the residual connection to obtain the second output feature.

6. The dynamic prior diffusion intrinsic image decomposition method according to claim 1, characterized in that: The prior implicit feature is converted into noise in the forward diffusion process, and the noise is removed by combining the conditional vector in the reverse diffusion process to obtain the diffusion implicit feature, including: In the diffusion process, the prior implicit features are transformed into pure noise; Determine the estimated noise at each time step, and use the estimated noise and conditional vector to update the image representation of the current time step; when time step t=1, the image representation of the current time step is pure noise; In the inverse diffusion process, the image representation of the current time step is denoised to obtain the image representation of the previous time step; The image representation of the previous time step is used as the image representation of the current time step. The estimated noise of the current time step is determined again based on the conditional vector and the image representation of the current time step. The next iteration is performed until the total number of iterations is reached to obtain the diffusion implicit feature.

7. The dynamic prior diffusion intrinsic image decomposition method according to claim 1, characterized in that: The second original image feature prior network has the same structure as the first original image feature prior network.

8. A dynamic prior diffusion intrinsic image decomposition device, characterized in that: include: The first original image feature prior network is used to extract prior features of the original image and obtain prior implicit features; A second original image feature prior network is used to extract features from the original image to obtain a conditional vector; A diffusion model is used to convert the prior implicit features into noise during the diffusion process, and to remove the noise by combining the conditional vector in the inverse diffusion process to obtain the diffusion implicit features; A U-Net network is used to guide the decomposition of the original image using the diffusion implicit features to obtain an illumination map or a reflection map; The intrinsic image decomposition device is obtained through pre-training, and the loss function in the training process of the intrinsic image decomposition device includes Charbonnier loss, edge loss and perceptual loss.