Intrinsic image decomposition method based on interactive image semantic information constraint

Through the interactive image semantic information constraint method, using the cross-attention mechanism and multi-layer perceptron, the problem of inaccurate separation of reflection characteristics and illumination characteristics in complex scenes is solved, the precise reconstruction of reflection maps and illumination maps is achieved, and the accuracy and consistency of image decomposition are improved.

CN120655911APending Publication Date: 2025-09-16NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510677280.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing methods are not accurate enough in separating reflection characteristics from illumination characteristics in complex scenes and lack clear semantic information constraints, resulting in large errors in the reconstruction of reflection maps and illumination maps.

Method used

Through interactive image semantic information constraints, semantic segmentation and feature extraction are adopted, and the cross-attention mechanism and multi-layer perceptron are used to extract reflection features and illumination features respectively, and the reconstruction results are optimized through end-to-end training.

Benefits of technology

It achieves fine and accurate feature separation in complex lighting environments and texture-rich scenes, improves the global consistency and local detail integrity of image decomposition, and ensures the precise decoupling and coordinated optimization of reflection features and illumination features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655911A_ABST
    Figure CN120655911A_ABST
Patent Text Reader

Abstract

The invention discloses an intrinsic image decomposition method based on interactive image semantic information constraint, and particularly relates to the field of computer vision and image processing. Comprising the steps of performing semantic segmentation and semantic feature extraction on an original image to obtain semantic segmentation features; performing feature extraction on the original image to obtain original image features; performing feature enhancement on the original image features to obtain enhanced original image features; obtaining reflection features by using common features of the original image features and the semantic segmentation features; obtaining illumination features by using difference features of the original image features and the semantic segmentation features; reconstructing the enhanced original image features and the reflection features to obtain a reflection image; and reconstructing the enhanced original image features and the illumination features to obtain an illumination image. Based on the method, the accuracy of reconstructing the illumination image and the reflection image can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer vision and image processing, and in particular to an intrinsic image decomposition method based on interactive image semantic information constraints. Background Art

[0002] During image acquisition, factors such as lighting variations, object material, relative position, and acquisition angle can significantly impact the image results. This is especially true under extreme conditions such as strong light or shadows, which can lead to loss of image detail. These factors complicate image recognition and understanding tasks in complex environments. Intrinsic Image Decomposition (IID) aims to decompose images into a reflectance map (reflecting the inherent properties of objects, such as color and texture) and an illumination map (reflecting the scene's lighting distribution), thereby eliminating environmental interference and improving the robustness of computer vision models.

[0003] In recent years, deep learning-based intrinsic image decomposition methods have made significant progress. However, existing methods still have shortcomings in reconstructing accurate images, especially in complex scenes, where the separation of reflectance and illumination characteristics is not precise enough. Some studies have attempted to improve performance by jointly training semantic segmentation and intrinsic image decomposition, but the lack of clear semantic information constraints prevents the full utilization of semantic prior knowledge. Furthermore, existing networks do not adequately exploit global information, and the detail and consistency of the reconstruction results need to be improved. Summary of the Invention

[0004] The main purpose of this application is to provide an intrinsic image decomposition method based on interactive image semantic information constraints, aiming to solve the problem that existing methods cannot fully utilize semantic prior knowledge, resulting in large errors in the reconstruction of reflectance maps and illumination maps.

[0005] To achieve the above-mentioned objectives, the present application provides an intrinsic image decomposition method based on interactive image semantic information constraints, including: performing semantic segmentation on the original image to obtain semantic labels; performing multi-scale feature extraction on the semantic labels to obtain semantic segmentation features; performing feature extraction on the original image to obtain original image features; performing feature enhancement on the original image features to obtain enhanced original image features; determining common features of the original image features and the semantic segmentation features, and fusing the common features with the original image features and the semantic segmentation features to obtain reflection features; determining difference features between the original image features and the semantic segmentation features, and fusing the difference features with the original image features and the semantic segmentation features to obtain illumination features; reconstructing the enhanced original image features and the reflection features to obtain a reflection map; and reconstructing the enhanced original image features and the illumination features to obtain an illumination map.

[0006] Optionally, the common features of the original image features and the semantic segmentation features are determined, and the common features are fused with the original image features and the semantic segmentation features to obtain reflection features, including: fusing the common information of the original image features and the semantic segmentation features to obtain the common features; fusing the common features with the original image features to obtain the first reflection feature; fusing the common information of the first reflection feature with the semantic segmentation feature to obtain the second reflection feature; and fusing the second reflection feature with the difference feature to obtain the reflection feature.

[0007] Optionally, the common features of the original image features and the semantic segmentation features are determined, and the common features are fused with the original image features and the semantic segmentation features to obtain reflected features, including: performing a first cross-attention operation on the query matrix of the original image features and the key matrix and value matrix of the semantic segmentation features to obtain the common features of the original image features and the semantic segmentation features; performing a first cross-attention operation on the query matrix of the common features and the key matrix and value matrix of the original image features to obtain a first reflected feature; performing a first cross-attention operation on the query matrix of the first reflected feature and the key matrix and value matrix of the semantic segmentation features to obtain a second reflected feature; and performing element-by-element addition on the query matrix of the second reflected feature and the query matrix of the first fused feature to obtain the reflected feature.

[0008] Optionally, the difference features between the original image features and the semantic segmentation features are determined, and the difference features are fused with the original image features and the semantic segmentation features to obtain the illumination features, including: determining the difference features between the original image features and the semantic segmentation features; fusing the difference features with the original image features for common information to obtain a first illumination feature; fusing the first common feature with the semantic segmentation feature for common information to obtain a second illumination feature; and fusing the second common feature with the difference features to obtain the illumination feature.

[0009] Optionally, the difference features between the original image features and the semantic segmentation features are determined, and the difference features are fused with the original image features and the semantic segmentation features to obtain illumination features, including: performing a second cross-attention operation on the query matrix of the original image features and the key matrix and value matrix of the semantic segmentation features to obtain the difference features between the original image features and the semantic segmentation features; performing a first cross-attention operation on the query matrix of the difference features and the key matrix and value matrix of the original image features to obtain a first illumination feature; performing a first cross-attention operation on the query matrix of the first illumination feature and the key matrix and value matrix of the semantic segmentation feature to obtain a second illumination feature; and performing element-by-element addition on the query matrix of the second illumination feature and the query matrix of the difference feature to obtain an illumination feature.

[0010] Optionally, the first cross-attention operation process includes: determining the similarity between the query matrix and the key matrix, and normalizing them to obtain attention weights; applying the attention weights to the value matrix to generate weighted features; adding the weighted features and the query vector element by element to obtain residual connection features; normalizing the residual connection features and performing feature enhancement through a multi-layer perceptron to obtain enhanced features; adding the enhanced features and the residual connection features to obtain the first cross-attention operation result.

[0011] Optionally, the second cross-attention operation process includes: determining the similarity between the query matrix of the original image feature and the key matrix of the semantic segmentation feature, and normalizing it to obtain the attention weight; applying the attention weight to the value matrix of the semantic segmentation feature to generate a weighted feature; determining the first difference feature between the original image feature and the semantic feature based on the weighted feature and the value matrix of the semantic segmentation feature; performing a residual connection on the first difference feature and the query vector of the original image feature to obtain a second difference feature; normalizing the second difference feature and performing feature enhancement through a multi-layer perceptron to obtain a third difference feature; adding the third difference feature to the second difference feature to obtain a difference feature.

[0012] To achieve the above-mentioned purpose, the present application also provides an intrinsic image decomposition device based on interactive image semantic information constraints, including: a semantic feature extraction module, which is used to perform semantic segmentation on the original image to obtain semantic labels; and perform multi-scale feature extraction on the semantic labels to obtain semantic segmentation features; an encoder, which is used to extract features from the original image to obtain original image features; a feature enhancement module, which is used to perform feature enhancement on the original image features to obtain enhanced original image features; a reflection attention module, which is used to determine the common features of the original image features and the semantic segmentation features, and fuse the common features with the original image features and the semantic segmentation features to obtain reflection features; an illumination attention module, which is used to determine the difference features of the original image features and the semantic segmentation features, and fuse the difference features with the original image features and the semantic segmentation features to obtain illumination features; a decoder, which is used to reconstruct the enhanced original image features and the reflection features to obtain a reflection map; and also used to reconstruct the enhanced original image features and the illumination features to obtain an illumination map; wherein the intrinsic image decomposition device is obtained through pre-training, and the loss function during the training process is the sum of Charbonnier loss, perceptual loss and CCR loss.

[0013] Optionally, the reflection attention module includes: a first common information module, used to perform common information fusion on the original image features and the semantic segmentation features to obtain a common feature; a second common information module, used to perform common information fusion on the common feature and the original image features to obtain a first reflection feature; a third common information module, used to perform common information fusion on the first reflection feature and the semantic segmentation feature to obtain a second reflection feature; a first fusion module, used to fuse the second reflection feature with the difference feature to obtain a reflection feature.

[0014] Optionally, the illumination attention module includes: a difference information module, used to determine the difference features between the original image features and the semantic segmentation features; a fourth common information module, used to perform common information fusion of the difference features and the original image features to obtain a first illumination feature; a fifth common information module, used to perform common information fusion of the first common feature and the semantic segmentation feature to obtain a second illumination feature; and a second fusion module, used to fuse the second common feature with the difference feature to obtain an illumination feature.

[0015] Compared with the prior art, the present invention has the following advantages: The intrinsic image decomposition method based on interactive image semantic information constraints of the present invention captures the global semantic information shared by reflection and illumination through the common information module, thereby enhancing the network's understanding of the overall structure and semantic relationship in complex scenes; and focuses on mining the unique detail features of reflection and illumination through the difference information module, ensuring fine and accurate feature separation in complex lighting environments or scenes with rich textures, thereby effectively realizing the precise decoupling and collaborative optimization of image reflection features and illumination features.

[0016] By alternately using the cross-attention mechanism to obtain the common features of the original image features and semantic segmentation features, and fusing the common features with the original image features and semantic segmentation features, it is possible to effectively fuse semantic information and high-frequency reflection features, thereby improving the global consistency and local detail integrity of the decomposition results; and through residual connection and normalization technology, the stability and expression ability of the reflection features are enhanced, and the nonlinear transformation ability of the features is further improved by using the multi-layer perceptron, making the reconstruction of the reflection map more refined and accurate, thereby fully utilizing the semantic prior knowledge.

[0017] Through the common features of semantic segmentation features, the difference features between the original image features and the semantic segmentation features are determined, and the difference features are used to constrain the reconstruction of the illumination map, thereby realizing the constraint of semantic information on the reconstructed image. It can ensure that the reflection characteristics and illumination characteristics complement each other and further improve the accuracy of the reconstructed illumination map; and design a pixel-level + perception + contrast structure multi-scale loss, and perform end-to-end joint training on the reflection-illumination branch, so that the reconstruction result is better in both detail texture and structural similarity. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flowchart of the intrinsic image decomposition method based on interactive image semantic information constraints in this application; Figure 2 This is a flowchart for extracting reflection features in the intrinsic image decomposition method based on interactive image semantic information constraints in this application; Figure 3 This is a flowchart for extracting illumination features in the intrinsic image decomposition method based on interactive image semantic information constraints in this application; Figure 4 This is a flowchart of the first cross-attention operation in the intrinsic image decomposition method based on interactive image semantic information constraints in this application; Figure 5 This is a flowchart of the second cross-attention operation in the intrinsic image decomposition method based on interactive image semantic information constraints in this application; Figure 6 This is the decomposition result diagram of Example 1 in the intrinsic image decomposition method based on interactive image semantic information constraints of this application.

[0019] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0021] The first embodiment of the present invention provides an intrinsic image decomposition method based on interactive image semantic information constraints, such as Figure 1 As shown, the specific steps include: Step S1: semantically segment the original image to obtain semantic labels (semantic prior knowledge); perform multi-scale feature extraction on the semantic labels to obtain semantic segmentation features; Specifically, an interactive image semantic segmentation algorithm is used to perform semantic segmentation on the original image to obtain semantic labels, and a first encoder is used to perform multi-scale feature extraction on the semantic labels to obtain semantic segmentation features.

[0022] Step S2, extracting features from the original image to obtain original image features; Specifically, the second encoder extracts features from the original image to obtain original image features. Since this embodiment requires obtaining a reflectance map and an illumination map, which are two independent processes, two second encoders are required to extract features from the original image to obtain two original image features. Furthermore, the second encoder and corresponding decoder in this embodiment are those of a traditional UNet network. Their structures are not within the scope of protection of the present invention, and the specific processing of the encoder and decoder will not be further described here.

[0023] Step S3, performing feature enhancement on the original image features to obtain enhanced original image features; Specifically, the feature enhancement module enhances the original image features to obtain enhanced original image features. Furthermore, the feature enhancement module can be a compression and activation network (SE-Net), which sequentially enhances key channel features, suppresses redundant information, and improves the quality of the original image features through global average pooling, weight learning, and feature recalibration.

[0024] Step S4, determining the common features of the original image features and the semantic segmentation features, fusing the common features with the original image features and the semantic segmentation features to obtain the reflection features; specifically, the reflection features are obtained through the reflection attention module, such as Figure 2 As shown in Figure 2, the operation flow of the reflective attention module is as follows.

[0025] In step S41, common information fusion is performed on the original image features and the semantic segmentation features to obtain common features, namely, initial reflection features. The common information fusion is achieved through the first cross-attention operation, the purpose of which is to obtain the common features between the two features. Therefore, the specific process of step S41 is as follows: Perform linear transformations on the original image features and semantic segmentation features to generate the original image feature query matrix Q, key matrix K, and value matrix V, as well as the original image feature query matrix Q, key matrix K, and value matrix V of the semantic segmentation features. Perform the first cross-attention operation on the original image feature query matrix Q, the key matrix K, and the value matrix V of the semantic segmentation features through the first common information module (CIM) to obtain the common features of the original image features and the semantic segmentation features. Step S42: performing common information fusion on the common features and the original image features to obtain a first reflected feature. Specifically, performing a linear transformation on the common features to generate a query matrix Q of the common features. A first cross-attention operation is performed on the query matrix Q of the common features and the key matrix K and value matrix V of the original image features by a second common information module to obtain a first reflected feature, i.e., a common feature between the common features and the original image features. Step S43: performing common information fusion on the first reflection feature and the semantic segmentation feature to obtain a second reflection feature. Specifically, performing a linear transformation on the first reflection feature to generate a query matrix Q of the first reflection feature. Performing a first cross-attention operation on the query matrix Q of the first reflection feature and the key matrix K and value matrix V of the semantic segmentation feature through a third common information module to obtain a second reflection feature, i.e., a common feature of the first reflection feature and the semantic segmentation feature. Step S44: Fusing the second reflection feature with the difference feature to obtain a reflection feature. Based on the results of the above steps, the process of step S44 is to add the query matrix Q of the second reflection feature and the query matrix Q of the first fused feature element by element to obtain the reflection feature.

[0026] It is worth noting that the first cross-attention operation method in this embodiment is the same, except that the specific values ​​of the query matrix Q, key matrix K and value matrix V are different; the following takes common features as an example to specifically introduce the process of the first cross-attention operation.

[0027] like Figure 3 As shown, determine the original image features The query matrix Q and semantic segmentation features The similarity between the key matrices K and normalize them to obtain the attention weights, where , , ; The calculation formula of attention weight is: ; Where, Representing the similarity between the query matrix Q and the key matrix K, the attention score can be obtained. Represents the scaling factor, which is the dimension of the key vector and is used to scale the dot product result to prevent the gradient from vanishing or exploding due to excessive values. The Softmax operation normalizes the attention score to a probability distribution so that the weight is between 0 and 1 and the sum is 1.

[0028] Applying attention weights to semantic segmentation features The value matrix V of , generates weighted features , that is, the initial reflection characteristic, the formula is as follows: ; Weighted features Add element-by-element to the query vector Q to obtain the residual connection feature This not only preserves the original input information, but also integrates the reflection features extracted by the cross-attention mechanism. The calculation formula is as follows: ; Normalize the residual connection features and enhance them through the multi-layer perceptron MLP to obtain enhanced features , which can realize the nonlinear expression ability of residual connection features; ; Enhance the features and residual connection features Add them together to get the common features ; .

[0029] In this embodiment, by alternately using the cross-attention mechanism, the common features of the original image features and the semantic segmentation features are determined, and the common features are fused with the original image features and the semantic segmentation features, the accuracy of reflection feature extraction can be enhanced; not only can semantic information and high-frequency features be effectively fused, but the stability and expression ability of the model are also enhanced through residual connection and normalization technology; the introduction of the multi-layer perceptron further enhances the nonlinear transformation capability of the features, making the reconstruction of the reflection map more refined and accurate.

[0030] Step S5: Determine the difference features between the original image features and the semantic segmentation features, fuse the difference features with the original image features and the semantic segmentation features, and obtain low-frequency illumination characteristics (such as brightness distribution). Specifically, the illumination characteristics are obtained through the illumination special attention module. The operation process of the illumination special attention module is as follows.

[0031] Specifically, such as Figure 4 As shown, step S51 determines the difference features between the original image features and the semantic segmentation features through the Difference Information Module (DIM). The specific process is to perform a second cross-attention operation on the query matrix Q of the original image features and the key matrix K and value matrix V of the semantic segmentation features to obtain the difference features between the original image features and the semantic segmentation features. Furthermore, the second cross-attention operation process includes: like Figure 5 As shown, the similarity between the query matrix Q of the original image feature and the key matrix K of the semantic segmentation feature is determined and normalized to obtain the attention weight; the attention weight is applied to the value matrix V of the semantic segmentation feature to generate the initial reflection feature; since the reflection feature and the illumination feature are complementary, the initial illumination feature can be determined based on the initial reflection feature extracted above; specifically, the first difference feature between the original image feature and the semantic feature is determined based on the value matrix V of the weighted reflection feature and the semantic segmentation feature , that is, the initial illumination feature, the formula is as follows: ; First difference feature Perform residual connection with the query vector Q of the original image feature to obtain the second difference feature ; ; The second difference feature Normalize LN and perform feature enhancement through multi-layer perceptron MLP to obtain the third difference feature ;

[0032] The third difference characteristic The second difference feature Add them together to get the difference features: .

[0033] Step S52: The fourth common information module performs common information fusion on the difference feature and the original image feature to obtain a first illumination feature. The specific process is to perform a first cross-attention operation on the query matrix Q of the difference feature and the key matrix K and value matrix V of the original image feature to obtain the first illumination feature, that is, the common feature of the difference feature and the original image feature. Step S53: performing common information fusion on the first common feature and the semantic segmentation feature through a fifth common information module to obtain a second illumination feature. Specifically, the query matrix Q of the first illumination feature is subjected to a first cross-attention operation with the key matrix K and value matrix V of the semantic segmentation feature to obtain the second illumination feature, i.e., the common feature of the first illumination feature and the semantic segmentation feature. Step S54: Fusing the second common feature with the difference feature to obtain the illumination feature. Based on the results of the above steps, the process of step S54 is to add the query matrix Q of the second illumination feature and the query matrix Q of the difference feature element by element to obtain the illumination feature.

[0034] In this embodiment, the illumination features are extracted by semantic segmentation features and combining the complementarity of reflection characteristics and illumination characteristics, and the extracted illumination features are fused with the original image features by alternating use of the cross-attention mechanism to obtain the final illumination features. By constraining the reconstruction of the illumination map by illumination characteristics, it can be ensured that the reflection characteristics and illumination characteristics complement each other, further improving the details and consistency of the reconstructed image.

[0035] Step S6: reconstruct the enhanced original image features and the reflection features to obtain a reflection map; reconstruct the enhanced original image features and the illumination features to obtain an illumination map.

[0036] Specifically, a decoder is used to reconstruct the enhanced original image features and reflection features, and the enhanced original image features and illumination features, respectively, to obtain a reflection map and an illumination map respectively.

[0037] An intrinsic image decomposition device based on interactive image semantic information constraints includes: a semantic feature extraction module, which is used to perform semantic segmentation on the original image to obtain semantic labels; and perform multi-scale feature extraction on the semantic labels to obtain semantic segmentation features; an encoder, which is used to extract features from the original image to obtain original image features; a feature enhancement module, which is used to perform feature enhancement on the original image features to obtain enhanced original image features; a reflection attention module, which is used to determine the common features of the original image features and the semantic segmentation features, and fuse the common features with the original image features and the semantic segmentation features to obtain reflection features; an illumination attention module, which is used to determine the difference features between the original image features and the semantic segmentation features, and fuse the difference features with the original image features and the semantic segmentation features to obtain illumination features; a decoder, which is used to reconstruct the enhanced original image features and the reflection features to obtain a reflection map; and also to reconstruct the enhanced original image features and the illumination features to obtain an illumination map.

[0038] Among them, the intrinsic image decomposition device is obtained through pre-training, and the loss function in the training process is Charbonnier loss , Perceptual Loss and CCR loss The sum of , the formula is as follows: ; in, (Pixel-level loss) was first used in super-resolution tasks to replace L1 loss and L2 loss. The formula is: ; Where, x is the tensor output by the decomposition network, is the reference value (ground-truth) paired with it.

[0039] Perceptual loss extracts high-level features based on the VGG16 network and optimizes content similarity. The formula is: ; CCR loss (contrast structure loss) can constrain the shape and material of the image and improve the robustness of the decomposed image. CCR loss first uses the CCR formula to calculate the real image and the generated image respectively, and then performs absolute value error processing, as follows: ; in, Represents the CCR value of the true label and the predicted image respectively. CCR loss is based on the color ratio of adjacent pixels and enhances the robustness of the reflection characteristics. For an RGB image, two adjacent pixels on the image are defined as and The CCR (Cross Color Ratio) value can be defined as: ; in, and Represents two adjacent pixels and The R, G, and B values ​​of the image.

[0040] Specifically, the reflection attention module includes: a first common information module, which is used to perform common information fusion on the original image features and the semantic segmentation features to obtain the common features; a second common information module, which is used to perform common information fusion on the common features and the original image features to obtain the first reflection features; a third common information module, which is used to perform common information fusion on the first reflection features and the semantic segmentation features to obtain the second reflection features; and a first fusion module, which is used to fuse the second reflection features with the difference features to obtain the reflection features.

[0041] The illumination attention module includes: a difference information module, which is used to determine the difference features between the original image features and the semantic segmentation features; a fourth common information module, which is used to perform common information fusion of the difference features and the original image features to obtain a first illumination feature; a fifth common information module, which is used to perform common information fusion of the first common feature and the semantic segmentation feature to obtain a second illumination feature; and a second fusion module, which is used to fuse the second common feature with the difference feature to obtain an illumination feature.

[0042] Example 1: Network training and testing (1) Dataset The MIT Essential Image Dataset and ShapeNet Dataset are used for training and testing. The MIT dataset contains 20 objects, each captured under 10 lighting conditions, and the ShapeNet dataset contains more than 3,000 categories.

[0043] (2) Experimental environment Table 4-1 Experimental environment configuration information

[0044] For all experiments in this example, a fixed random seed is used to initialize the parameters to maintain consistency. The network is optimized by the Adam optimizer, whose parameters are and The weight decay rate was set to 1e-8. The network training cycle was set to 250 epochs, the initial learning rate was set to 0.0002, and the batch size for each training was 8. The image resolution used for both training and testing was 256×256×3.

[0045] (3) Experimental results Figure 6 The results of this example are shown in the figure. As can be seen, the method of the present invention achieves excellent performance across multiple dimensions: color accuracy of the reflectance image, cleanliness of the illumination image, sharpness of detail, and semantic consistency. The reflectance image achieves complete lines and texture fidelity, with natural transitions between light and dark without abrupt discontinuities. It approaches GT performance closely at the edges of painted images and at color boundaries across all objects.

[0046] Table 1 shows the experimental results comparison of the essential image decomposition method based on semantic information constraints proposed in this paper and other methods on the MIT dataset.

[0047] Table 1 Experimental results on the MIT dataset

[0048] As shown in Table 1, the semantically constrained intrinsic image decomposition method proposed in this paper outperforms several other representative methods on all evaluation metrics (MSE, LMSE, and DSSIM) on the MIT dataset. The errors in all metrics are significantly reduced for both the reflectance and illumination branches, demonstrating that the proposed method effectively preserves the texture details and structural information in the intrinsic image while significantly reducing structural distortion and artifacts during the decomposition process, demonstrating strong generalization performance and visual quality.

[0049] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An intrinsic image decomposition method based on interactive image semantic information constraints, characterized in that: include: Perform semantic segmentation on the original image to obtain semantic labels; Performing multi-scale feature extraction on the semantic labels to obtain semantic segmentation features; Performing feature extraction on the original image to obtain original image features; Performing feature enhancement on the original image features to obtain enhanced original image features; Determining common features of the original image features and the semantic segmentation features, and fusing the common features with the original image features and the semantic segmentation features to obtain reflection features; Determining a difference feature between the original image feature and the semantic segmentation feature, and fusing the difference feature with the original image feature and the semantic segmentation feature to obtain an illumination feature; Reconstructing the enhanced original image features and the reflection features to obtain a reflection image; The enhanced original image features and illumination features are reconstructed to obtain an illumination map.

2. The intrinsic image decomposition method based on interactive image semantic information constraint according to claim 1, characterized in that: The determining of common features between the original image features and the semantic segmentation features, and fusing the common features with the original image features and the semantic segmentation features to obtain reflection features includes: Performing common information fusion on the original image features and the semantic segmentation features to obtain common features; Performing common information fusion on the common features and the original image features to obtain a first reflection feature; Performing common information fusion on the first reflection feature and the semantic segmentation feature to obtain a second reflection feature; The second reflection feature is fused with the difference feature to obtain a reflection feature.

3. The intrinsic image decomposition method based on interactive image semantic information constraint according to claim 2, characterized in that: The determining of common features between the original image features and the semantic segmentation features, and fusing the common features with the original image features and the semantic segmentation features to obtain reflection features includes: Perform the first cross-attention operation on the query matrix of the original image features and the key matrix and value matrix of the semantic segmentation features to obtain the common features of the original image features and the semantic segmentation features; Perform the first cross attention operation on the query matrix of the common features and the key matrix and value matrix of the original image features to obtain the first reflected feature; Perform a first cross attention operation on the query matrix of the first reflection feature and the key matrix and value matrix of the semantic segmentation feature to obtain a second reflection feature; The query matrix of the second reflection feature and the query matrix of the first fusion feature are added element by element to obtain the reflection feature.

4. The intrinsic image decomposition method based on interactive image semantic information constraint according to claim 1, characterized in that: The determining of the difference feature between the original image feature and the semantic segmentation feature, and fusing the difference feature with the original image feature and the semantic segmentation feature to obtain the illumination feature includes: Determining the difference between the original image features and the semantic segmentation features; Performing common information fusion on the difference feature and the original image feature to obtain a first illumination feature; Performing common information fusion on the first common feature and the semantic segmentation feature to obtain a second illumination feature; The second common feature is fused with the difference feature to obtain the illumination feature.

5. The intrinsic image decomposition method based on interactive image semantic information constraint according to claim 4, characterized in that: The determining of the difference feature between the original image feature and the semantic segmentation feature, and fusing the difference feature with the original image feature and the semantic segmentation feature to obtain the illumination feature includes: Perform a second cross-attention operation on the query matrix of the original image features and the key matrix and value matrix of the semantic segmentation features to obtain the difference features between the original image features and the semantic segmentation features; Perform the first cross attention operation on the query matrix of the difference feature and the key matrix and value matrix of the original image feature to obtain the first illumination feature; Perform a first cross-attention operation on the query matrix of the first illumination feature and the key matrix and value matrix of the semantic segmentation feature to obtain a second illumination feature; The query matrix of the second illumination feature is added element by element to the query matrix of the difference feature to obtain the illumination feature.

6. The intrinsic image decomposition method based on interactive image semantic information constraint according to claim 3 or 5, characterized in that: The first cross attention operation process includes: Determine the similarity between the query matrix and the key matrix and normalize it to obtain the attention weight; Applying the attention weights to the value matrix to generate weighted features; Adding the weighted features to the query vector element by element to obtain a residual connection feature; Normalizing the residual connection features and performing feature enhancement through a multi-layer perceptron to obtain enhanced features; The enhanced features and the residual connection features are added to obtain the first cross attention operation result.

7. The intrinsic image decomposition method based on interactive image semantic information constraint according to claim 5, characterized in that: The second cross-attention operation process includes: Determine the similarity between the query matrix of the original image features and the key matrix of the semantic segmentation features, and normalize them to obtain the attention weight; Applying the attention weights to the value matrix of the semantic segmentation features to generate weighted features; Determining a first difference feature between the original image feature and the semantic feature according to the value matrix of the weighted feature and the semantic segmentation feature; Performing a residual connection on the query vector of the first difference feature and the original image feature to obtain a second difference feature; Normalizing the second difference feature and performing feature enhancement through a multi-layer perceptron to obtain a third difference feature; The third difference feature is added to the second difference feature to obtain a difference feature.

8. An intrinsic image decomposition device based on interactive image semantic information constraints, characterized in that: include: Semantic feature extraction module, used to perform semantic segmentation on the original image and obtain semantic labels; and performing multi-scale feature extraction on the semantic labels to obtain semantic segmentation features; An encoder, configured to extract features from the original image to obtain original image features; A feature enhancement module, configured to enhance the features of the original image to obtain enhanced original image features; A reflection attention module is used to determine the common features of the original image features and the semantic segmentation features, and fuse the common features with the original image features and the semantic segmentation features to obtain reflection features; an illumination attention module, configured to determine a difference feature between the original image feature and the semantic segmentation feature, and fuse the difference feature with the original image feature and the semantic segmentation feature to obtain an illumination feature; A decoder, configured to reconstruct the enhanced original image features and the reflection features to obtain a reflection map; and further configured to reconstruct the enhanced original image features and the illumination features to obtain an illumination map; The intrinsic image decomposition device is obtained through pre-training, and the loss function during the training process is the sum of Charbonnier loss, perceptual loss and CCR loss.

9. The intrinsic image decomposition device based on interactive image semantic information constraint according to claim 8, characterized in that: The reflection attention module includes: A first common information module is used to perform common information fusion on the original image features and the semantic segmentation features to obtain common features; A second common information module is used to perform common information fusion on the common features and the original image features to obtain a first reflection feature; A third common information module is used to perform common information fusion on the first reflection feature and the semantic segmentation feature to obtain a second reflection feature; The first fusion module is used to fuse the second reflection feature with the difference feature to obtain a reflection feature.

10. The intrinsic image decomposition device based on interactive image semantic information constraint according to claim 8, characterized in that: The illumination attention module includes: A difference information module, used to determine the difference features between the original image features and the semantic segmentation features; a fourth common information module, configured to perform common information fusion on the difference feature and the original image feature to obtain a first illumination feature; a fifth common information module, configured to perform common information fusion on the first common feature and the semantic segmentation feature to obtain a second illumination feature; The second fusion module is used to fuse the second common feature with the difference feature to obtain the illumination feature.

Citation Information

Cited By

  • Automatic artifact detection and restoration method and system based on endoscope image

    CN121392066A

  • An automatic artifact detection and correction method and system based on endoscopic images

    CN121392066B