Image decomposition method and device, equipment, medium and program product
By employing feature extraction and cross-attention feature concatenation, the problems of low efficiency and insufficient accuracy in image decomposition in existing technologies are solved, achieving efficient and accurate decomposition of reflectance and illumination maps.
Patent Information
- Application Number
- CN202411083985.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-06
AI Technical Summary
In existing technologies, image decomposition into reflection and illumination maps requires the use of multiple neural network models, resulting in low efficiency and insufficient decomposition accuracy, and a lack of correlation between different component images during decomposition.
By extracting features from the image to be decomposed, the first reflection map features, the first illumination map features, and the first original image features are obtained. These features are then associated and concatenated using cross-attention features to obtain the second reflection map features and the second illumination map features, which are then used to decompose the image.
It improves the efficiency and accuracy of image decomposition, enhances the correlation between features of images with different components, simplifies the model structure, and improves decomposition performance.
Smart Images

Figure CN121481835A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and particularly relates to an image decomposition method and device, equipment, medium and program product. BACKGROUND
[0002] Eigen image decomposition is a basic task in computer vision, which usually refers to decomposing an image into different components, such as decomposing into a reflection map and an illumination map, which helps to better analyze the image. In the related art, eigen image decomposition can be implemented based on a neural network model. Usually, a neural network model is trained for a reflection map or an illumination map, respectively, to decompose an original image into a reflection map or an illumination map. However, this method needs to use multiple neural network models, which is relatively complex and inefficient. SUMMARY
[0003] The present application provides an image decomposition method and device, equipment, medium and program product to improve the efficiency and accuracy of image decomposition.
[0004] In a first aspect, the present application provides an image decomposition method, comprising:
[0005] performing feature extraction on a to-be-decomposed image to obtain a first reflection map feature, a first illumination map feature and a first original image feature of the to-be-decomposed image;
[0006] determining a second reflection map feature according to a first cross-attention feature between the first reflection map feature and the first illumination map feature and the first reflection map feature, and determining a second illumination map feature according to the first cross-attention feature and the first illumination map feature;
[0007] splicing the second reflection map feature, the second illumination map feature and the first original image feature to obtain a first aggregation feature;
[0008] decomposing the to-be-decomposed image according to the second reflection map feature, the second illumination map feature and the first aggregation feature to obtain a reflection map and an illumination map of the to-be-decomposed image.
[0009] In a second aspect, the present application provides an image decomposition device, comprising:
[0010] an encoding module configured to perform feature extraction on a to-be-decomposed image to obtain a first reflection map feature, a first illumination map feature and a first original image feature of the to-be-decomposed image;
[0011] The feature extraction module is used to determine a second reflection map feature based on a first cross-attention feature between the first reflection map feature and the first illumination map feature, and the first reflection map feature; and to determine a second illumination map feature based on the first cross-attention feature and the first illumination map feature.
[0012] The feature extraction module is further configured to concatenate the second reflection map feature, the second illumination map feature, and the first original image feature to obtain a first aggregated feature;
[0013] The decomposition module is used to decompose the image to be decomposed based on the second reflection map feature, the second illumination map feature and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed.
[0014] Thirdly, this application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the above-described image decomposition method.
[0015] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described image decomposition method.
[0016] Fifthly, this application provides a computer program product including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above-described image decomposition method.
[0017] In this embodiment, feature extraction can be performed on the image to be decomposed to obtain the first reflection map feature, the first illumination map feature, and the first original image feature. Then, the features of the image to be decomposed are analyzed, and the second reflection map feature and the second illumination map feature are obtained based on the first cross-attention feature between the first reflection map feature and the first illumination map feature. These features are then concatenated with the first original image feature. In this way, through features cross-attention and concatenation aggregation, the correlation of feature representations of different component images is enhanced. Then, based on the concatenated first aggregated feature, the second reflection map feature, and the second illumination map feature, the image to be decomposed is decomposed, which can simultaneously decompose the reflection map and the illumination map of the image to be decomposed, thus improving the accuracy of image decomposition.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the present application and form part of the specification. They are used together with the embodiments of the present application to explain the application and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed example embodiments described with reference to the accompanying drawings, in which:
[0020] Figure 1 A flowchart of an image decomposition method provided in an embodiment of this application;
[0021] Figure 2 This is a schematic diagram of the architecture of the image decomposition model in an embodiment of this application;
[0022] Figure 3 This is a flowchart of the training method for the image decomposition model in the embodiments of this application;
[0023] Figure 4 A block diagram of an image decomposition apparatus provided in an embodiment of this application;
[0024] Figure 5 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] To enable those skilled in the art to better understand the technical solutions of this application, exemplary embodiments of this application are described below in conjunction with the accompanying drawings, including various details of the embodiments of this application to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Where there is no conflict, the various embodiments of this application and the features thereof may be combined with each other.
[0027] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0028] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0029] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0030] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in this application comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example, appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely identifying specific individuals.
[0031] To facilitate understanding of the embodiments of this application, several concepts will be briefly introduced below:
[0032] Intrinsic image: refers to the decomposition of an image into different components, such as a reflection image and an illumination image. The introduction of the concept of intrinsic image is beneficial for better image analysis and processing. In addition to decomposing into reflection images and illumination images, intrinsic image decomposition can also be decomposed into other component images based on different decomposition methods. The image decomposition method in the embodiments of this application is mainly aimed at the scenario of decomposing into reflection images and illumination images. Of course, it can also be extended to be applicable to the decomposition of other component images, and there is no limitation thereto.
[0033] Albedo image: Albedo image features primarily reflect the color and material information of the original image (i.e., the image to be decomposed). In the albedo image, the boundaries between highlight and non-highlight areas are eliminated, while the boundaries between objects still exist. This is because the albedo image removes the influence of lighting, making it easier to segment different targets. This processing helps to better analyze and understand objects in the image, especially when it is necessary to accurately identify and segment different objects in the image.
[0034] Shading image: Also known as a specular map, a shading image represents the illumination of an image. It includes information such as the color and intensity distribution of the light source, as well as the shadows and shading information on the object's structure. The relationship between the original image and its corresponding decomposed reflection and shading images is: I = R * S, where I represents the original image, R represents the reflection image, and S represents the shading image.
[0035] In related technologies, intrinsic image decomposition can be implemented based on neural network models. Typically, a separate neural network model is trained for each of the reflection or illumination maps to decompose the original image into reflection or illumination maps. However, this method requires the use of multiple neural network models, which is complex and inefficient. Furthermore, in related technologies, the decomposition of different component images is independent and unrelated, making it difficult to guarantee that the reflection and illumination maps generated based on different neural network models can be aggregated into the original image. This may lead to incorrect aggregation results and reduce the accuracy of the decomposition.
[0036] Therefore, this application provides an image decomposition method. For an image to be decomposed, a first reflection map feature, a first illumination map feature, and a first original image feature are obtained. Based on a first cross-attention feature between the first reflection map feature and the first illumination map feature, a second reflection map feature and a second illumination map feature are obtained. The second reflection map feature, the second illumination map feature, and the first original image feature are then concatenated to obtain a first aggregated feature. Subsequently, based on the second reflection map feature, the second illumination map feature, and the first aggregated feature, the image to be decomposed is decomposed to obtain the reflection map and illumination map of the image to be decomposed. In this way, the decomposed reflection map and illumination map can be obtained simultaneously by processing multiple features corresponding to the image to be decomposed, without needing to use two separate models, thus improving efficiency and simplifying implementation. Furthermore, the processing of the cross-attention feature enhances the feature representations of the reflection map and the illumination map, improves the correlation between feature representations, and thus improves the accuracy of image decomposition.
[0037] To facilitate understanding of the embodiments of this application, a detailed description of the image decomposition method disclosed in this application is provided first. The execution subject of the image decomposition method provided in this application is generally an electronic device with a certain computing power. This electronic device may include, for example, a terminal device, a server, or other processing device. The terminal device may be an in-vehicle device, user equipment (UE), mobile device, personal digital assistant (PDA), handheld device, computing device, wearable device, etc. The server may be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. In some possible implementations, the image decomposition method can be implemented by a processor calling computer-readable instructions stored in memory.
[0038] The image decomposition method provided in this application embodiment will be described below using a server as the execution subject as an example.
[0039] Figure 1 A flowchart of an image decomposition method provided in this application embodiment is shown below. Figure 1 The method includes:
[0040] S101: Extract features from the image to be decomposed to obtain the first reflection map features, the first illumination map features, and the first original image features of the image to be decomposed.
[0041] Three different encoders—an albedo encoder, a shading encoder, and an image encoder—can be used to extract features from the image to be decomposed, yielding the first albedo feature, the first shading feature, and the first image feature. The encoder takes the image to be decomposed as input and outputs a feature vector. This feature vector represents the image features captured by the encoder. The specific image features the encoder captures can be defined (e.g., semantic features). The output feature vector is a high-dimensional feature vector; that is, the first albedo feature, the first shading feature, and the first image feature are three feature vectors in a high-dimensional space.
[0042] For example, based on the reflectance map feature encoder, features are extracted from the image to be decomposed to obtain the feature vector of the first reflectance map feature of the image to be decomposed; based on the illumination map feature encoder, features are extracted from the image to be decomposed to obtain the feature vector of the first illumination map feature of the image to be decomposed; based on the original image feature encoder, features are extracted from the image to be decomposed to obtain the feature vector of the first original image feature of the image to be decomposed. This application embodiment does not impose limitations on the aforementioned reflectance map feature encoder, illumination map feature encoder, and original image feature encoder; existing mature models can be used and trained to obtain them.
[0043] In the embodiments of this application, the first reflection map feature is used to represent the image features of the reflection map of the image to be decomposed, the first illumination map feature is used to represent the image features of the illumination map of the image to be decomposed, and the first original image feature is used to represent the image features of the image to be decomposed. The image features may include semantic features, texture features, color features, shape features, illumination features, etc., of the image, without any limitation.
[0044] Specifically, the semantic features of an image refer to the information that can express the high-level meaning and content of the image. They are usually used to understand and interpret the scenes, objects and their relationships contained in the image. Semantic features focus more on the actual content and meaning of the image.
[0045] Image color features are global features that describe the color information of objects in an image or image region. They are based on pixel-level features and are insensitive to changes in the orientation, size, or other characteristics of the image or image region.
[0046] Texture features of an image are also a type of global feature, describing the texture information of the scene corresponding to the image or image region. Unlike color features, texture features require statistical calculations in a region containing multiple pixels. When the image resolution changes, the texture is affected by lighting and reflection.
[0047] Image shape features refer to the information extracted from an image that describes the shape and geometric structure of an object. They can reflect information such as the image's edges, contours, shape descriptions, shape context, corners, and shape skeleton.
[0048] The lighting features of an image describe the distribution of light and the characteristics of light interaction with objects in the image. Lighting has a great influence on the visual effect of an image. Lighting features include information such as brightness, contrast, lighting direction, lighting intensity, shadows, and highlights.
[0049] In summary, the first reflection map feature, the first illumination map feature, and the first original image feature are obtained based on the image to be decomposed and different encoders. The first original image feature focuses on reflecting the content and structure of the image to be decomposed; the first reflection map feature may vary depending on the reflection process and focuses on reflecting the shape information of objects in the image to be decomposed; the first illumination map feature focuses more on the interaction between the image to be decomposed and the external environment (such as lighting conditions, projection errors, etc.). The first reflection map feature comes from the reflection map after decomposition of the image to be decomposed, the first illumination map feature comes from the illumination map after decomposition of the image to be decomposed, and the first original image feature comes from the image to be decomposed. There is a correlation among the three. By learning and strengthening the correlation among the three, it is beneficial to obtain more accurate reflection maps and illumination maps based on the decomposition of the image to be decomposed.
[0050] S102: Determine the second reflection map feature based on the first cross-attention feature between the first reflection map feature and the first illumination map feature, and the first reflection map feature; determine the second illumination map feature based on the first cross-attention feature and the first illumination map feature.
[0051] In practice, an image and its decomposed reflection map and illumination map have a certain correlation. Therefore, in this embodiment of the application, in order to improve the decomposition accuracy, a first cross-attention feature between the first reflection map feature and the first illumination map feature can be mined and obtained. Then, different image components can be associated through the first cross-attention feature to obtain the second reflection map feature and the second illumination map feature.
[0052] S103: The second reflection map feature, the second illumination map feature, and the first original map feature are concatenated to obtain the first aggregated feature.
[0053] In this embodiment, the second reflection map features and the second illumination map features after enhanced correlation, as well as the first original image features, are stitched together, which also enhances the correlation between image features of each dimension. As a result, the accuracy of image decomposition can be further improved due to the improved accuracy of feature representation.
[0054] S104: Based on the second reflection map features, the second illumination map features, and the first aggregation features, the image to be decomposed is decomposed to obtain the reflection map and illumination map of the image to be decomposed.
[0055] Thus, in this embodiment, for the initially extracted first reflection map features, first illumination map features, and first original image features corresponding to the image to be decomposed, the association is strengthened through the first cross-attention feature to obtain the second reflection map features and the second illumination map features. This enhances the correlation between the feature representations corresponding to the illumination map and the reflection map, improving the accuracy and effectiveness of the feature representation. Furthermore, the second reflection map features, the second illumination map features, and the first original image features are concatenated to obtain the first aggregated feature. Then, based on the second reflection map features, the second illumination map features, and the first aggregated feature, the image to be decomposed is decomposed to obtain the reflection map and the illumination map. This allows for the acquisition of reflection maps and illumination maps of different component dimensions based on more accurate and more correlated feature representations of different component dimensions, improving the accuracy of image decomposition. Moreover, it allows for the decomposition of reflection maps and illumination maps of different components based on feature representations of different component dimensions, improving efficiency and decomposition performance.
[0056] The image decomposition method in this application embodiment can be implemented using a pre-trained image decomposition model. Specifically, this application provides a possible architecture for an image decomposition model, see [link / reference]. Figure 2 The diagram shown is a schematic representation of the image decomposition model in an embodiment of this application.
[0057] like Figure 2 As shown, the image decomposition model in this embodiment includes a feature extraction module and a decomposition module, specifically:
[0058] 1) The feature extraction module can be connected to three different encoders, namely the original image feature encoder E. image E, a reflection map feature encoder albedo and illumination map feature encoder E shading The feature extraction module is mainly used to enhance and aggregate the first reflection map features, the first illumination map features, and the first original image features output by three different encoders to obtain three different constraints for the decomposition module. The constraints can represent the basis for model processing. For example, in the embodiments of this application, the constraints can be used to control the noise addition process and the noise reduction process of the decomposition module to accurately obtain the decomposed reflection map and illumination map.
[0059] like Figure 2 As shown, the feature extraction module further includes an attention layer and a connection layer. The attention layer is used to determine the second reflection map feature based on the first cross-attention feature between the first reflection map feature and the first illumination map feature, as well as the first reflection map feature; and to determine the second illumination map feature based on the first cross-attention feature and the first illumination map feature.
[0060] The connecting layer is used to stitch together the second reflection map features, the second illumination map features, and the first original map features to obtain the first aggregated feature.
[0061] The three constraints ultimately used for the decomposition module can be the second reflection map feature, the second illumination map feature, and the first aggregation feature.
[0062] 2) The decomposition module is mainly used to decompose the reflection map and illumination map of different components based on the three different constraints output by the feature extraction module. Specifically, it is used to decompose the image to be decomposed according to the second reflection map feature, the second illumination map feature and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed.
[0063] The decomposition module can be implemented using a diffusion model, which offers good controllability, interpretability, and reversibility. The diffusion model mainly includes a diffusion process and a reverse diffusion process. The diffusion process can be understood as a noise-adding process, where Gaussian noise is gradually added to corrupt the image until it becomes Gaussian white noise. The reverse diffusion process can be understood as a noise-removing process, where noise is gradually removed from Gaussian white noise until a clean, noise-free image is obtained. During training, noise of any level can be added to the training image, and the neural network is optimized to predict this noise. Through training, the neural network learns to predict noise in an image from a noisy one, thus acquiring the ability to remove noise. Of course, the decomposition module in this embodiment is not limited to a diffusion model and can be implemented using other models; no specific limitations are imposed.
[0064] Thus, this application embodiment provides an architecture for an image decomposition model including a decomposition module and a feature extraction module. The feature extraction model enhances the feature representations of images with different components, improving the correlation and accuracy of these representations. The decomposition module uses the enhanced feature representations as constraints to decompose images with different components under different constraints, i.e., decomposing reflection and illumination maps. For example, the decomposition module can be a diffusion model. In this case, the algorithm architecture provided in this application embodiment can effectively integrate the diffusion model into image decomposition. Obtaining reflection and illumination maps based on the diffusion model effectively utilizes the correlation between images with different components, increases controllability, and improves image decomposition performance.
[0065] based on Figure 2 Regarding the image decomposition model architecture shown, and for steps S102-S104 above, this application also provides possible implementation methods, specifically:
[0066] like Figure 2As shown, the feature extraction module in this embodiment includes an attention layer and a concat layer. The attention layer can adopt a cross-attention mechanism, etc., and this embodiment does not impose any restrictions.
[0067] In this embodiment of the application, regarding the determination method of the first cross-attention feature in step S102 above, a possible implementation method is also provided, including:
[0068] 1) Perform a linear transformation on the first reflection map features to determine the query features of the first reflection map features.
[0069] 2) Perform a linear transformation on the features of the first illumination map to determine the key features and value features of the first illumination map.
[0070] 3) Multiply and activate the query features of the first reflection map feature and the key features of the first illumination map feature to determine the cross-attention weight between the first reflection map feature and the first illumination map feature.
[0071] 4) Determine the first cross-attention feature based on the cross-attention weights and the value features of the first illumination map feature.
[0072] For example, the first cross-attention features are obtained through an attention layer. For the image to be decomposed X0, the reflection map feature encoder E... albedo Obtain the first reflection map feature E albedo (X0), via the illumination map feature encoder E shading Obtain the first illumination map feature E shading (X0), which is then input into the image decomposition model. Through the attention layer of the feature extraction model in the image decomposition model, the first cross-attention feature Attention(E) between the first illumination map feature and the first reflection map feature is obtained. albedo (X0),E shading (X0)), where Attention() represents the attention layer method. Of course, in this embodiment, there are no restrictions on the method of obtaining the first cross-attention feature.
[0073] Therefore, regarding step S102 above, a second reflection map feature can be obtained based on the first cross-attention feature and the first reflection map feature. For example, the second reflection map feature is F. albedo , specifically:
[0074] F albedo =E albedo (X0)+Attention(E albedo (X0),E shading (X0))
[0075] Based on the first cross-attention feature and the first illumination map feature, the second illumination map feature is obtained. For example, the second illumination map feature is F. shading , specifically:
[0076] F shading =E shading (X0)+Attention(E albedo (X0),E shading (X0))
[0077] In one possible implementation, step S103 above, which involves stitching together the second reflection map feature, the second illumination map feature, and the first original image feature to obtain a first aggregated feature, includes: stitching together the second reflection map feature, the second illumination map feature, and the first original image feature through a connecting layer to obtain a first aggregated feature.
[0078] For example, for an image X0 to be decomposed, the original image feature encoder E... image Obtain the first original image feature F image =E image (X0), and then input into the image decomposition model, based on the connection layer in the feature extraction module, the second reflection map feature, the second illumination map feature, and the first original image feature are concatenated to obtain the first aggregated feature F. union Then F union =concat(F image ,F albedo ,F shading ), where concat() represents the connection layer method.
[0079] Furthermore, such as Figure 2 As shown, the feature extraction module in this embodiment may further include a random noise coding layer. For example, the random noise coding layer can be implemented using an unconditional denoising diffusion implicit model, without limitation. This random noise coding layer is mainly used to add random noise, which can increase randomness during model training. For example, it can sample noise from a Gaussian distribution. This random noise coding layer can be used or not, and there is no limitation in this embodiment. For example, based on the random noise coding layer, the first aggregated feature can also be represented as:
[0080] F union =concat(F image ,F albedo ,F shading ,α*noise)
[0081] Here, α is a control parameter used to control whether random noise is used. Noise represents the random noise output by the random noise coding layer. When α = 0, it means that no random noise is added, and when α = 1, it means that random noise is added.
[0082] Thus, in this embodiment of the application, the first reflection map features, the first illumination map features, and the first original image features are associated and aggregated through various methods to obtain the second reflection map features, the second illumination map features, and the first aggregated features. Through association and aggregation processing, the feature representation can be enhanced, thereby improving the accuracy of subsequent image decomposition.
[0083] In one possible implementation, regarding step S104 above, which involves decomposing the image to be decomposed based on the second reflection map features, the second illumination map features, and the first aggregation features to obtain the reflection map and illumination map of the image to be decomposed, this application also provides a possible implementation, including:
[0084] 1) Using the second reflection map feature, the second illumination map feature and the first aggregation feature as constraints, noise is added to the image to be decomposed according to the preset noise parameters to obtain a noisy image.
[0085] For example, step S104 above is implemented through a decomposition module, and the second reflection map feature F albedo The corresponding constraint is represented as C albedo Second illumination map feature F shading The corresponding constraint is represented as C shading The first aggregation feature F union The corresponding constraint is represented as C union Let the image to be decomposed be X0, and the decomposition module be, for example, a conditional denoising diffusion implicit model, denoised as D. Through the decomposition module, during the noise addition process, the number of sampling steps is, for example, T. Then, under three different constraints, the image to be decomposed yields a noisy image X. T X T =D forward (X0,(C union C albedo C shading ), where D forward This indicates the addition of noise to the decomposition module.
[0086] In this process, through noise processing by the decomposition module, the image to be decomposed can obtain three corresponding noise images under three different constraints. These three noise images are then averaged or weighted to obtain the final noise image X. T .
[0087] Specifically, this application provides a possible implementation method, which uses a second reflection map feature, a second illumination map feature, and a first aggregation feature as constraints, and adds noise to the image to be decomposed according to preset noise parameters to obtain a noisy image, including: adding noise to the image to be decomposed according to the second reflection map feature as a constraint and according to noise parameters to obtain a first noisy image; adding noise to the image to be decomposed according to the second illumination map feature as a constraint and according to noise parameters to obtain a second noisy image; adding noise to the image to be decomposed according to the first aggregation feature as a constraint and according to noise parameters to obtain a third noisy image; and determining a noise image based on the first noise image, the second noise image, and the third noise image.
[0088] 2) Using the second reflection map feature, the second illumination map feature, and the first aggregation feature as constraints, the noisy image is denoised to obtain the reflection map and illumination map of the image to be decomposed.
[0089] The denoising process can be understood as the inverse of the noise-adding process in the decomposition module. Through this inverse process, the corresponding image can be recovered from the noisy image based on different constraints. For example, the original image recovered by the decomposition module (i.e., the input image to be decomposed) is represented as X. image The recovered reflection map is represented as X. albedo The recovered illumination map is represented as X. shading The noise reduction process can then be expressed as:
[0090] X image ,X albedo ,X shading =D reverse (X T ,(C union C albedo C shading ))
[0091] Among them, D reverse The noise reduction process of the decomposition module, that is, in the embodiments of this application, the reflection map and illumination map corresponding to the image to be decomposed can be obtained by the noise addition and noise reduction processes of the decomposition module under three different constraints.
[0092] Thus, in this embodiment of the application, the decomposition tasks of reflection map and illumination map can be effectively combined, and the decomposition of images with different components can be achieved based on a single model, thereby improving the performance and efficiency of image decomposition.
[0093] Furthermore, based on the image decomposition method in this application, after obtaining the reflection map and illumination map of the image to be decomposed, it can be applied to a variety of different scenarios. Specifically, this application provides several possible embodiments:
[0094] In one possible embodiment, a first editing operation is performed on the reflection map to obtain an updated reflection map, and a second editing operation is performed on the illumination map to obtain an updated illumination map; a target image is generated based on the updated reflection map and the updated illumination map.
[0095] In this embodiment, the decomposed reflection map and illumination map can also be edited in local areas, such as recoloring, relighting, and retexturing. Then they can be synthesized to generate the edited target image. This allows for editing and modifying different components of the image, expanding or enriching the image content. For example, it can be applied to advertising content scenarios, without any limitations.
[0096] In this embodiment, editing operations can also be performed only on the reflection map or the illumination map to allow for more flexible editing and modification of the image.
[0097] In one possible embodiment, the target object is detected based on the reflection map, and the target object is segmented from the image to be decomposed or the reflection map based on the detection result.
[0098] In this embodiment of the application, the image to be decomposed is decomposed to obtain a reflection map and an illumination map. Since the illumination effect can be removed from the reflection map, it is easier to identify and segment different target objects based on the reflection map, such as license plate recognition extraction scene, surveillance scene, etc., without limitation.
[0099] In this way, intrinsic image decomposition can be applied to various computer vision tasks, improving the performance and accuracy of computer vision tasks.
[0100] Of course, the embodiments of this application are not limited to the above-mentioned possible application scenarios, but can also be applied to scenarios such as 3D reconstruction, virtual reality, and augmented reality. The specific embodiments of this application are not limited.
[0101] Based on the above embodiments, the training process of the image decomposition model in this application embodiment will be briefly described below. (See reference...) Figure 3 The diagram shown is a flowchart of the training method for the image decomposition model in this embodiment of the application, including:
[0102] S301: Extract features from the original sample image to obtain the third reflection map features, the third illumination map features, and the second original image features of the original sample image.
[0103] For example, such as Figure 2 The architecture shown can be implemented using a reflection map feature encoder E. albedo Illumination map feature encoder E shading Original Image Feature Encoder E imageFeature extraction is performed on the original image of the sample to obtain the features of the third reflection image, the third illumination image, and the second original image.
[0104] S302: Determine the fourth reflection map feature based on the second cross-attention feature between the third reflection map feature and the third illumination map feature, and the third reflection map feature; determine the fourth illumination map feature based on the second cross-attention feature and the third illumination map feature.
[0105] In this embodiment of the application, during the training process, a second cross-attention feature between the third reflection map feature and the third illumination map feature can be obtained through the attention layer. Then, based on the second cross-attention feature and the third reflection map feature, a fourth reflection map feature is obtained, and based on the second cross-attention feature and the third illumination map feature, a fourth illumination map feature is obtained.
[0106] S303: The fourth reflection map feature, the fourth illumination map feature, and the second original map feature are spliced together to obtain the second aggregated feature.
[0107] In this embodiment of the application, the fourth reflection map feature, the fourth illumination map feature, and the second original map feature are spliced together by a splicing layer to obtain the second aggregated feature.
[0108] S304: Based on the fourth reflectance map features, the fourth illumination map features, and the second aggregation features, decompose the original sample image to obtain the predicted original image, the predicted reflectance map, and the predicted illumination map of the original sample image.
[0109] Regarding this step, this application provides a possible implementation method, including:
[0110] 1) Based on the features of the fourth reflection map, the features of the fourth illumination map, and the second aggregation features, the original sample image is denoised by the decomposition module to obtain the final predicted noisy image.
[0111] Specifically, this application provides possible implementation methods, including: obtaining a predicted reflection noise image based on the fourth reflection map features and the original sample image through a decomposition module, obtaining a predicted illumination noise image based on the fourth illumination map features and the original sample image, and obtaining a predicted original noise image based on the second aggregation features and the original sample image; and obtaining a final predicted noise image based on the predicted reflection noise image, the predicted illumination noise image, and the predicted original noise image.
[0112] In this embodiment of the application, the decomposition module can perform noise addition processing under different constraints to obtain predicted noise images with different components, namely, predicted reflection noise image, predicted illumination noise image and predicted original noise image, and then obtain the final predicted noise image, for example, by averaging.
[0113] 2) Based on the final predicted noisy image, as well as the features of the fourth reflection map, the features of the fourth illumination map, and the second aggregation feature, the final predicted noisy image is denoised by the decomposition module to obtain the original predicted image, the predicted reflection map, and the predicted illumination map.
[0114] S305: Train the image decomposition model based on the predicted original image, predicted reflectance map, predicted illumination map, and the original sample image.
[0115] Specifically, based on the predicted original image, predicted reflectance map, predicted illumination map, and the original sample image, the image decomposition model is trained until the preset number of iterations or the target loss function converges.
[0116] The target loss function includes a first loss function based on the predicted reflection noise image and target noise, the predicted illumination noise image and target noise, and the predicted original noise image and target noise, as well as a second loss function based on the product of the predicted reflection image and the predicted illumination image and the original sample image.
[0117] For example, the original image of the sample is X * The noise prediction network used for noise addition processing in the decomposition module is ∈ θ The constraint condition corresponding to the fourth reflection map feature is represented as C1. albedo The constraint condition for the fourth illumination pattern feature is expressed as C1. shading The constraint condition corresponding to the second aggregation feature is represented as C1. union The predicted reflectance map is represented as X1 albedo The predicted illumination map is represented as X1 shading Then the first loss function can be expressed as L1:
[0118]
[0119] The second loss function can be expressed as L2, L2 = SSIM(X) * -X1 albedo *X1 shading )
[0120] The target loss function can be expressed as Loss, Loss = L1 + L2
[0121] Where t represents the number of sampling steps, ∈ tThe target noise is represented by the ground truth noise. Through a noise prediction network, predicted reflection noise images, predicted illumination noise images, and predicted original noise images can be obtained under different constraints. Then, the target noise is combined to obtain the first loss function, which characterizes the noise prediction loss function. The Structure Similarity Index Measure (SSIM) is used to measure the similarity between images. Since the decomposed reflection and illumination images have a certain correlation with the original image, when training the image decomposition model, the target loss function is obtained based on the first loss function that characterizes the noise prediction loss and the second loss function that characterizes the correlation between the reflection, illumination, and original images. Then, the image decomposition model is trained based on the target loss function, which can improve the training accuracy and effectiveness.
[0122] In this embodiment, a novel image decomposition model architecture is provided, and the image decomposition model is trained so that it can be used to simultaneously decompose reflection and illumination maps, achieving a combination of reflection and illumination map decomposition without the need to train multiple models separately, which is simpler and more efficient, improving image decomposition efficiency. Furthermore, based on the feature extraction module in the image decomposition model, the feature representations of different component images can be correlated and aggregated, and then used as constraints to decompose images of different components, thus improving accuracy.
[0123] It is understood that the various method embodiments mentioned above in this application can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this application will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0124] In addition, this application also provides an image decomposition apparatus, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the image decomposition methods provided in this application. The corresponding technical solutions and descriptions are described in the corresponding section of the method and will not be repeated here.
[0125] Figure 4 This is a block diagram of an image decomposition apparatus provided in an embodiment of this application.
[0126] Reference Figure 4 This application provides an image decomposition apparatus, which includes:
[0127] Encoding module 41 is used to extract features from the image to be decomposed, and obtain the first reflection map feature, the first illumination map feature and the first original image feature of the image to be decomposed;
[0128] The feature extraction module 42 is used to determine a second reflection map feature based on a first cross-attention feature between the first reflection map feature and the first illumination map feature, and the first reflection map feature; and to determine a second illumination map feature based on the first cross-attention feature and the first illumination map feature.
[0129] The feature extraction module 42 is further configured to concatenate the second reflection map feature, the second illumination map feature and the first original image feature to obtain a first aggregated feature;
[0130] The decomposition module 43 is used to decompose the image to be decomposed according to the second reflection map feature, the second illumination map feature and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed.
[0131] In some possible embodiments, the feature extraction module 42 is further configured to:
[0132] Perform a linear transformation on the first reflection map feature to determine the query feature of the first reflection map feature;
[0133] A linear transformation is performed on the first illumination map features to determine the key features and value features of the first illumination map features;
[0134] The query features of the first reflection map feature and the key features of the first illumination map feature are multiplied and activated to determine the cross-attention weight between the first reflection map feature and the first illumination map feature.
[0135] The first cross-attention feature is determined based on the cross-attention weights and the value features of the first illumination map feature.
[0136] In some possible embodiments, when decomposing the image to be decomposed based on the second reflection map feature, the second illumination map feature, and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed, the decomposition module 43 is used to:
[0137] Using the second reflection map feature, the second illumination map feature, and the first aggregation feature as constraints, noise is added to the image to be decomposed according to preset noise parameters to obtain a noisy image;
[0138] Using the second reflection map feature, the second illumination map feature, and the first aggregation feature as constraints, the noisy image is denoised to obtain the reflection map and illumination map of the image to be decomposed.
[0139] In some possible embodiments, when the image to be decomposed is denoised according to preset noise parameters based on the second reflection map feature, the second illumination map feature, and the first aggregation feature as constraints to obtain a noisy image, the decomposition module 43 is used to:
[0140] Using the second reflection map features as constraints, the image to be decomposed is denoised according to the noise parameters to obtain a first noisy image;
[0141] Using the second illumination map features as constraints, noise is added to the image to be decomposed according to the noise parameters to obtain a second noisy image;
[0142] Using the first aggregation feature as a constraint, the image to be decomposed is denoised according to the noise parameters to obtain a third noisy image;
[0143] The noise image is determined based on the first noise image, the second noise image, and the third noise image.
[0144] In some possible embodiments, after obtaining the reflection map and illumination map of the image to be decomposed, the method further includes an editing module 44 for:
[0145] A first editing operation is performed on the reflection map to obtain an updated reflection map, and a second editing operation is performed on the illumination map to obtain an updated illumination map;
[0146] Generate a target image based on the updated reflectance map and the updated illumination map.
[0147] In some possible embodiments, the image decomposition model includes an attention layer, a connection layer, and a decomposition module, wherein:
[0148] The attention layer is used to determine a second reflection map feature based on a first cross-attention feature between the first reflection map feature and the first illumination map feature, and the first reflection map feature; and to determine a second illumination map feature based on the first cross-attention feature and the first illumination map feature.
[0149] The connecting layer is used to stitch together the second reflection map feature, the second illumination map feature, and the first original image feature to obtain a first aggregated feature;
[0150] The decomposition module is used to decompose the image to be decomposed based on the second reflection map feature, the second illumination map feature, and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed.
[0151] In some possible embodiments, a training module 45 is also included, for:
[0152] Feature extraction is performed on the original sample image to obtain the third reflection map feature, the third illumination map feature, and the second original image feature of the original sample image;
[0153] A fourth reflection map feature is determined based on a second cross-attention feature between the third reflection map feature and the third illumination map feature, and the third reflection map feature; a fourth illumination map feature is determined based on the second cross-attention feature and the third illumination map feature.
[0154] The fourth reflection map feature, the fourth illumination map feature, and the second original image feature are spliced together to obtain the second aggregated feature;
[0155] Based on the fourth reflectance map feature, the fourth illumination map feature, and the second aggregation feature, the original sample image is decomposed to obtain the predicted original image, the predicted reflectance map, and the predicted illumination map of the original sample image;
[0156] The image decomposition model is trained based on the predicted original image, the predicted reflectance map, the predicted illumination map, and the sample original image.
[0157] Figure 5 This is a block diagram of an electronic device provided in an embodiment of this application.
[0158] Reference Figure 5 This application provides an electronic device, which includes: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 to enable the at least one processor 501 to perform the above-described image decomposition method.
[0159] This application also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the image decomposition method described above. The computer-readable storage medium may be volatile or non-volatile.
[0160] This application also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described image decomposition method.
[0161] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0162] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0163] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0164] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.
[0165] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0166] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0167] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0168] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0170] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for general illustrative purposes only and should not be construed as limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this application as set forth by the appended claims.
Claims
1. An image decomposition method, characterized in that, include: Feature extraction is performed on the image to be decomposed to obtain the first reflection map feature, the first illumination map feature, and the first original image feature of the image to be decomposed. A second reflection map feature is determined based on a first cross-attention feature between the first reflection map feature and the first illumination map feature, and the first reflection map feature; a second illumination map feature is determined based on the first cross-attention feature and the first illumination map feature. The second reflection map feature, the second illumination map feature, and the first original image feature are concatenated to obtain the first aggregated feature; The image to be decomposed is decomposed based on the second reflection map feature, the second illumination map feature, and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed.
2. The image decomposition method according to claim 1, characterized in that, The method further includes: Perform a linear transformation on the first reflection map feature to determine the query feature of the first reflection map feature; A linear transformation is performed on the first illumination map features to determine the key features and value features of the first illumination map features; The query features of the first reflection map feature and the key features of the first illumination map feature are multiplied and activated to determine the cross-attention weight between the first reflection map feature and the first illumination map feature. The first cross-attention feature is determined based on the cross-attention weights and the value features of the first illumination map feature.
3. The image decomposition method according to claim 1, characterized in that, The step of decomposing the image to be decomposed based on the second reflection map feature, the second illumination map feature, and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed includes: Using the second reflection map feature, the second illumination map feature, and the first aggregation feature as constraints, noise is added to the image to be decomposed according to preset noise parameters to obtain a noisy image; Using the second reflection map feature, the second illumination map feature, and the first aggregation feature as constraints, the noisy image is denoised to obtain the reflection map and illumination map of the image to be decomposed.
4. The image decomposition method according to claim 3, characterized in that, The step of adding noise to the image to be decomposed based on preset noise parameters, using the second reflection map feature, the second illumination map feature, and the first aggregation feature as constraints, to obtain a noisy image, includes: Using the second reflection map features as constraints, the image to be decomposed is denoised according to the noise parameters to obtain a first noisy image; Using the second illumination map features as constraints, noise is added to the image to be decomposed according to the noise parameters to obtain a second noisy image; Using the first aggregation feature as a constraint, the image to be decomposed is denoised according to the noise parameters to obtain a third noisy image; The noise image is determined based on the first noise image, the second noise image, and the third noise image.
5. The image decomposition method according to claim 1, characterized in that, After obtaining the reflection map and illumination map of the image to be decomposed, the method further includes: A first editing operation is performed on the reflection map to obtain an updated reflection map, and a second editing operation is performed on the illumination map to obtain an updated illumination map; Generate a target image based on the updated reflectance map and the updated illumination map.
6. The image decomposition method according to claim 1, characterized in that, The method is implemented using an image decomposition model, which includes an attention layer, a connection layer, and a decomposition module, wherein: The attention layer is used to determine a second reflection map feature based on a first cross-attention feature between the first reflection map feature and the first illumination map feature, and the first reflection map feature; and to determine a second illumination map feature based on the first cross-attention feature and the first illumination map feature. The connecting layer is used to stitch together the second reflection map feature, the second illumination map feature, and the first original image feature to obtain a first aggregated feature; The decomposition module is used to decompose the image to be decomposed based on the second reflection map feature, the second illumination map feature, and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed.
7. The image decomposition method according to claim 6, characterized in that, The image decomposition model is trained through the following steps: Feature extraction is performed on the original sample image to obtain the third reflection map feature, the third illumination map feature, and the second original image feature of the original sample image; A fourth reflection map feature is determined based on a second cross-attention feature between the third reflection map feature and the third illumination map feature, and the third reflection map feature; a fourth illumination map feature is determined based on the second cross-attention feature and the third illumination map feature. The fourth reflection map feature, the fourth illumination map feature, and the second original image feature are spliced together to obtain the second aggregated feature; Based on the fourth reflectance map feature, the fourth illumination map feature, and the second aggregation feature, the original sample image is decomposed to obtain the predicted original image, the predicted reflectance map, and the predicted illumination map of the original sample image; The image decomposition model is trained based on the predicted original image, the predicted reflectance map, the predicted illumination map, and the sample original image.
8. An image decomposition apparatus, characterized in that, include: The encoding module is used to extract features from the image to be decomposed, and obtain the first reflection map feature, the first illumination map feature, and the first original image feature of the image to be decomposed. The feature extraction module is used to determine a second reflection map feature based on a first cross-attention feature between the first reflection map feature and the first illumination map feature, and the first reflection map feature; and to determine a second illumination map feature based on the first cross-attention feature and the first illumination map feature. The feature extraction module is further configured to concatenate the second reflection map feature, the second illumination map feature, and the first original image feature to obtain a first aggregated feature; The decomposition module is used to decompose the image to be decomposed based on the second reflection map feature, the second illumination map feature and the first aggregation feature to obtain the reflection map and illumination map of the image to be decomposed.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1-7.
11. A computer program product, characterized in that, Includes computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the method as described in any one of claims 1-7.