Image processing method and device

Through the spatial three-plane model and adaptive feature fusion technology, realistic spectral BRDF parameters are generated, which solves the problem of scarce spectral BRDF data and realizes efficient spectral image synthesis.

CN120689215APending Publication Date: 2025-09-23BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510613931.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

It is difficult to obtain high-quality spectral BRDF data, and existing datasets are scarce and limited in scale, making it difficult to achieve realistic spectral image synthesis.

Method used

A spatial three-plane model is used to extract image features, feature vectors are generated by projection of illumination parameters, weights are adaptively selected for fusion, and spectral response data is generated by mapping. The RGB-spectral joint training strategy is combined to improve data utilization.

Benefits of technology

Generating realistic spectral BRDF parameters by a small number or even a single image solves the problem of scarce spectral BRDF data and improves the efficiency and quality of spectral image synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689215A_ABST
    Figure CN120689215A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, and relates to the technical field of computer vision. The method comprises the following steps: extracting image features from an image by using a spatial three-plane model, and generating spatial three planes and spectral three planes through the image features; projecting illumination parameters to the three spatial planes and the three spectral planes to obtain a plurality of feature vectors; adaptively selecting the weight of each feature vector, and fusing the plurality of feature vectors by adopting the weights; and decoding the fusion result through mapping to generate spectral response data of the image. According to the invention, the feature vectors of multiple planes can be adaptively combined, and vivid spectral BRDF parameters can be generated through a small number of images or even a single image by using a spatial three-plane model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to an image processing method and device. Background Art

[0002] Real light is composed of a range of wavelengths, and the synthesis of realistic images requires capturing the wavelength-dependent reflectance properties of materials under different lighting and viewing conditions. Spectral bidirectional reflectance distribution functions (BRDFs) play a key role in this process, enabling the synthesis of hyperspectral images for applications in remote sensing, materials analysis, and virtual reality.

[0003] However, obtaining high-quality spectral BRDF data is challenging because the measurement process requires high-resolution scanning of the four-dimensional domain, which is a tedious and time-consuming task, resulting in the existing spectral BRDF datasets being sparse and limited in size. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to provide an image processing method, device, electronic device and storage medium, which can specifically solve the existing problems.

[0005] Based on the above-mentioned purpose, in a first aspect, the present disclosure proposes a method for processing an image, comprising: extracting image features from an image using a spatial three-plane model, and generating a spatial three-plane and a spectral three-plane through the image features; projecting illumination parameters onto the spatial three-plane and the spectral three-plane to obtain a plurality of eigenvectors; adaptively selecting a weight for each eigenvector, and fusing the plurality of eigenvectors using the weight; decoding the fusion result through mapping to generate spectral response data of the image.

[0006] In the second aspect, an image processing device is also provided, including: an extraction unit, configured to use a spatial three-plane model to extract image features from an image, and generate a spatial three-plane and a spectral three-plane through the image features; a projection unit, configured to project illumination parameters onto the spatial three-plane and the spectral three-plane to obtain multiple feature vectors; a selection unit, configured to adaptively select the weight of each feature vector, and use the weight to fuse the multiple feature vectors; a generation unit, configured to decode the fusion result through mapping, and generate spectral response data of the image.

[0007] In a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect.

[0008] In a fourth aspect, a computer-readable storage medium is further provided, on which a computer program is stored, and the computer program is executed by a processor to implement any method described in the first aspect.

[0009] In a fifth aspect, a computer program product is also provided, comprising a computer program, wherein the computer program is executed by a processor to implement any one of the methods described in the first aspect.

[0010] In general, the present disclosure has at least the following beneficial effects: it provides a new spectral BRDF generation method, adaptively combines feature vectors of multiple planes, and can generate realistic spectral BRDF parameters using a small number or even a single image using a spatial three-plane model. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present disclosure and should not be regarded as limiting the scope of the present disclosure.

[0012] Figure 1 A flowchart of an image processing method according to an embodiment of the present disclosure is shown; Figure 2 A flow chart of a method for training a spatial three-plane model according to an embodiment of the present disclosure is shown; Figure 3a Another flowchart of the image processing method according to an embodiment of the present disclosure is shown; Figure 3b A comparison diagram of a reconstructed image obtained by applying the image processing method of the present application according to an embodiment of the present disclosure and other hyperspectral image reconstructed images is shown; Figure 4 A schematic diagram showing an image processing apparatus according to an embodiment of the present disclosure is shown; Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown; Figure 6 A schematic diagram of a storage medium provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0013] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.

[0014] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0015] Figure 1 The image processing method of the present disclosure is shown. In an embodiment of the present disclosure, the method includes: Step S101 : extracting image features from an image, and generating spatial three-planes and spectral three-planes based on the image features.

[0016] Step S102 : Projecting the illumination parameters onto the spatial three planes and the spectral three planes to obtain a plurality of eigenvectors.

[0017] Step S103 : adaptively selecting a weight for each feature vector, and fusing the multiple feature vectors using the weight.

[0018] Step S104 : decoding the fusion result through mapping to generate spectral response data of the image.

[0019] In this embodiment, the image is an RGB image. The image processing method executes the image processing method (an arbitrary electronic device) by projecting the wavelength and the incident-exit angle (expressed in the Rusinkiewicz coordinate system) onto three spatial planes and three spectral planes to generate six eigenvectors.

[0020] The mapping here can be done in various ways, such as a preset mapping model or a BRDF mapping module based on MLP, to ultimately generate the corresponding spectral response data. .

[0021] The present invention discloses a spectral BRDF generation method, which adaptively combines feature vectors of multiple planes and can generate realistic spectral BRDF parameters through a small number or even a single image using a spatial three-plane model.

[0022] In some optional implementations of any embodiment of the present disclosure, the adaptively selecting the weight of each feature vector includes: determining an attention-based adaptive score for each feature vector, and determining the weight of each feature vector based on the adaptive score.

[0023] In these optional implementations, the execution entity may determine the weight of each feature vector based on the adaptive score in various ways. For example, the adaptive score may be directly used as the weight of the feature vector, or the adaptive score may be subjected to a preset processing, such as multiplying it by a preset coefficient, to obtain the weight of the feature vector.

[0024] Optionally, each feature vector is a plurality of one-dimensional feature vectors; determining the attention-based adaptive score for each feature vector includes: converting the plurality of one-dimensional feature vectors into a three-dimensional feature map of a specific shape; mixing the three-dimensional feature maps of the plurality of feature vectors to obtain a mixed feature; determining a channel statistic for the mixed feature, the channel statistic being used to indicate global information of the mixed feature, the global information including contextual information; reducing the channel dimension of the channel statistic, and determining the attention-based adaptive score for each feature vector based on the reduction result.

[0025] Specifically, reducing the channel dimension of the channel statistics includes: using a preset convolution kernel, expanding the channel statistics into a descriptor and activating it to obtain an activated descriptor; for each feature vector, using a corresponding convolution kernel, convolving the activated descriptor to obtain a reduction result corresponding to the feature vector.

[0026] Optionally, the adopting the weights to fuse the feature vectors includes: weighting and dimensionality conversion of the feature vectors according to their respective adaptive scores.

[0027] Figure 2 The training method of the spatial three-plane model according to the embodiment of the present disclosure is shown. The spatial three-plane is the spatial three-plane in any of the above implementations. Figure 2 As shown, the training method includes: generating RGB loss values ​​through RGB image samples; generating spectral BRDF loss values ​​through spectral BRDF samples; training the spatial three-plane model to be trained through the RGB loss values, and training the spatial three-plane to be trained through the BRDF loss values ​​to obtain the spatial three-plane.

[0028] In some optional implementations of any embodiment of the present disclosure, the spectral BRDF sample includes the true value of the spectral response data; the generation of the spectral BRDF loss value through the spectral BRDF sample includes: obtaining the estimated values ​​of the three planes of the space to be trained during the training process; using a logarithmic relative mapping method to process the true value and the estimated value, the data before processing is high dynamic range data, and the data after processing is low dynamic range data; determining the difference between the processed true value and the estimated value; and determining the loss value of the spectral BRDF based on the difference.

[0029] Optionally, determining the loss value of the spectral BRDF according to the difference includes: determining a total variation loss value according to the estimated value; and determining the loss value of the spectral BRDF according to the total variation loss value and the difference.

[0030] Optionally, the feature vector includes a spectral feature vector and a spatial feature vector; generating an RGB loss value through an RGB image sample includes: at a position where the target spectral BRDF response occurs, averaging all spectral feature vectors along the dimension at that position to obtain a target spectral feature vector; and generating an RGB loss value based on the target spectral feature vector.

[0031] Figure 3a The image processing method of the present disclosure is shown. In the embodiment of the present disclosure, the method can be summarized as follows: given an input RGB image, the spatial three-plane model, i.e., the SSTA network, uses an encoder-decoder architecture to extract two sets of three-planes: a spatial three-plane for encoding the reflection response of the incident and exit light angles; and a spectral three-plane for capturing the response at different angles and wavelengths. Next, we project the wavelength and incident-exit angle (expressed in the Rusinkiewicz coordinate system) onto these two three-planes to generate six feature vectors. Subsequently, the AFF module fuses these features and generates a potential BRDF feature, which is input into the MLP-based BRDF mapping module to finally predict the corresponding spectral response data such as spectral reflectance. .

[0032] 1. First, spectral-spatial three-plane aggregation can be performed: The dynamic neural radiation field can be represented as a function of the three-dimensional spatial position and one-dimensional timestamp Therefore, K-Planes is a dynamic NeRF method that converts the radiance of dynamic scenes into Modeled as:

[0033] In order to decompose the static standard scene and the dynamic motion field, a plane decomposition method can be introduced to express the four-dimensional function as characteristic planes. Among them, three planes Encodes spatial domain information, while the other three planes Capturing spatiotemporal variations allows for a clear decoupling of the static radiation field and the temporal motion field.

[0034] 2. Secondly, the spectral response data, also known as spectral BRDF, can be expressed: Based on the Rusinkiewicz coordinate system, the spectral BRDF response can also be modeled as a function of the incident-outgoing angle and wavelength Specifically, given the incident light direction and the direction of the emitted light , the half vector can be calculated as , where the angle express Relative to a fixed surface normal The polar angle, Is the emission direction With half vector The polar angle between It is around Similar to the 4D function in Dynamic NeRF, the spectral BRDF response It can be expressed as:

[0035] Decompose the 4D spectral response into two triplanes, where the three planes It depends only on the angle parameter and is called the three planes of space. We denote these three planes as 、 and At the same time, the other three planes Including wavelength dependence, known as the spectral three planes, respectively denoted as 、 and Since the spatial triplanes are only dependent on the incident and exit angles, they can be shared between RGB and spectral BRDFs. As described in the RGB-spectral joint training strategy, this decomposed representation can leverage RGB data to enhance the generation of spectral BRDFs.

[0036] 3. Perform image to three-plane mapping We use a CNN-based encoder-decoder network architecture to generate spatial and spectral triplanars. The network input is an RGB image of the target material rendered on a sphere, and the output is a 6-channel feature map consisting of two triplanars: the spatial triplanar , and the spectral triplane To obtain the spectral BRDF response value, we first calculate the coordinates , and project the coordinates onto the two three planes, resulting in a dimension of of feature vectors.

[0037]

[0038]

[0039] in Represents the projection operator. and are associated with spatial and spectral parameters, respectively. -dimensional feature vector. The following sections describe how to fuse these spatial and spectral features to predict spectral response data.

[0040] 4. Perform adaptive feature fusion (AFF) Different from the previous method of feature fusion through dot product, this paper proposes an adaptive feature fusion (AFF) module to dynamically select appropriate weights to fuse different features. Figure 3a As shown, we first transform each eigenvector, i.e. The one-dimensional feature vector of Then, we apply Convolution and element-by-element addition promote feature interaction, thereby obtaining fused mixed features. .

[0041]

[0042] Next, we perform global average pooling To obtain channel statistics that capture the global context :

[0043] To expand the key feature representation, inspired by the bottleneck structure, we use Convolution will Expanded to descriptor (where d=32), and then activated with ReLU, we get . Subsequently, we apply six branch-specific convolutions with kernels , to reduce the channel dimension, strengthen the inter-channel dependency, and extract multi-branch feature representation. The specific process is as follows:

[0044]

[0045] In order to adaptively weight different feature branches, we compute an attention-based adaptive score for each branch. The reduction result of the scale branch Weight Obtained through the softmax-based attention mechanism:

[0046] Finally, the input spatial and spectral features are weighted according to their respective adaptation scores and aggregated into the final fusion output , which effectively captures the relative importance of different features in the fusion:

[0047] The reshape operator transforms a two-dimensional feature map into a one-dimensional feature vector. In this way, our AFF module effectively enhances the interaction between features by adaptively selecting receptive fields and assigning dynamic weights, thereby ensuring the optimal fusion strategy and achieving better spectral BRDF reconstruction.

[0048] 5. BRDF Mapping Module: Given fusion features , we designed the BRDF mapping module, which is a multi-layer perceptron (MLP) with three hidden layers and ReLU activation function to map the corresponding coordinates Decode and generate spectrum :

[0049] in, Represents the learnable parameters in the MLP network.

[0050] Through SSTA feature representation, selective feature fusion based on AFF module and BRDF mapping network based on MLP, we obtain spectral reflectance at a single coordinate. By repeating this process for different incident-outgoing directions and wavelengths, we achieve full spectral BRDF reconstruction, enabling flexible spectral image rendering under different lighting and shape conditions.

[0051] 6. RGB spectrum joint training strategy The spatial and spectral decomposition introduced by the SSTA module aims to leverage the abundance of RGB BRDF data, addressing the scarcity of spectral BRDF data during spectral BRDF generation. Building on this, we propose a joint RGB-spectral training strategy that utilizes both RGB and spectral BRDF data for training. Below, we detail the workflow and loss function used for training with both RGB and spectral BRDF data.

[0052] 7. Spectral BRDF Training Since the measured BRDF usually has a high dynamic range (HDR), this will lead to large fluctuations in the value, especially in the highlight part. -law, a logarithmic relative mapping method that compresses HDR data into a range that is easier to train:

[0053] in, represents the scaled BRDF value, is the compression parameter that controls the compression strength. sampling coordinates , we are passing by After -law processing, the spectral response between the true value and the estimated value was measured and difference.

[0054]

[0055] At the same time, we adopt the total variation (TV) loss To constrain the smoothness of the generated spectral BRDF. The total loss of the spectral BRDF can be expressed as:

[0056] Among them, the coefficient Set to 2 to balance the total variation loss.

[0057] 8. About RGB and BRDF training RGB BRDF and spectral BRDF share the same spatial three-plane characteristics , but in the spectral three-plane characteristics Specifically, for each set of spectral reflectance and its corresponding coordinates in the spectral BRDF , we sample the corresponding spectral three-plane features using the projection operation described above.

[0058] In order to obtain the spectral characteristics of the RGB-BRDF case, we Along the three planes of the spectrum All features of the dimension are averaged.

[0059]

[0060] in, Represents the spectrum in three planes Dimensions, represents the spectral characteristics of the RGB BRDF data. Intuitively, this averaging operation converts the spectral response into a grayscale average reflectance response. Given the spatial and spectral characteristics, we also generate the response , and supervise it using the grayscale value of RGB BRDF.

[0061]

[0062] To train the model, we use the loss function and (i.e. ) to process RGB and spectral BRDF data, obtaining RGB loss and spectral BRDF loss respectively. This joint training strategy enables us to utilize large-scale RGB BRDF data, significantly improving the generalization ability of the model.

[0063] The following table shows the experimental results on ambient light and parallel light test data:

[0064] Figure 3b A comparison diagram of the reconstructed image obtained by applying the image processing method of the present application and other hyperspectral image reconstructed images is shown.

[0065] The present disclosure provides an image processing device, which is used to execute the image processing method described in the above embodiment. Figure 4 As shown, the device includes: an extraction unit 401, configured to extract image features from an image using a spatial three-plane model, and generate a spatial three-plane and a spectral three-plane through the image features; a projection unit 402, configured to project illumination parameters onto the spatial three-plane and the spectral three-plane to obtain a plurality of feature vectors; a selection unit 403, configured to adaptively select a weight for each feature vector, and fuse the plurality of feature vectors using the weight; a generation unit 404, configured to decode the fusion result through mapping, and generate spectral response data of the image.

[0066] The image processing device provided by the above-mentioned embodiment of the present disclosure and the image processing method provided by the embodiment of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0067] The embodiment of the present disclosure further provides an electronic device corresponding to the image processing method provided in the above embodiment, so as to execute the above image processing method.

[0068] Please refer to Figure 5 , which shows a schematic diagram of an electronic device provided by some embodiments of the present disclosure. Figure 5 As shown, the electronic device 50 includes: a processor 500, a memory 501, a bus 502 and a communication interface 503, and the processor 500, the communication interface 503 and the memory 501 are connected via the bus 502; the memory 501 stores a computer program that can be run on the processor 500, and when the processor 500 runs the computer program, it executes the method provided in any of the aforementioned embodiments of the present disclosure.

[0069] Memory 501 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between the system network element and at least one other network element is achieved through at least one communication interface 503 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.

[0070] Bus 502 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, etc. Memory 501 is used to store programs, and processor 500 executes the programs upon receiving execution instructions. The image processing method disclosed in any of the aforementioned embodiments of the present disclosure may be applied to or implemented by processor 500.

[0071] The processor 500 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in the processor 500. The processor 500 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 501 , and the processor 500 reads the information in the memory 501 and completes the steps of the above method in combination with its hardware.

[0072] The electronic device provided by the embodiment of the present disclosure and the image processing method provided by the embodiment of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented by them.

[0073] The present disclosure also provides a computer-readable storage medium corresponding to the image processing method provided in the above embodiment. Figure 6 The computer-readable storage medium shown is a CD 60 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the image processing method provided by any of the aforementioned embodiments is executed.

[0074] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.

[0075] The computer-readable storage medium provided by the above-mentioned embodiment of the present disclosure and the image processing method provided by the embodiment of the present disclosure are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0076] It should be noted that: In the above text, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present disclosure is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0077] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present disclosure.

[0078] The embodiments of the present disclosure are described above in conjunction with the accompanying drawings, which are only specific implementation methods of the present disclosure. However, the present disclosure is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present disclosure, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present disclosure and the claims, which are all within the protection of the present disclosure.

Claims

1. A method for processing an image, characterized in that: include: Extracting image features from the image using a spatial three-plane model, and generating a spatial three-plane and a spectral three-plane based on the image features; Projecting the illumination parameters onto the spatial three planes and the spectral three planes to obtain a plurality of eigenvectors; Adaptively selecting a weight for each feature vector, and fusing the multiple feature vectors using the weight; The fusion result is decoded through mapping to generate spectral response data of the image.

2. The method according to claim 1, characterized in that The adaptively selecting the weight of each feature vector includes: For each feature vector, an attention-based adaptive score is determined, and a weight of each feature vector is determined based on the adaptive score.

3. The method according to claim 2, characterized in that Each eigenvector is a number of one-dimensional eigenvectors; Determining the attention-based adaptive score for each feature vector includes: Converting the multiple one-dimensional feature vectors into three-dimensional feature maps of a specific shape; Mixing the three-dimensional feature maps of the multiple feature vectors to obtain mixed features; Determining a channel statistic for the hybrid feature, where the channel statistic is used to indicate global information of the hybrid feature, where the global information includes context information; The channel dimensionality is reduced for the channel statistics, and an attention-based adaptive score is determined for each feature vector based on the reduction result.

4. The method according to claim 3, characterized in that The reducing the channel dimension of the channel statistics includes: Using a preset convolution kernel, the channel statistics are expanded into a descriptor and activated to obtain an activated descriptor; For each feature vector, a corresponding convolution kernel is used to convolve the activation descriptor to obtain a reduction result corresponding to the feature vector.

5. The method according to claim 2, characterized in that The step of fusing the feature vectors using the weights includes: Each feature vector is weighted and dimensionally transformed according to its adaptive score.

6. A method for training a spatial three-plane model, characterized in that: The three spatial planes are the three spatial planes in any one of claims 1 to 5; The training method comprises: Generate RGB loss values ​​from RGB image samples; Generate spectral BRDF loss value through spectral BRDF sample; The spatial three-plane model to be trained is trained by the RGB loss value, and the spatial three-plane to be trained is trained by the BRDF loss value to obtain the spatial three-plane.

7. The method according to claim 6, characterized in that The spectral BRDF samples include true values ​​of spectral response data; Generating a spectral BRDF loss value through a spectral BRDF sample includes: Obtaining estimated values ​​of the three planes of the space to be trained during the training process; The true value and the estimated value are processed using a logarithmic relative mapping method, wherein the data before the processing is high dynamic range data and the data after the processing is low dynamic range data; Determine the difference between the true value and the estimated value after processing; A loss value of the spectral BRDF is determined according to the difference.

8. The method according to claim 7, characterized in that Determining the loss value of the spectral BRDF according to the difference includes: determining a total variation loss value based on the estimated value; A loss value of the spectral BRDF is determined according to the total variation loss value and the difference value.

9. The method according to claim 6, characterized in that The feature vector includes a spectral feature vector and a spatial feature vector; Generating an RGB loss value using an RGB image sample includes: At the location of the target spectral BRDF response, all spectral feature vectors along the mid-dimension of the location are averaged to obtain the target spectral feature vector; Generate an RGB loss value according to the target spectral feature vector.

10. An image processing device, characterized in that: include: an extraction unit configured to extract image features from the image using a spatial three-plane model, and generate a spatial three-plane and a spectral three-plane according to the image features; A projection unit is configured to project the illumination parameters onto the spatial three planes and the spectral three planes to obtain a plurality of eigenvectors; A selection unit is configured to adaptively select a weight of each feature vector, and fuse the multiple feature vectors using the weight; The generating unit is configured to decode the fusion result through mapping to generate spectral response data of the image.