Diffraction partition coding element modeling method, element, device and snapshot hyperspectral imaging method

By dividing the diffractive optical elements in the polar coordinate system and performing iterative optimization, the problems of low coding efficiency and low resolution in the diffraction imaging system are solved, achieving efficient hyperspectral imaging, which is suitable for compact snapshot hyperspectral imaging systems.

CN121937622APending Publication Date: 2026-04-28BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2025-11-24
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing diffraction imaging systems suffer from low coding efficiency, significant quantization noise and manufacturing errors, and low spatial resolution, which limits the development of hyperspectral imaging systems.

Method used

The diffraction partition coding element modeling method is adopted. The design plane of the diffraction optical element is divided into multiple equiangular sectors in the polar coordinate system to generate a diffraction partition coding matrix. The quantization and manufacturing errors are processed through iterative optimization steps. The coded image is generated by combining the spectral response and noise of the RGB sensor. The image is then input into the hyperspectral image reconstruction model for optimization to generate the optimal DPE element and reconstruction model.

Benefits of technology

It improves coding efficiency, reduces the impact of quantization noise and manufacturing errors, enhances spatial resolution, and achieves optimal system performance through hardware-software co-optimization, ensuring that DPE design is directly related to imaging effect, thereby improving the resolution and reliability of hyperspectral images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937622A_ABST
    Figure CN121937622A_ABST
Patent Text Reader

Abstract

The invention provides a diffraction partition coding element modeling method, element and device, and a snapshot hyperspectral imaging method, and the modeling method comprises the steps: carrying out the polar coordinate partitioning of a diffraction optical element, generating a DPE matrix, carrying out the rotation transformation of the DPE matrix into an initial height map, carrying out the iterative optimization of a full-precision height map, and obtaining an initial height map; each round of optimization comprises the following steps: quantifying the height map based on the learnable depth and introducing a manufacturing error to obtain a simulated manufacturing height map; calculating a spread function (PSF) set of each wavelength point; the PSF and an original hyperspectral image are subjected to convolution, and a coded image is generated in combination with sensor response; processing the coded image and the PSF through a reconstruction network to obtain a reconstructed image; and calculating the total loss to update the full-precision height map, learn the depth and reconstruct the network, and outputting a DPE modeling result and a target hyperspectral image reconstruction model after iteration is completed. The problems of low coding efficiency, large influence of quantization noise and manufacturing errors, low spatial resolution and the like in an existing diffraction hyperspectral imaging system can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of optoelectronic technology, and in particular to diffraction partition coding element modeling methods, elements, devices, and snapshot hyperspectral imaging methods. Background Technology

[0002] Hyperspectral imaging can capture high-dimensional spatial and spectral information, revealing the intrinsic properties of matter, and has broad application prospects in many fields such as biomedical detection, food industry, agriculture, and ecology. Traditional hyperspectral imaging techniques usually use line scanning or band scanning strategies to acquire three-dimensional hyperspectral data cubes, but these methods are often time-consuming and require a large number of components, resulting in high system complexity and high cost, making them unsuitable for dynamic scenarios.

[0003] Currently, compact lensless imaging systems based on point spread function (PSF), such as diffraction imaging systems, provide a pathway for the development of low-cost, high-resolution integrated hyperspectral imaging systems.

[0004] However, current technological bottlenecks severely restrict the practical development of diffraction imaging systems. First, the impact of height maps on spectral coding efficiency in diffractive optical elements has not been fully studied, leading to low coding efficiency. Second, manufacturing defects, particularly quantization noise and manufacturing errors, cause significant discrepancies between simulation models and experimental results, thus reducing coding fidelity. Third, existing diffraction coding mechanisms introduce systematic ambiguity effects, reducing the spatial resolution of reconstructed hyperspectral images and limiting the system's imaging quality. Summary of the Invention

[0005] In view of this, embodiments of this application provide a diffraction partition coding element modeling method, element, device, and snapshot hyperspectral imaging method to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of this application provides a method for modeling diffraction partitioned coding elements, including: The design plane of the diffractive optical element is divided into multiple equiangular sectors in the polar coordinate system to generate a diffraction partition coding (DPE) matrix. This DPE matrix is ​​then rotated and transformed to generate an initial DPE height map. The initial DPE height map is used as the current full-precision height map, and an optimization step of at least one iteration is performed on the full-precision height map. The optimization step includes: quantizing and processing the full-precision height map corresponding to the current iteration based on the learnable depth corresponding to the current iteration to obtain a simulated manufacturing height map; obtaining the point spread function (PSF) corresponding to the wavelength of each sector based on the simulated manufacturing height map to form a PSF set corresponding to the current iteration; convolving the PSF set with the pre-simulated original hyperspectral image, and combining the spectral response and noise of the RGB sensor to generate the encoded image corresponding to the current iteration; inputting the encoded image and the PSF set into a hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs the reconstructed hyperspectral image corresponding to the current iteration; determining the total loss of the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, the learnable depth, and the hyperspectral image reconstruction model based on the total loss. After the iteration is completed, the latest full-precision height map is used as the modeling result data of the DPE element, and the latest hyperspectral image reconstruction model is used as the target hyperspectral image reconstruction model. The modeling result data of the DPE element and the target hyperspectral image reconstruction model are then output.

[0007] In some embodiments of this application, the step of dividing the design plane of the diffractive optical element into multiple equiangular sectors in a polar coordinate system, generating a diffraction partition coding (DPE) matrix, and rotating and transforming the DPE matrix to generate an initial DPE height map includes: Acquire imaging band requirement data, which includes spectral range and number of bands; The design plane of the diffractive optical element is placed in a polar coordinate system, and the design plane is divided into multiple equiangular sectors according to the number of wavebands; Based on the spectral range, the wavelength corresponding to each band of each sector is determined, and initial height distribution parameters are generated for each sector to obtain an initialized DPE matrix, wherein the DPE matrix is ​​used to define the initial height distribution parameters of each sector in the radial direction. By rotating and transforming the DPE matrix, an initial DPE height map in the Cartesian coordinate system corresponding to the current iteration is generated, and the design plane is used as the height map plane.

[0008] In some embodiments of this application, the step of quantizing and processing manufacturing errors on the full-precision heightmap corresponding to the current iteration based on the learnable depth corresponding to the current iteration to obtain the simulated manufacturing heightmap includes: Based on the learnable depth corresponding to the current iteration round, the full-precision height map corresponding to the current iteration round is scaled to obtain the scaled height map; wherein, the learnable depth corresponding to the first iteration round is a predefined parameter, and the value of the learnable depth is between π and 4π; The step height is determined based on the learnable depth corresponding to the current iteration round and the preset number of quantization steps; The scaled height map is uniformly quantized based on the step height to obtain a quantized height map with quantization noise. Add manufacturing errors to the quantized height map to simulate errors generated during the lithography and etching processes, and obtain the simulated manufacturing height map corresponding to the current iteration.

[0009] In some embodiments of this application, obtaining the point spread function (PSF) corresponding to the wavelength of each sector based on the simulated height map to form the PSF set corresponding to the current iteration round includes: According to Fresnel's theorem, the incident light field is modeled as a spherical wave to obtain the corresponding incident light field; Based on the height map after the simulation manufacturing, the phase delay of the wavelength corresponding to each sector is determined. Based on the phase delay of each wavelength, the output light field obtained after the incident light field passes through the simulated height map is determined: The light field reaching the RGB sensor plane is determined based on the emitted light field and the distance between the RGB sensor plane and the aperture. The square of the light field reaching the sensor plane is determined to obtain the point spread function (PSF) for each wavelength in the current iteration round. The PSFs for each wavelength constitute the PSF set for the current iteration round.

[0010] In some embodiments of this application, the step of convolving the PSF set with the pre-simulated original hyperspectral image and combining it with the spectral response and noise of the RGB sensor to generate the encoded image corresponding to the current iteration round includes: Based on the PSF set corresponding to the current iteration round and the original hyperspectral image obtained in advance based on the true value simulation, the hyperspectral image corresponding to each band of each sector is convolved with the PSF to obtain the convolved image. The convolutional image is integrated onto the three channels of the RGB sensor, and the spectral response and noise of the RGB sensor are combined to obtain the encoded image corresponding to the current iteration.

[0011] In some embodiments of this application, the step of inputting the encoded image and PSF set into a hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs the reconstructed hyperspectral image corresponding to the current iteration includes: The encoded image and the PSF set corresponding to the current iteration are input into the feature extraction module of the hyperspectral image reconstruction model, so that the feature extraction module extracts the preprocessed feature map and inputs the preprocessed feature map into the reconstruction network of the hyperspectral image reconstruction model, so that the reconstruction network reconstructs the preprocessed feature map to output the reconstructed hyperspectral image corresponding to the current iteration. The feature extraction module includes: The first branch is used to input the encoded image into the multi-feature spectral attention network MFSA-Net in the reconstruction network, the first deconvolution unit in the second branch, and the second deconvolution unit in the third branch, respectively. The second branch is used to map the PSF set to the spectral feature space based on the first RGB mapping unit, and to deconvolve the PSF set mapped to the spectral feature space and the encoded image based on the first deconvolution unit, so as to output the spatial features corresponding to the encoded image. The third branch is used to extract the restoration space features corresponding to the encoded image based on multiple convolution units, and to map the restoration space features to the spectral feature space based on the second RGB mapping unit. Then, based on the second deconvolution unit, the restoration space features mapped to the spectral feature space and the encoded image are deconvolved to output the spectral reconstruction features corresponding to the encoded image. A dual-branch feature fusion unit is used to fuse the spatial features and the spectral reconstruction features to obtain a preprocessed feature map, and input the preprocessed feature map into the Res-Unet in the reconstruction network; The reconstructed network includes: Multi-head Spectral Attention Network (MFSA-Net), Res-Unet, and an addition operation module; The multi-head spectral attention network MFSA-Net includes: The first convolutional unit is used to perform preliminary feature extraction and channel dimension transformation on the encoded image to obtain a first feature map, and the first feature map is input into the multi-level embedding MLE unit and the feature fusion unit respectively. The multi-level embedding (MLE) unit is used to achieve multi-scale feature embedding of the first feature map through multi-level convolution and upsampling operations to obtain the second feature map. The first spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the second feature map through the spectral attention mechanism to obtain the third feature map, and input the third feature map into the second multi-channel residual fusion (MCRF) unit and the first downsampling unit respectively. The first downsampling unit is used to reduce the spatial resolution of the third feature map and increase the channel dimension to obtain the fourth feature map; The second spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the fourth feature image through the spectral attention mechanism to obtain the fifth feature image, and input the fifth feature image into the first MCRF unit and the second downsampling unit respectively. The second downsampling unit is used to reduce the spatial resolution of the fifth feature map and increase the channel dimension to obtain the sixth feature map; The third spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the sixth feature map through the spectral attention mechanism to obtain the seventh feature map. The first upsampling unit is used to restore the spatial resolution of the seventh feature map and reduce the channel dimension to obtain the eighth feature map; The first multi-channel residual fusion MCRF unit is used to fuse the fifth feature map and the eighth feature map to obtain the ninth feature map; The second convolutional unit is used to extract features from the ninth feature map to obtain the tenth feature map; The fourth spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the tenth feature map through the spectral attention mechanism to obtain the eleventh feature map; The second upsampling unit is used to restore the eleventh feature map to the original input resolution to obtain the twelfth feature map; The second multi-channel residual fusion MCRF unit is used to fuse the third feature map and the twelfth feature map to obtain the thirteenth feature map; The third convolutional unit is used to extract features from the thirteenth feature map to obtain the fourteenth feature map; The fifth spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the fourteenth feature map through a spectral attention mechanism to obtain the fifteenth feature map; A feature fusion unit is used to perform residual connection on the first feature map and the fifteenth feature map to obtain the sixteenth feature map; The fourth convolutional unit is used to map the sixteenth feature map to the target spectral dimension to obtain a reconstructed hyperspectral map, and input the reconstructed hyperspectral map into the summation operation module; The Res-Unet is used to reconstruct the hyperspectral data cube corresponding to the preprocessed feature map, and the hyperspectral data cube is input into the summation operation module; the summation operation module is used to perform addition and convolution on the reconstructed hyperspectral map and the hyperspectral data cube to output the reconstructed hyperspectral image corresponding to the current iteration round; The first, second, third, fourth, and fifth spectral attention SA units are each equipped with a multi-head spectral attention (MHSA) unit and a multi-scale feedforward network (MS FFN) unit.

[0012] In some embodiments of this application, determining the total loss for the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, the learnable depth, and the hyperspectral image reconstruction model based on the total loss, includes: Based on the PSF set corresponding to the current iteration round, the average absolute difference value of the point spread function PSF corresponding to the wavelength of each sector is obtained by weighted summation, and the average absolute difference value is mapped through the activation function to obtain the spectral difference regularization loss corresponding to the current iteration round. Based on the PSF set corresponding to the current iteration round, obtain the corresponding Fisher information matrix and determine the sum of all elements in the Fisher information matrix, and determine the Fisher information regularization loss corresponding to the current iteration round based on the sum of all elements in the Fisher information matrix. In addition, the mean absolute error between the reconstructed hyperspectral image and the original hyperspectral image is determined, and the mean absolute error is used as the reconstruction loss of the reconstruction network in the hyperspectral image reconstruction model; The total loss for the current iteration round is determined based on the spectral difference regularization loss, the Fisher information regularization loss, and the reconstruction loss corresponding to the current iteration round. The total loss corresponding to the current iteration is backpropagated to update the weights of the reconstruction network in the full-precision height map, learnable depth, and hyperspectral image reconstruction model based on the total loss, thereby obtaining the latest full-precision height map, learnable depth, and hyperspectral image reconstruction model in sequence.

[0013] Another aspect of this application provides a diffraction partition coding element, comprising: An optical substrate; and Micro-nanostructures formed on the optical substrate, wherein the surface of the micro-nanostructures has a height distribution based on polar coordinate partitioning; The polar coordinate partitioning divides the optical aperture into multiple equiangular sectors, and the height distribution of each sector is formed by rotating around the polar coordinate center. The height distribution is defined by the modeling result data of the DPE element. The DPE element modeling result data is obtained by the diffraction partition coding element modeling method provided in the first aspect above, such that the total depth of the height distribution is in the range of 0.768 μm to 3.072 μm.

[0014] A third aspect of this application provides an encoded image acquisition device, comprising: an RGB sensor and the diffraction partitioning encoding element as described in the second aspect above; The diffraction partition coding element is installed at a preset distance in front of the RGB sensor, and the diffraction partition coding element is configured to perform phase modulation on the incident light field from the target scene; the preset distance includes 50mm; The RGB sensor is configured to receive a light signal that has been phase-modulated by the diffraction partition coding element and convert the light signal into a coded image of the target scene.

[0015] A fourth aspect of this application provides a snapshot hyperspectral imaging method, comprising: The encoding image of the target scene acquired by the encoding image acquisition device as described in the third aspect above is obtained, and optical simulation is performed on the modeling result data of the DPE element to obtain the PSF set corresponding to the diffraction partition encoding element. The encoded image of the target scene and the PSF set are input into the target hyperspectral image reconstruction model so that the target hyperspectral image reconstruction model outputs the reconstructed hyperspectral image of the target scene.

[0016] A fifth aspect of this application provides a diffraction partition coding element modeling apparatus, comprising: The initialization module is used to divide the design plane of the diffractive optical element into multiple equiangular sectors in the polar coordinate system, generate a diffraction partition coding (DPE) matrix, and rotate and transform the DPE matrix to generate an initial DPE height map. An iterative training module is used to take the initial DPE height map as the current full-precision height map and perform at least one iteration of optimization steps on the full-precision height map. The optimization steps include: quantizing and processing the full-precision height map corresponding to the current iteration based on the learnable depth corresponding to the current iteration to obtain a simulated manufacturing height map; obtaining the point spread function (PSF) corresponding to the wavelength of each sector based on the simulated manufacturing height map to form a PSF set corresponding to the current iteration; convolving the PSF set with the pre-simulated original hyperspectral image and combining it with the spectral response and noise of the RGB sensor to generate an encoded image corresponding to the current iteration; inputting the encoded image and the PSF set into a hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs a reconstructed hyperspectral image corresponding to the current iteration; determining the total loss of the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, the learnable depth, and the hyperspectral image reconstruction model based on the total loss. The result output module is used to, after the iteration is completed, take the latest full-precision height map as the DPE element modeling result data, take the latest hyperspectral image reconstruction model as the target hyperspectral image reconstruction model, and output the DPE element modeling result data and the target hyperspectral image reconstruction model.

[0017] A sixth aspect of this application provides a snapshot hyperspectral imaging system, comprising: The image acquisition and dot diffusion module is used to acquire the encoded image of the target scene acquired by the encoded image acquisition device as provided in the third aspect above, and to perform optical simulation on the modeling result data of the DPE element to obtain the PSF set corresponding to the diffraction partition encoding element. The hyperspectral image reconstruction module is used to input the encoded image of the target scene and the PSF set into the target hyperspectral image reconstruction model, so that the target hyperspectral image reconstruction model outputs the reconstructed hyperspectral image of the target scene.

[0018] The seventh aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method.

[0019] The eighth aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method.

[0020] The ninth aspect of this application provides a computer program product comprising a computer program that, when executed by a processor, implements the diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method.

[0021] The diffraction partition coding element modeling method provided in this application divides the design plane of the diffraction optical element into multiple equiangular sectors in a polar coordinate system to generate a diffraction partition coding (DPE) matrix. This DPE matrix is ​​then rotated and transformed to generate an initial DPE height map. The initial DPE height map is used as the current full-precision height map, and at least one iteration of optimization steps is performed on the full-precision height map. The optimization steps include: quantizing and processing manufacturing errors on the full-precision height map corresponding to the current iteration based on the learnable depth corresponding to the current iteration, to obtain a simulated manufacturing height map; and then, based on the simulated manufacturing height map... The point spread function (PSF) corresponding to the wavelength of each sector is obtained to form the PSF set corresponding to the current iteration round. The PSF set and the original hyperspectral image obtained by pre-simulation are convolved, and the spectral response and noise of the RGB sensor are combined to generate the encoded image corresponding to the current iteration round. The encoded image and the PSF set are input into the hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs the reconstructed hyperspectral image corresponding to the current iteration round. The total loss of the current iteration round is determined according to the PSF set, the reconstructed hyperspectral image and the original hyperspectral image, and the full-precision height map and the learnable depth are updated according to the total loss. The system includes a hyperspectral image reconstruction model. After iteration, the latest full-precision height map is used as the modeling result data for the DPE element, and the latest hyperspectral image reconstruction model is used as the target hyperspectral image reconstruction model. The modeling result data for the DPE element and the target hyperspectral image reconstruction model are output. An initial height map is generated through polar coordinate partitioning and rotation transformation, which can provide a good data foundation for subsequent DPE element modeling optimization and effectively improve coding efficiency. By introducing a learnable deep quantization mechanism, the actual manufacturing process of the DPE element can be effectively simulated, thereby effectively reducing the impact of quantization noise and manufacturing errors on the modeling result of the DPE element. Furthermore, the DPE element modeling results have a simple structure, enabling it to be manufactured using simple methods, thereby effectively improving spatial resolution. By jointly optimizing the hardware (DPE) and software (hyperspectral image reconstruction model), optimal system performance can be achieved. The end-to-end optimization framework ensures that the DPE design is directly related to the actual imaging effect, effectively improving the execution efficiency of the complete closed-loop feedback link of encoding and decoding. It also ensures that the optical encoding mode of the final modeled DPE is one of the "optimal solutions" in information theory, thereby effectively improving the resolution, reliability, and effectiveness of reconstructing hyperspectral images using DPE elements and hyperspectral image reconstruction models.

[0022] Additional advantages, objectives, and features of this application will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon review of the following description, or may be learned by practice of the application. The objectives and other advantages of this application can be realized and obtained by means of the structures specifically pointed out in the specification and drawings.

[0023] Those skilled in the art will understand that the purposes and advantages that can be achieved with this application are not limited to those specifically described above, and that the above and other purposes that this application can achieve will be more clearly understood from the following detailed description. Attached Figure Description

[0024] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. The components in the drawings are not drawn to scale but are merely for illustrating the principles of this application. For ease of illustration and description of certain parts of this application, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to this application. In the drawings: Figure 1 This is a schematic diagram of the first process of a diffraction partition coding element modeling method in one embodiment of this application.

[0025] Figure 2 This is a schematic diagram of a second process for modeling a diffraction partition coding element in one embodiment of this application.

[0026] Figure 3 This is a flowchart illustrating the optimization steps in one embodiment of this application.

[0027] Figure 4 This is a schematic diagram of a hyperspectral image reconstruction model in one embodiment of this application.

[0028] Figure 5 This is a schematic diagram of the multi-head spectral attention network MFSA-Net in one embodiment of this application.

[0029] Figure 6 This is a schematic diagram of a spectral attention (SA) unit in one embodiment of this application.

[0030] Figure 7 This is a schematic diagram of a multi-level embedded MLE unit in one embodiment of this application.

[0031] Figure 8 This is a schematic diagram of a multi-channel residual fusion (MCRF) unit in one embodiment of this application.

[0032] Figure 9 This is a schematic diagram of a multi-head spectral attention (MHSA) unit in one embodiment of this application.

[0033] Figure 10 This is a schematic diagram of a multi-scale feedforward network (MS FFN) unit in one embodiment of this application.

[0034] Figure 11 This is a schematic diagram of the diffraction partition coding element modeling device in one embodiment of this application.

[0035] Figure 12 This is a schematic diagram of the structure of the encoded image acquisition device in one embodiment of this application.

[0036] Figure 13 This is a schematic flowchart of a snapshot hyperspectral imaging method according to an embodiment of this application.

[0037] Figure 14 This is a schematic diagram of the learning-based diffraction partition coding joint optimization framework in an application example of this application.

[0038] Figure 15 This is a schematic diagram of the DPE matrix generation process in an application example of this application.

[0039] Figure 16 This is a schematic diagram of the initial DPE height map in an application example of this application.

[0040] Figure 17 This is a schematic diagram of the spectral imaging results data of a hyperspectral system in an application example of this application.

[0041] Figure 18 This is a schematic diagram of the experimental results data of system prototype construction and performance verification in an application example of this application.

[0042] Figure 19 This is a schematic diagram of the application experimental data in an application example of this application.

[0043] Figure 20 This is a schematic diagram of dynamic hyperspectral imaging verification data in a real-world scenario, as shown in an application example of this application. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit it.

[0045] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0046] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0047] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0048] In the following description, embodiments of the present application will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0049] It should be noted that in recent years, computational hyperspectral imaging techniques based on compressed sensing theory have been developed, such as Coded Aperture Snapshot Spectral Imaging (CASSI). These methods capture encoded, comprehensive information in a single shot and reconstruct high-dimensional data using compressed sensing or deep learning algorithms, significantly reducing acquisition time. However, the individual optical components in these systems, such as dispersive prisms and relay lenses, result in large sizes and difficulties in integration. Depositing pixel-level narrowband filter arrays directly onto a panchromatic camera is a feasible approach to achieving integrated multispectral imaging, but this method requires a trade-off between spatial and spectral resolution and has extremely low optical throughput. Recently, micro- and nanostructures with tunable spectral filtering capabilities, including metasurfaces, metallic lenses, photonic crystals, and Fabry-Perot filters, have played a key role in advancing compact hyperspectral imaging systems. Despite the technological promise of these techniques, they are limited by high manufacturing complexity, resulting in relatively low spatial resolution, typically rarely exceeding 2048×2048 pixels, far below 5 megapixels. Furthermore, these technologies are limited by high production costs and low yield rates, hindering large-scale production and commercial application.

[0050] In contrast, compact lensless imaging systems based on the point spread function (PSF), such as diffraction imaging systems, offer a pathway for the development of low-cost, high-resolution integrated hyperspectral imaging systems. However, these diffraction hyperspectral imaging systems currently suffer from technical problems such as low coding efficiency, significant quantization noise and manufacturing errors, and low spatial resolution.

[0051] Based on this, in order to solve the above problems, the embodiments of this application provide a diffraction partition coding element modeling method, a diffraction partition coding element, a post-coded image acquisition device, a snapshot hyperspectral imaging method, a diffraction partition coding element modeling device for executing the diffraction partition coding element modeling method, a snapshot hyperspectral imaging system for executing the snapshot hyperspectral imaging method, a corresponding electronic device, a computer-readable storage medium, and a program product, which can achieve high-resolution, high-coding-efficiency, and low-noise-affected hyperspectral imaging.

[0052] The following examples will provide a detailed description.

[0053] Based on this, embodiments of this application provide a method for modeling diffraction partition coding elements that can be implemented by a diffraction partition coding element modeling device, see [link to relevant documentation]. Figure 1 The diffraction partitioning coding element modeling method specifically includes the following: Step 100: Divide the design plane of the diffractive optical element into multiple equiangular sectors in the polar coordinate system, generate a diffraction partition coding (DPE) matrix, and rotate and transform the DPE matrix to generate an initial DPE height map.

[0054] It should be noted that the design plane of a diffractive optical element (DOE) refers to the two-dimensional planar region where optical design is performed, corresponding to the aperture of the actual diffractive-partitioned encoding (DPE) element.

[0055] In step 100, polar coordinates (radius r and angle α) are used instead of Cartesian coordinates (x and y) to describe the optical plane, which better reflects the symmetry of a circular diffractive optical element (DOE). An isoangular sector refers to dividing the 360-degree polar coordinate plane into equal angles. A sector-shaped region. For example, when When the value is 31, the angle of each sector is approximately 11.6 degrees.

[0056] The diffraction partition coding (DPE) matrix, which can be abbreviated as DPE matrix, is a matrix of size R× A two-dimensional matrix, where R represents the number of sampling points in the radial direction. This represents the number of bands, which is also the number of sectors. Each element of this matrix is ​​used to store the height distribution parameters of the corresponding sector at a specific radius.

[0057] Correspondingly, rotation transformation refers to converting the DPE matrix in polar coordinates into a two-dimensional height distribution map in Cartesian coordinates through coordinate transformation. This two-dimensional height distribution map is the initial DPE height map.

[0058] Step 200: Using the initial DPE height map as the current full-precision height map, perform at least one iteration of optimization steps on the full-precision height map; wherein, the optimization steps include: quantizing and processing the full-precision height map corresponding to the current iteration based on the learnable depth corresponding to the current iteration to obtain a simulated manufacturing height map; obtaining the point spread function (PSF) corresponding to the wavelength of each sector according to the simulated manufacturing height map to form a PSF set corresponding to the current iteration; convolving the PSF set and the pre-simulated original hyperspectral image, and combining the spectral response and noise of the RGB sensor to generate the encoded image corresponding to the current iteration; inputting the encoded image and the PSF set into the hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs the reconstructed hyperspectral image corresponding to the current iteration; determining the total loss of the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, the learnable depth, and the hyperspectral image reconstruction model according to the total loss.

[0059] It is understandable that a full-precision height map refers to a height distribution map with continuously varying values, normalized to [0,1]. The learnable depth refers to a parameter that is automatically adjusted during optimization, with values ​​corresponding to the equivalent phase from π to 4π; traditional methods typically fix it at 2π. The quantization and manufacturing error processing can be achieved through scaling, uniform quantization (discretizing the continuous height into L steps), and manufacturing error injection, which will be explained in detail in subsequent embodiments. Correspondingly, the height map after simulated manufacturing refers to the height map after the above processing, which more closely approximates the shape of the actually manufactured component.

[0060] The point spread function (PSF) refers to the response of an optical system to a point light source, used to characterize the spatial coding characteristics of the system. A PSF set refers to the complete set of PSFs covering the target band range (e.g., 31 bands), that is, the PSF corresponding to the wavelength of each sector, which is a complete description of the DPE optical coding characteristics. In one or more embodiments of this application, the wavelength of a sector can refer to the center wavelength of the sector.

[0061] The original hyperspectral image refers to the simulated hyperspectral data cube used as training samples. The spectral response of the RGB sensor refers to the sensitivity curves used to describe the red (R), green (G), and blue (B) channels of the RGB sensor to different wavelengths of light. The noise may include sensor noise models such as shot noise and / or readout noise. The encoded image refers to a single RGB image acquired by a system simulating an actual imaging system (such as the encoded image acquisition device mentioned in subsequent embodiments).

[0062] In one or more embodiments of this application, the hyperspectral image reconstruction model refers to a neural network model used to reconstruct an encoded image. This hyperspectral image reconstruction model may include a feature extraction module and a reconstruction network, as detailed in subsequent embodiments. Correspondingly, the reconstructed hyperspectral image refers to the hyperspectral data cube recovered by the hyperspectral image reconstruction model from the encoded image.

[0063] In step 200, the total loss can refer to the weighted sum of reconstruction loss, spectral difference regularization loss, and Fisher information regularization loss. The reconstruction loss is used to measure the pixel-level difference between the reconstructed image and the original image; the spectral difference regularization loss is used to ensure that the PSFs of different bands have sufficient discriminative power; and the Fisher information regularization loss is used to maximize the information transmission capability of the system, which will be described in detail in subsequent embodiments.

[0064] Specifically, backpropagation can be performed based on the total loss to update the full-precision height map, learnable depth, and hyperspectral image reconstruction model. If the model in the optimization step of the current iteration has converged or the current iteration has reached a preset threshold, the iteration stops, and the latest full-precision height map, learnable depth, and hyperspectral image reconstruction model obtained at the current update are sequentially used as the DPE element modeling result data, target learnable depth, and target hyperspectral image reconstruction model. If the model in the optimization step of the current iteration has not converged and the current iteration has not reached the preset threshold, the latest full-precision height map, learnable depth, and hyperspectral image reconstruction model obtained at the current update are sequentially used as the full-precision height map, learnable depth, and hyperspectral image reconstruction model corresponding to the next iteration, and the optimization steps of the next iteration are continued.

[0065] Step 300: After the iteration is completed, the latest full-precision height map is used as the modeling result data of the DPE element, and the latest hyperspectral image reconstruction model is used as the target hyperspectral image reconstruction model. The modeling result data of the DPE element and the target hyperspectral image reconstruction model are output.

[0066] In step 300, the DPE element modeling result data refers to the final optimized full-precision height map, which is used to guide the manufacturing of actual DPE elements; the target hyperspectral image reconstruction model is used for actual online imaging applications.

[0067] In other words, this embodiment achieves a perfect match between diffraction coding element design and reconstruction algorithm through hardware-software co-optimization. Compared with traditional methods, the DPE element designed in this embodiment has higher coding efficiency and better manufacturing robustness. Simultaneously, the reconstruction network is specifically optimized for specific coding characteristics, resulting in a significant improvement in overall system performance. The results obtained in this embodiment can be directly used to manufacture compact snapshot hyperspectral imaging systems, with broad application prospects in industrial inspection, medical diagnosis, and remote sensing monitoring.

[0068] As described above, the diffraction partitioning coding element modeling method provided in this application generates an initial height map through polar coordinate partitioning and rotation transformation, which can provide a good data foundation for subsequent DPE element modeling optimization to effectively improve coding efficiency. By introducing a learnable deep quantization mechanism, it can effectively simulate the real manufacturing process of DPE elements, thereby effectively reducing the impact of quantization noise and manufacturing errors on the DPE element modeling results. Moreover, the DPE element modeling results have a simple structure, which can be realized through a simple manufacturing method, thereby effectively improving spatial resolution. By jointly optimizing the hardware (DPE) and software (hyperspectral image reconstruction model), optimal system performance can be achieved. The end-to-end optimization framework can ensure that the DPE design is directly related to the actual imaging effect, effectively improve the execution efficiency of the complete closed-loop feedback link of encoding-decoding, and ensure that the optical coding mode of the finally modeled DPE is one of the "optimal solutions" in information theory. This can effectively improve the resolution, reliability, and effectiveness of reconstructing hyperspectral images using DPE elements and hyperspectral image reconstruction models.

[0069] To further address the issues of inaccurate initial DPE matrix generation and mismatch with imaging requirements, a diffraction partitioning coding element modeling method is provided in this application embodiment, see [link to relevant documentation]. Figure 2 Step 100 in the diffraction partition coding element modeling method specifically includes the following: Step 110: Obtain imaging band requirement data, which includes spectral range and number of bands.

[0070] If the entire spectral range (e.g., 400 nm - 700 nm) is divided into 31 bands, and each band has a wavelength (e.g., 400 nm, 410 nm, ..., 700 nm) as its center wavelength, this wavelength is used as the representative of that band during design and optimization. That is, in DPE modeling, the first sector is optimized for its wavelength 400 nm, the second sector for its wavelength 410 nm, and so on.

[0071] Step 120: Place the design plane of the diffractive optical element in a polar coordinate system, and divide the design plane into multiple equiangular sectors according to the number of wavebands.

[0072] Specifically, the design plane of the diffractive optical element (DOE) is placed in a polar coordinate system and divided into... An isometric sector.

[0073] Step 130: Determine the wavelength corresponding to each band of each sector according to the spectral range, and generate initial height distribution parameters for each sector to obtain an initialized DPE matrix, wherein the DPE matrix is ​​used to define the initial height distribution parameters of each sector in the radial direction.

[0074] Initial height distribution parameters can be generated for each sector based on its corresponding wavelength, thus obtaining an initialized DPE matrix (size R × 10⁻⁶). The DPE matrix is ​​used to define the initial height distribution parameters of each sector along the radius R direction.

[0075] Specifically, for each sector, the initial height distribution parameters can be generated according to the following formula: in, Represents the initial height distribution parameters. It refers to the substrate thickness (usually 1mm or 2mm thick; a 2mm thickness was used in the demonstration experiment for RIE lithography). For all angles within the nth sector; It represents the nth imaging spectral band (up to 31 bands, ranging from 400-700nm, with a resolution of 10nm). It is a substrate material whose wavelength is predetermined by human intervention. The refractive index; m is a freely selectable value, mainly used to ensure... The item value is a positive number, and r represents the radial distance, indicating the positional encoding within the polar coordinate partitioning framework; f This indicates the system focal length, which is usually 30mm-50mm.

[0076] Step 140: By rotating and transforming the DPE matrix, generate an initial DPE height map in the Cartesian coordinate system corresponding to the current iteration, and make the design plane the current height map plane.

[0077] Specifically, the DPE matrix can be rotated to generate an initial DPE height map in Cartesian coordinates, so that the design plane is currently used as the height map plane.

[0078] As can be seen from the above description, the diffraction partitioning coding element modeling method provided in this application partitions the image based on specific imaging band requirements, which can ensure the design is targeted; the polar coordinate division is more in line with the rotational symmetry characteristics of the optical system, and the rotation transformation can ensure the accuracy and continuity of the height map generation.

[0079] To further address the impact of manufacturing process limitations on design performance, a diffraction partitioning coding element modeling method is provided in this application embodiment, see [link to relevant documentation]. Figure 3 The step 200 of the diffraction partition coding element modeling method, which involves quantizing and processing the full-precision height map corresponding to the current iteration based on the learnable depth of the current iteration to obtain the simulated manufacturing height map, specifically includes the following: Step 211: Based on the learnable depth corresponding to the current iteration round, scale the full-precision height map corresponding to the current iteration round to obtain the scaled height map; wherein, the learnable depth corresponding to the first iteration round is a predefined parameter, and the value of the learnable depth is between π and 4π.

[0080] Specifically, based on the learnable depth d corresponding to the current iteration round total The full-precision heightmap h corresponding to the current iteration round is generated. continuous (x,y) is scaled to physical height to obtain a physically continuous and variable scaled height map h. scaled (x,y); where d is the learnable depth corresponding to the initial iteration round. total The parameters are predefined and their values ​​are between π and 4π, which means that they are allowed to learn within the equivalent phase range of π to 4π.

[0081] Step 212: Determine the step height based on the learnable depth corresponding to the current iteration round and the preset number of quantization steps.

[0082] Specifically, based on the learnable depth d corresponding to the current iteration round total And the step height is calculated based on the preset number of steps L: Δh=d total / L. The number of quantization steps L is pre-set based on manufacturing capabilities, for example, L=8 (3 lithography steps) or L=16 (4 lithography steps).

[0083] Step 213: Perform uniform quantization processing on the scaled height map based on the step height to obtain a quantized height map with quantization noise.

[0084] Specifically, the scaled height map is based on the step height Δh. Uniform quantization is performed to obtain a quantized height map that introduces quantization noise. The quantized height map at this point The continuous surface was transformed into discrete steps; the round operation is non-differentiable, requiring a pass-through estimator during backpropagation, i.e., quantization during forward propagation, and the gradient directly passed to the scaled height map during backpropagation. .

[0085] Step 214: Add manufacturing errors to the quantized height map to simulate errors generated during the lithography and etching processes, and obtain the simulated manufacturing height map corresponding to the current iteration.

[0086] Specifically, in the quantized height map h quantized Adding a manufacturing error N(x,y) to (x,y) to simulate the manufacturing process of photolithography and etching, we obtain the simulated height map after manufacturing, which includes quantization noise and manufacturing error, corresponding to the current iteration. =h quantized (x,y)+N(x,y). Where N(x,y) is a random noise field. It can be modeled as Gaussian white noise with a mean of 0 and a standard deviation of σME (e.g., 10 nm); N(x,y) can also be modeled as more complex spatially correlated noise to better match the error characteristics of a specific process.

[0087] As can be seen from the above description, the diffraction partitioning coding element modeling method provided in this application embodiment can learn to adapt to different manufacturing process requirements; quantization noise and manufacturing error simulation can improve design robustness; and the pass-through estimator can ensure the feasibility of gradient backpropagation.

[0088] To further address the issues of inaccurate PSF calculations and discrepancies with actual physical processes, a diffraction partitioning coding element modeling method is provided in this application embodiment, see [link to relevant documentation]. Figure 3 The step 200 of the diffraction partitioning coding element modeling method, which involves obtaining the point spread function (PSF) corresponding to the wavelength of each sector based on the simulated height map to form the PSF set corresponding to the current iteration, specifically includes the following: Step 221: According to Fresnel's theorem, model the incident light field as a spherical wave to obtain the corresponding incident light field.

[0089] Specifically, according to Fresnel's theorem, the incident light field is modeled as a spherical wave (or, depending on the actual situation, a plane wave). Equation (1) describes a diverging spherical wave: (1) in, This represents the complex amplitude of the incident light field; Represents the Cartesian coordinates on the DPE plane; The wavelength of light is represented by z; the distance from the light source (or scene) to the DPE plane is represented by z, such as 1 meter. i represents the imaginary unit.

[0090] Equation (1) simulates the wavefront shape of light emitted from a point source, after traveling a distance z, and reaching the DPE plane. This wavefront is a sphere. In broader imaging models, a complex scene can be viewed as a collection of countless point sources; therefore, simulating the response (PSF) of a single point source is fundamental to modeling the entire imaging system.

[0091] Step 222: Based on the height map after simulation manufacturing, determine the phase delay of the wavelength corresponding to each sector.

[0092] Specifically, based on Fresnel diffraction theory, for each wavelength λ, the height map is based on the simulated manufacturing process. The phase delay of each wavelength λ is calculated as shown in the following formula (2): (2) in, This represents the phase delay of wavelength λ; This is the difference between the refractive index of glass and that of air.

[0093] Step 223: Determine the outgoing light field obtained by passing the incident light field through the simulated height map based on the phase delay of each wavelength.

[0094] Specifically, based on the phase delay of each wavelength λ, the output light field obtained after the incident light field passes through the simulated height map is calculated, as shown in the following formula (3): (3) in, This represents the complex amplitude of the outgoing light field.

[0095] Step 224: Determine the light field reaching the RGB sensor plane based on the emitted light field and the distance between the RGB sensor plane and the aperture.

[0096] Specifically, the light field reaching the RGB sensor plane is calculated based on the emitted light field. Assuming the distance between the RGB sensor plane and the aperture is f (i.e., the focal length f mentioned above), such as 50 mm, according to Fresnel's law, the light field reaching the RGB sensor plane can be modeled as the following formula (4): (4) in, This represents the complex amplitude of the light field reaching the plane of the RGB sensor; This represents the two-dimensional inverse Fourier transform. yes The abbreviation for is the complex amplitude of the emitted light field; and These represent the frequency variables in the x and y directions, respectively; This represents a two-dimensional Fourier transform.

[0097] Step 225: Determine the square of the light field reaching the sensor plane to obtain the point spread function (PSF) corresponding to each wavelength in the current iteration round, wherein the PSF corresponding to each wavelength constitutes the PSF set corresponding to the current iteration round.

[0098] Specifically, the square of the light field intensity reaching the RGB sensor plane is calculated to obtain the point spread function (PSF) for each wavelength in the current iteration, as shown in formula (5): (5) in, The point spread function (PSF) represents the wavelength λ, and the PSFs corresponding to each wavelength constitute a complete set of PSFs.

[0099] As can be seen from the above description, the diffraction partition coding element modeling method provided in the embodiments of this application is based on accurate optical simulation of Fresnel's theorem; the complete optical field propagation model ensures the accuracy of PSF calculation; and it can provide a reliable optical characteristic description for subsequent coding and reconstruction.

[0100] To further address the issues of unrealistic simulation and mismatch with sensor characteristics during the encoding process, a diffraction partitioning encoding element modeling method is provided in this application embodiment, see [link to relevant documentation]. Figure 3 The step 200 of the diffraction partition coding element modeling method, which involves convolving the PSF set with the pre-simulated original hyperspectral image and combining it with the spectral response and noise of the RGB sensor to generate the coded image corresponding to the current iteration, specifically includes the following: Step 231: Based on the PSF set corresponding to the current iteration round and the original hyperspectral image obtained in advance based on the true value simulation, convolve the hyperspectral image corresponding to each band of each sector with the PSF to obtain the convolved image.

[0101] Specifically, based on the PSF set corresponding to the current iteration round and the original hyperspectral image obtained in advance from the truth simulation. The hyperspectral image of each band is convolved with the corresponding PSF, as shown in formula (6): (6) in For convolution; This is the image after convolution.

[0102] Step 232: Integrate the convolutional image onto the three channels of the RGB sensor, and combine the spectral response and noise of the RGB sensor to obtain the encoded image corresponding to the current iteration round.

[0103] Specifically, the convolved image Integrating to the three channels c of the RGB sensor and combining them with the spectral response of the RGB sensor, the image on the RGB sensor corresponding to the current iteration (a blurred encoded image) can be represented as the simulated encoded image. : (7) in, For noise (used to simulate RGB sensor noise). This is the lower limit of the hyperspectral wavelength that needs to be measured. This represents the upper limit of the hyperspectral wavelength that needs to be measured. This is the spectral sensitivity function of the three channels of the RGB sensor, that is, the response curves of the three channels to different wavelengths.

[0104] As can be seen from the above description, the diffraction partition coding element modeling method provided in this application embodiment can accurately simulate the coding process of the actual imaging system; by considering the sensor spectral response and noise characteristics, it can generate training data that conforms to the real situation.

[0105] To further address the issues of low reconstruction quality and incomplete detail recovery in hyperspectral images, a diffraction partitioning coding element modeling method is provided in this application embodiment, see [link to relevant documentation]. Figure 3 The step 200 of the diffraction partition coding element modeling method, which involves inputting the encoded image and PSF set into a hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs the reconstructed hyperspectral image corresponding to the current iteration, specifically includes the following: Step 241: Input the encoded image corresponding to the current iteration round and the PSF set corresponding to the current iteration round into the feature extraction module in the hyperspectral image reconstruction model, so that the feature extraction module extracts the preprocessed feature map and inputs the preprocessed feature map into the reconstruction network in the hyperspectral image reconstruction model, so that the reconstruction network reconstructs the preprocessed feature map to output the reconstructed hyperspectral image corresponding to the current iteration round.

[0106] See Figure 4 The feature extraction module specifically includes the following components: (1) The first branch is used to input the encoded image into the multi-feature spectral attention network MFSA-Net in the reconstruction network, the first deconvolution unit in the second branch and the second deconvolution unit in the third branch respectively; (2) The second branch is used to map the PSF set to the spectral feature space based on the first RGB mapping unit, and to deconvolve the PSF set mapped to the spectral feature space and the encoded image based on the first deconvolution unit, so as to output the spatial features corresponding to the encoded image; (3) The third branch is used to extract the restoration space features corresponding to the encoded image based on multiple convolution units, and to map the restoration space features to the spectral feature space based on the second RGB mapping unit, and then to deconvolve the restoration space features mapped to the spectral feature space and the encoded image based on the second deconvolution unit, so as to output the spectral reconstruction features corresponding to the encoded image. (4) Dual-branch feature fusion unit, used to fuse the spatial features and the spectral reconstruction features to obtain a preprocessed feature map, and input the preprocessed feature map into the Res-Unet in the reconstruction network.

[0107] Specifically, the encoded image corresponding to the current iteration round (It can also be written as) The PSF set corresponding to the current iteration round is input into the feature extraction module in the reconstruction network, so that the feature extraction module passes through the deconvolution module DECONV and a series of convolution modules CONV to extract the preprocessed feature map input into the neural network.

[0108] The deconvolution module is calculated using the following formula (8): (8) in, For measured values, Here are the point spread functions for each band. For inverse Fourier transform, This is a Fourier transform, where n is noise, typically taken as... .

[0109] See Figure 4The reconstruction network specifically includes a multi-head spectral attention network (MFSA-Net), a Res-Unet, and an addition operation module. The preprocessed feature map is input into a reconstruction network composed of MFSA-Net and Res-Unet, which outputs the reconstructed hyperspectral image corresponding to the current iteration. The reconstruction network includes Multi-Feature Spectral Attention-Net (MFSA-Net) and Res-UNet (residual-Unet). MFSA-Net adopts a U-shaped architecture, consisting of Multi-Level Embedding (MLE) units, Spectral Attention (SA) units (including Multi-Head Spectral Attention (MHSA) units and Multi-Scale Feed-Forward Network (MS FFN) units), and Multi-Channel Residual Fusion (MCRF) units. Res-UNet is used to reconstruct the hyperspectral data cube from the PSF feature extraction module. The outputs of the two sub-networks are combined through an addition operation and then passed through convolutional layers to generate the final reconstructed hyperspectral image.

[0110] Specifically, see Figure 5 The multi-head spectral attention network MFSA-Net specifically includes the following components: (1) The first convolutional unit is used to perform preliminary feature extraction and channel dimension transformation on the encoded image to obtain a first feature map, and input the first feature map into the multi-level embedding MLE unit and the feature fusion unit respectively; (2) Multi-level embedding MLE unit, used to achieve multi-scale feature embedding of the first feature map through multi-level convolution and upsampling operations to obtain the second feature map; (3) The first spectral attention SA unit is used to extract the spectral and spatial features corresponding to the second feature map through the spectral attention mechanism to obtain the third feature map, and input the third feature map into the second multi-channel residual fusion MCRF unit and the first downsampling unit respectively; (4) The first downsampling unit is used to reduce the spatial resolution of the third feature map and increase the channel dimension to obtain the fourth feature map; (5) The second spectral attention SA unit is used to extract the spectral and spatial features corresponding to the fourth feature image through the spectral attention mechanism to obtain the fifth feature image, and input the fifth feature image into the first MCRF unit and the second downsampling unit respectively; (6) A second downsampling unit is used to reduce the spatial resolution of the fifth feature map and increase the channel dimension to obtain a sixth feature map; (7) The third spectral attention SA unit is used to extract the spectral and spatial features corresponding to the sixth feature map through the spectral attention mechanism to obtain the seventh feature map; (8) The first upsampling unit is used to restore the spatial resolution of the seventh feature map and reduce the channel dimension to obtain the eighth feature map; (9) A first multi-channel residual fusion MCRF unit is used to fuse the fifth feature map and the eighth feature map to obtain a ninth feature map; (10) The second convolutional unit is used to extract features from the ninth feature map to obtain the tenth feature map; (11) The fourth spectral attention SA unit is used to extract the spectral and spatial features corresponding to the tenth feature map through the spectral attention mechanism to obtain the eleventh feature map; (12) The second upsampling unit is used to restore the eleventh feature map to the original input resolution to obtain the twelfth feature map; (13) A second multi-channel residual fusion MCRF unit is used to fuse the third feature map and the twelfth feature map to obtain a thirteenth feature map; (14) The third convolutional unit is used to extract features from the thirteenth feature map to obtain the fourteenth feature map; (15) The fifth spectral attention SA unit is used to extract the spectral and spatial features corresponding to the fourteenth feature map through the spectral attention mechanism to obtain the fifteenth feature map; (16) Feature fusion unit, used to perform residual connection on the first feature map and the fifteenth feature map to obtain the sixteenth feature map; (17) The fourth convolutional unit is used to map the sixteenth feature map to the target spectral dimension to obtain a reconstructed hyperspectral map, and input the reconstructed hyperspectral map into the summation operation module.

[0111] The Res-Unet is used to reconstruct the hyperspectral data cube corresponding to the preprocessed feature map, and the hyperspectral data cube is input into the summation operation module. The summation operation module is used to perform addition and convolution on the reconstructed hyperspectral map and the hyperspectral data cube to output the reconstructed hyperspectral image corresponding to the current iteration round.

[0112] Among them, the first, second, third, fourth, and fifth spectral attention SA units can all employ the same type of spectral attention SA unit, such as... Figure 6As shown, the spectral attention SA unit is equipped with a multi-head spectral attention MHSA unit and a multi-scale feedforward network MS FFN unit.

[0113] The first SA unit and the first downsampling unit constitute the first stage of the encoder; the second SA unit and the second downsampling unit constitute the second stage of the encoder; the third SA unit constitutes the bottleneck part; and the first upsampling unit, the second convolutional unit, the fourth SA unit, the second upsampling unit, the second MCRF unit, the third convolutional unit, and the fifth SA unit constitute the decoder part.

[0114] See Figure 7 The multi-level embedding MLE unit consists of multiple convolutional blocks (i.e., convolutional units), upsampling layers (i.e., upsampling units), and downsampling layers (i.e., downsampling units). Both the first and second multi-channel residual fusion MCRF units can employ the same type of multi-channel residual fusion MCRF unit; see [link to relevant documentation]. Figure 8 The multi-channel residual fusion (MCRF) unit replaces skip connections, achieving the fusion of feature maps at the same scale. The MCRF unit consists of two branches: one branch directly concatenates two sets of feature maps and extracts features through convolution, while the other branch uses a residual structure to extract deep features. The features from the two branches are then fused together.

[0115] See Figure 9 Multi-head self-attention (MHSA) units calculate the multi-head self-attention of spectral channels; see [link to documentation]. Figure 10 The multi-scale feedforward network (MS FFN) unit uses convolutional kernels of different sizes to extract multi-scale features. Specifically, in... Figures 4 to 10 middle, , and These are key hyperparameters that control the different depths and feature extraction capabilities of the MESA-Net network. Their values ​​(e.g.) =2) represents the consecutive use of two SA units in a specific processing stage, enabling the network to learn more complex and abstract feature representations at each layer, ultimately working together to improve the reconstruction quality of hyperspectral images. H represents height, which refers to the number of pixels or feature points in the vertical direction of the image or feature map; W represents width, which refers to the number of pixels or feature points in the horizontal direction of the image or feature map; C represents the number of channels, which refers to the dimension of the feature vector at each spatial location (H, W).

[0116] As can be seen from the above description, the diffraction partitioning coding element modeling method provided in this application embodiment can improve the spatial effect and spectral imaging effect of diffraction imaging through neural network algorithm, alleviate the inherent blurring problem of diffraction imaging, fully utilize different feature information through multi-branch feature extraction, effectively capture spectral correlation through spectral attention mechanism, maintain multi-scale spatial information through U-shaped network structure, promote gradient flow and feature reuse through residual connection, and improve reconstruction accuracy and robustness through multi-path fusion.

[0117] To further address the problem of a single optimization objective and the tendency to get trapped in local optima, this application provides a diffraction partitioning coding element modeling method, see [link to relevant documentation]. Figure 3 The step 200 of the diffraction partition coding element modeling method, which involves determining the total loss of the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, the learnable depth, and the hyperspectral image reconstruction model based on the total loss, specifically includes the following: Step 251: Based on the PSF set corresponding to the current iteration round, obtain the average absolute difference value of the point spread function PSF corresponding to the wavelength of each sector by weighted summation, and map the average absolute difference value through the activation function to obtain the spectral difference regularization loss corresponding to the current iteration round.

[0118] Specifically, based on the PSF set corresponding to the current iteration round, the average absolute difference of PSFs among all band pairs can be calculated by weighted summation, and the resulting average absolute difference can be mapped using the Sigmoid function to obtain the spectral difference regularization loss. : Wherein, the scaling factor of the Sigmoid function , can be adjusted as needed; m is the weighted average spectral difference degree; n represents the total number of spectral bands; and n and The index represents the spectral band; k represents the index variable; (x, y) are the Cartesian coordinates of the DPE plane. For the i-th wavelength, The nth wavelength is represented by the mean, which is the average pixel intensity. The point spread function can also be written as .

[0119] Step 252: Obtain the corresponding Fisher information matrix based on the PSF set corresponding to the current iteration round and determine the sum of all elements in the Fisher information matrix, and determine the Fisher information regularization loss corresponding to the current iteration round based on the sum of all elements in the Fisher information matrix.

[0120] Specifically, the Fisher Information Matrix (FIM) is calculated based on the PSF set corresponding to the current iteration round and the imaging model corresponding to the diffraction imaging coding (i.e., the mathematical model describing how a hyperspectral image is encoded into RGB measurements, including optical coding to generate the encoded image). Calculate the sum A of all elements in the Fisher Information Matrix (FIM); Calculate the Fisher information regularization loss based on A. : And, step 253: determine the mean absolute error between the reconstructed hyperspectral image and the original hyperspectral image, so as to use the mean absolute error as the reconstruction loss of the reconstruction network in the hyperspectral image reconstruction model.

[0121] Specifically, the reconstructed hyperspectral image is compared with the original hyperspectral image. Mean absolute error between This can be taken as the reconstruction loss, or simply referred to as L1.

[0122] Step 254: Determine the total loss for the current iteration based on the spectral difference regularization loss, the Fisher information regularization loss, and the reconstruction loss corresponding to the current iteration round.

[0123] Specifically, calculate the total loss corresponding to the current iteration round: Here, α and β are hyperparameters used to balance the three sometimes competing objectives of reconstruction accuracy, spectral difference, and information content.

[0124] Step 255: Backpropagate the total loss corresponding to the current iteration round to update the weights of the reconstruction network in the full-precision height map, learnable depth, and hyperspectral image reconstruction model based on the total loss, thereby obtaining the latest full-precision height map, learnable depth, and hyperspectral image reconstruction model in sequence.

[0125] Specifically, the total loss corresponding to the current iteration is backpropagated, and the full-precision heightmap h is updated based on this total loss. continuous (x,y), learnable depth d total And the weights of the reconstruction network in the hyperspectral image reconstruction model, so as to obtain the full-precision height map h corresponding to the next iteration. continuous (x,y), learnable depth d total And hyperspectral image reconstruction models.

[0126] As can be seen from the above description, the diffraction partition coding element modeling method provided in this application has a multi-objective loss function that can balance different optimization requirements; spectral difference regularization can ensure PSF discrimination; Fisher information regularization can improve system information capacity; and reconstruction loss can guarantee image quality. Therefore, joint optimization can achieve optimal global performance.

[0127] From a software perspective, this application also provides an apparatus for performing all or part of the diffraction partition coding element modeling method, see [link to relevant documentation]. Figure 11 The diffraction partition coding element modeling device specifically includes the following components: The initialization module 10 is used to divide the design plane of the diffractive optical element into multiple equiangular sectors in the polar coordinate system, generate a diffraction partition coding (DPE) matrix, and rotate and transform the DPE matrix to generate an initial DPE height map.

[0128] The iterative training module 20 is used to take the initial DPE height map as the current full-precision height map and perform at least one iteration of optimization steps on the full-precision height map. The optimization steps include: quantizing and processing the full-precision height map corresponding to the current iteration based on the learnable depth corresponding to the current iteration to obtain a simulated manufacturing height map; obtaining the point spread function (PSF) corresponding to the wavelength of each sector based on the simulated manufacturing height map to form a PSF set corresponding to the current iteration; convolving the PSF set with the pre-simulated original hyperspectral image and combining it with the spectral response and noise of the RGB sensor to generate an encoded image corresponding to the current iteration; inputting the encoded image and the PSF set into a hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs a reconstructed hyperspectral image corresponding to the current iteration; determining the total loss of the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, the learnable depth, and the hyperspectral image reconstruction model based on the total loss.

[0129] The result output module 30 is used to, after the iteration is completed, take the latest full-precision height map as the DPE element modeling result data and the latest hyperspectral image reconstruction model as the target hyperspectral image reconstruction model, and output the DPE element modeling result data and the target hyperspectral image reconstruction model.

[0130] The embodiments of the diffraction partition coding element modeling apparatus provided in this application can be used to execute the processing flow of the embodiments of the diffraction partition coding element modeling method described above. Its functions will not be repeated here, but can be referred to the detailed description of the embodiments of the diffraction partition coding element modeling method described above.

[0131] The diffraction partition coding element modeling device can perform the modeling of diffraction partition coding elements in either a server or a client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations in this regard. If all operations are performed in the client device, the client device may further include a processor for the specific processing of diffraction partition coding element modeling.

[0132] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0133] The server and the client device can communicate using any suitable network protocol, including those not yet developed as of the date of this application. Such network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Furthermore, such network protocols may also include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer Protocol) protocols used on top of the aforementioned protocols.

[0134] As described above, the diffraction partitioning coding element modeling device provided in this application generates an initial height map through polar coordinate partitioning and rotation transformation, which can provide a good data foundation for subsequent DPE element modeling optimization to effectively improve coding efficiency. By introducing a learnable deep quantization mechanism, it can effectively simulate the real manufacturing process of DPE elements, thereby effectively reducing the impact of quantization noise and manufacturing errors on the DPE element modeling results. Moreover, the DPE element modeling results have a simple structure, which can be realized through a simple manufacturing method, thereby effectively improving spatial resolution. By jointly optimizing the hardware (DPE) and software (hyperspectral image reconstruction model), optimal system performance can be achieved. The end-to-end optimization framework can ensure that the DPE design is directly related to the actual imaging effect, effectively improve the execution efficiency of the complete closed-loop feedback link of encoding-decoding, and ensure that the optical coding mode of the finally modeled DPE is one of the "optimal solutions" in information theory. This can effectively improve the resolution, reliability, and effectiveness of reconstructing hyperspectral images using DPE elements and hyperspectral image reconstruction models.

[0135] To further address the limitations of traditional diffraction elements in terms of performance and manufacturing difficulty, this application provides an embodiment of a diffraction partitioned coding element (DPE) based on the diffraction partitioned coding element modeling method provided in the aforementioned embodiments. The DPE specifically includes: an optical substrate; and a micro / nano structure formed on the optical substrate. The surface of the micro / nano structure has a height distribution based on polar coordinate partitioning. The polar coordinate partitioning divides the optical aperture into multiple equiangular sectors, and the height distribution of each sector is formed by rotating around the polar coordinate center. The height distribution is defined by the DPE modeling result data. The DPE modeling result data is obtained using the diffraction partitioned coding element modeling method provided in the aforementioned embodiments, ensuring that the total depth of the height distribution is within the range of 0.768 μm to 3.072 μm. The DPE provided in this embodiment features a polar coordinate partitioning structure that improves coding efficiency; an optimized height distribution with optimal performance; a total depth range that adapts to manufacturing process requirements; and a rotationally symmetric structure that facilitates processing and integration.

[0136] To further address the issues of large size and high cost in traditional hyperspectral imaging systems, based on the diffraction partition coding (DPE) element provided in the aforementioned embodiments, this application also provides an embodiment of a coded image acquisition device, see [link to embodiment]. Figure 12 The encoded image acquisition device specifically includes the following components: RGB sensor (such as) Figure 12The image shows a camera and a diffraction partition coding element as mentioned in the preceding embodiments. The diffraction partition coding element is installed at a preset distance in front of the RGB sensor and is configured to perform phase modulation on the incident light field from the target scene. The preset distance includes 50 mm. The RGB sensor is configured to receive the light signal phase-modulated by the diffraction partition coding element and convert the light signal into an encoded image of the target scene. That is, the designed DPE element is installed 50 mm in front of the RGB sensor, the distance between the experimental scene and the hyperspectral system is fixed at 1 m, and the imaging experiment is carried out under halogen lamp illumination.

[0137] Specifically, based on the DPE modeling results, corresponding DPE elements can be fabricated using micro-nano fabrication techniques such as photolithography and reactive ion etching. This application provides a post-encoded image acquisition device with a compact design for easy system integration; it is plug-and-play and compatible with existing sensors; a fixed installation distance ensures imaging stability; and it achieves hyperspectral imaging at low cost.

[0138] To further address the issues of slow hyperspectral imaging speed and inability to be applied in real time, based on the encoded image acquisition device mentioned in the foregoing embodiments, this application also provides an embodiment of a snapshot hyperspectral imaging method, see [link to embodiment]. Figure 13 The snapshot hyperspectral imaging method specifically includes the following: Step 400: Obtain the encoded image of the target scene acquired by the encoded image acquisition device, and perform optical simulation on the modeling result data of the DPE element to obtain the PSF set corresponding to the diffraction partition encoding element.

[0139] It is understood that the encoded image acquisition device in step 400 can be the encoded image acquisition device mentioned in the foregoing embodiments.

[0140] Step 500: Input the encoded image of the target scene and the PSF set into the target hyperspectral image reconstruction model so that the target hyperspectral image reconstruction model outputs the reconstructed hyperspectral image of the target scene.

[0141] Specifically, the DPE element can be installed in front of a commercial RGB sensor to build a compact imaging system; a single snapshot of the target scene is taken to capture an encoded image (a blurred RGB image); the encoded image and the point spread function (PSF) corresponding to the DPE modeling result data are input into a pre-trained target reconstruction network so that the target reconstruction network outputs the corresponding reconstructed hyperspectral image.

[0142] To further illustrate the above embodiments, this application also provides an ultra-high-definition snapshot hyperspectral imaging system and method based on learning-based diffraction partitioning, which is applicable to fields such as remote sensing, industrial and agricultural observation that require hyperspectral information. It aims to solve the technical problems of low coding efficiency, large influence of quantization noise and manufacturing errors, and low spatial resolution in existing diffraction hyperspectral imaging systems, and provides an ultra-high-definition snapshot hyperspectral imaging system and method based on learning-based diffraction partitioning to achieve high-resolution, high coding efficiency, and low noise impact hyperspectral imaging.

[0143] See Figure 14 The proposed learning-based diffraction partitioning encoding joint optimization framework includes diffraction partitioning modes, regularization methods, learnable structure depth, a diffraction partitioning hyperspectral system, a diffraction imaging encoder, an imaging reconstruction decoder, and a PSF feature extraction and reconstruction network. The hardware of the ultra-high-definition snapshot hyperspectral imaging system based on learning-based diffraction partitioning comprises two parts: a diffraction partitioned encoding (DPE) element and an RGB sensor. The system combines the DPE and the RGB camera via optomechanical components. When light of different wavelengths passes through the DPE, it is modulated and encoded by the DPE and then received by the RGB camera. The RGB camera receives the spectral scene encoded by the DPE.

[0144] The core improvements involved in the ultra-high resolution snapshot hyperspectral imaging method based on learned diffraction partitioning are as follows: 1. Diffraction Partition Encoding Element (DPE): A polar coordinate partitioning scheme is adopted to convert the height map plane into polar coordinates and divide it into isoangular sectors that match the spectral bands. Each sector uses the height map distribution of the spectral response corresponding to the sensitive wavelength as the initial height map. By adaptively optimizing the spectral response of each sector, multi-channel coding with spatially differentiated distribution is achieved. This component is manufactured using advanced photolithography and reactive ion etching (RIE) technologies. It has a diameter of 12.7 mm (this diameter is comparable to the size of the card used with the DPE and should be larger than the effective aperture to ensure the card does not obstruct the effective aperture; its typical range is "effective aperture < diameter ≤ 15 mm"). It has an effective aperture of 5.12 mm (the effective aperture can be set from 3 mm to 6 mm; when the effective aperture is smaller, the image is clearer, but the modulation effect for hyperspectral information weakens and the light transmission decreases; when the effective aperture is larger, the image is blurrier, but the modulation effect for hyperspectral information strengthens and the light transmission increases). The total depth is 1.758 μm (the total depth is limited to within 0.768 μm - 3.072 μm).

[0145] 2. The optimization design process of DPE includes: (1) Diffraction partition coding design: During the optimization process, the height distribution of the diffractive optical element is quantized into a dimension of... The matrix (this size can be changed according to the resolution requirements of the actual DPE produced). or The specific DPE matrix generation process is as follows: Figure 15 First, let's set one. A matrix of dimensions, such as Figure 14 As shown, it is named the DPE matrix. Then, according to... Figure 14 As shown, the angular range of the polar coordinate plane of the diffraction coding element is divided into n0=31 equally spaced sectors (each sector is n0=31). When the imaging wavelength is 400-700nm, there are 31 wavelength bands; when the imaging wavelength is 400-1000nm, there are 61 wavelength bands, forming n0 sectors. (The first sector corresponds to a 400nm response, the second sector corresponds to a 410nm response, and so on). An initial distribution is set for each sector, and then, according to... Figure 15 As shown, this polar coordinate distribution is mapped to Cartesian coordinates. This generates a DPE height distribution, which can then be used for the next step of the calculation. For example... Figure 16 As shown, this yields an initial DPE height map.

[0146] (2) Learnable structural depth quantization: In existing technologies, the total depth of diffractive optical elements (i.e., the deepest step depth) is usually set to a custom value, such as 1.536 μm, to meet the requirements of their maximum imaging wavelength (e.g., 700 nm). Phase distribution. Then, the height map steps are quantized (e.g., 4-step, 8-step, 16-step quantization) before final manufacturing. This is because diffractive optical elements can only be quantized to the power of 2 in photolithography, a limitation imposed on them by the manufacturing process. This technique defines the total depth of the structure as a trainable variable, uniformly quantizes the full-precision height map, and introduces manufacturing errors to generate quantized diffractive encoded elements containing quantization noise (QN) and manufacturing errors (ME) to more accurately represent the actual scene. This learnable total depth is changed from the previously fixed... The phase distribution is changed to It can learn to adapt to QN and ME, better aligning with real physical processing constraints.

[0147] (3) Regularization Constraints and Prior Knowledge Extraction: Two types of regularization prior knowledge constraints are introduced to promote DPE optimization. The first type of constraint It is mainly used to calculate the difference between PSFs of different spectral bands. By controlling and amplifying the difference between PSFs of different spectral bands and assigning different weights according to the distance between PSF bands, spectral correlation can be taken into account.

[0148] The second type of constraint By incorporating the Fisher information matrix (FIM) regularization term, a priori constraints are imposed on the PSF. The specific formula is as follows: , where A is the sum of the parameters of the Fisher matrix, and its calculation process is the existing method.

[0149] The system's total loss function is , This is the Mean Absolute Error (MAE).

[0150] (4) Hyperspectral Image Reconstruction: The reconstruction network convolves the real image with the PSF, then performs RGB mapping and noise injection to generate the encoded image. First, the convolution process of the PSF with the real image and the generation process of the encoded image are derived: For a beam of light incident from depth z, according to Fresnel's theorem, it can be modeled as formula (1); the light after passing through a designed DOE is represented by formula (3); where the phase delay... It is expressed as formula (2); the light passing through the DOE aperture reaches the sensor plane. Assuming the distance between the sensor plane and the aperture is d, according to Fresnel's law, the light reaching the sensor can be modeled as formula (4); thus, the PSF of the system can be expressed as formula (5).

[0151] After generating the PSF, the original hyperspectral image PSF modulation can be used, as in equation (6). Combining the spectral response of the sensor, the image on the sensor can be expressed as equation (7).

[0152] The encoded image and PSF are input into the hyperspectral image reconstruction model. The hyperspectral image reconstruction model consists of two parts: a PSF feature extraction module and a reconstruction network. First, the encoded image and point spread function are input into the PSF feature extraction module, which extracts the features input into the neural network through deconvolution (DECONV) and a series of convolutions (CONV). The deconvolution module is calculated based on formula (8). The encoded image is directly input into the designed multi-feature spectral attention network (MFSA-Net). MFSA-Net is a U-shaped structure consisting of an encoder, a bottleneck, and a decoder. MFSA-Net is constructed from a multi-level embedded attention (MLE) module and a multi-channel residual fusion (MCRF) module.

[0153] First, the encoded image After a 1×1 convolution, it can be represented as The data is then input into the MLE unit, which consists of multiple convolutional blocks (i.e., convolutional units), upsampling layers (i.e., upsampling units), and downsampling layers (i.e., downsampling units). The output feature map is represented as... Secondly, go through One SA unit, one downsampling unit, There are one SA unit and one downsampling unit. These units constitute the encoder. The downsampling unit is a convolutional layer with a stride of 3×3. The downsampling unit downsamples the feature map and doubles the number of channels. After the last downsampling unit, the feature map is represented as... .third, After by The bottleneck consisting of SA modules is represented as .

[0154] Next, a decoder was designed to form a U-shaped structure together with the encoder and decoder. The decoder consists of SA units, upsampling units, and MCRF units. The MCRF units replace skip connections, achieving the fusion of feature maps at the same scale. Each MCRF unit consists of two branches: one branch directly concatenates two sets of feature maps and extracts features through convolution, while the other branch uses a residual structure to extract deep features. The features from the two branches are then fused together. After the decoder, the feature map is represented as... Finally, by... and The summation yields the output, represented as Then, the hyperspectral image is reconstructed by passing it through a 1×1 convolutional layer.

[0155] The SA unit consists of two layers of normalization, a multi-head spectral attention (MHSA) unit, and a multi-scale feedforward network (MS FFN) unit. The MHSA unit computes the multi-head self-attention of the spectral channels, while the MS FFN unit uses convolutional kernels of different sizes to extract multi-scale features.

[0156] 3. RGB sensor: Used to capture optical signals encoded by DPE elements.

[0157] Based on this, the application examples in this application provide the following improvements: 1. Diffraction Partition Coding (DPE) Strategy.

[0158] 2. Can learn deep quantization methods.

[0159] 3. The Joint Optimization Framework (LDEC, Learned Diffractive-Partitioned Encoded Co-Optimization): This framework integrates coding mode distribution, hardware optimization, regularization techniques, and prior information to achieve joint optimization of coding modes. By introducing two types of regularized prior knowledge constraints, it controls and amplifies the differences between PSFs in different spectral bands. Furthermore, it applies prior constraints to the PSFs using Fisher information matrix regularization terms, and utilizes a deconvolution-based feature extraction method to extract prior knowledge, significantly improving the reconstruction quality of hyperspectral images.

[0160] 4. Combination of High Resolution and Compactness: An integrated diffraction-partitioned hyperspectral imaging system was developed, achieving a resolution exceeding 20 megapixels (4504×4504 pixels). It enables hyperspectral imaging with 31 spectral channels within the 400-700nm range, achieving a spatial resolution of 0.8mm at a distance of 1 meter. Simultaneously, it boasts an ultra-compact size (29mm×29mm×80mm) and light weight (100g). This system achieves the highest resolution among existing diffraction hyperspectral imaging systems while maintaining such a high number of spectral channels.

[0161] 5. Plug-and-play compatibility: DPE components can be directly integrated with commercial sensors without modifying their original structure, offering plug-and-play scalability. They can be seamlessly and flexibly integrated into most existing imaging systems without modifying their original components, significantly expanding their hyperspectral data acquisition capabilities and paving the way for the widespread application of high-resolution hyperspectral imaging in real-world scenarios.

[0162] To further illustrate the beneficial effects of the application examples of this application, the embodiments of this application also provide the following experimental content: The experimental setup used halogen lamps for illumination, and a scene (such as a resolution panel and various types of objects) was placed 1 meter away from the proposed camera system. The corresponding experimental results are as follows: (1) See Figure 17 The spectral imaging results data of the hyperspectral system shown are as follows: Figure 17 Part A of the diagram shows the point spread functions for each band obtained using the diffraction partition coding element modeling method of this application. Figure 17 Part B of the diagram shows the cross-correlation coefficients between point spread functions. It can be seen that the correlation coefficients between point spread functions of different bands are all below 0.4. Figure 17 Part C of the paper demonstrates an imaging experiment of the system on a resolution panel. This application can achieve an imaging result of 0.84 mm at a distance of 1 m. Figure 17Part D of the diagram shows the algorithm comparison of the system, including comparison curves with various algorithms. The algorithms compared are: Res-Unet, MST, and CAFormer. Among them, Res-Unet stands for Residual U-shaped Network; MST stands for Multispectral Transformer, which is a neural network architecture based on a self-attention mechanism; CAFormer is a channel attention Former, where Former is short for Transformer.

[0163] (2) See Figure 18 The system prototype construction and performance verification experimental results data shown are as follows: Figure 18 Part A in the diagram represents the hyperspectral camera prototype system. The hyperspectral system consists of a DPE element and an RGB sensor. Figure 18 Part B in the diagram represents the DPE and the reference diffraction coding element. Figure 18 Part C in the figure represents the results of optical microscopy (OM) and scanning electron microscopy (SEM) imaging. Figure 18 The D part in the figure represents the comparative experimental results. Synthetic RGB of reconstructed and example spectral images in the 400-700 nm range is given. Figure 18 The E section represents the spatial and spectral results. Spatial results for DPE and Ref. are shown. A spectral comparison of two selected points is displayed at the bottom. Correlation coefficients between the spectral curves and the ground truth are given.

[0164] (3) See Figure 19 The application experimental data shown are as follows: Figure 19 Part A in the diagram represents the classification results for green peppers. The images on the left side of Part A show RGB images and example spectral images with intensity variations. The graph on the right side of Part A shows a spectral comparison of two selected regions and their point classifications. Figure 19 Part B in the diagram represents the spatial and spectral results of traditional Chinese watercolor painting. The images on the left side of Part B show composite RGB images at different resolutions (1 million pixels, 4 million pixels, and 20 million pixels), while the chart on the right side of Part B shows a spectral comparison of three selected points in the composite RGB images.

[0165] (4) See Figure 20 The dynamic hyperspectral imaging verification data shown is from a real-world scenario. Figure 20 Part A in the graph represents the spectral imaging of the scene. The reconstructed composite RGB is shown on the left. In the graph on the right, the colored lines represent the spectral values ​​of the points, and the solid black lines represent the ground truth values ​​for the selected points. Figure 20 Part B in the figure represents the correlation coefficient, which illustrates the spectral correlation between two different materials and the ground conditions over time and from different perspectives. Figure 20 Section C in the image represents spectral video. Example spectral images from different frames are shown.

[0166] This application also provides an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method mentioned in the above embodiments. The processor and memory can be connected via a bus or other means, taking a bus connection as an example. The receiver can be connected to the processor and memory via wired or wireless means.

[0167] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0168] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method in the above method embodiments.

[0169] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0170] The one or more modules are stored in the memory, and when executed by the processor, they execute the diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method in the implementation embodiment.

[0171] In some embodiments of this application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected via a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0172] As one implementation method, the functions of the receiver and transmitter in this application can be implemented by transceiver circuits or dedicated transceiver chips, and the processor can be implemented by dedicated processing chips, processing circuits or general-purpose chips.

[0173] As another implementation approach, the server provided in this application embodiment can be implemented using a general-purpose computer. That is, the program code implementing the processor, receiver, and transmitter functions is stored in memory, and the general-purpose processor implements the processor, receiver, and transmitter functions by executing the code in memory.

[0174] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0175] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned diffraction partition coding element modeling method and / or snapshot hyperspectral imaging method.

[0176] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave.

[0177] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0178] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0179] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to the embodiments of this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for modeling diffraction partitioned coding elements, characterized in that, include: The design plane of the diffractive optical element is divided into multiple equiangular sectors in the polar coordinate system to generate a diffraction partition coding (DPE) matrix. This DPE matrix is ​​then rotated and transformed to generate an initial DPE height map. The initial DPE height map is used as the current full-precision height map, and an optimization step of at least one iteration is performed on the full-precision height map. The optimization step includes: quantizing and processing the full-precision height map corresponding to the current iteration based on the learnable depth corresponding to the current iteration to obtain a simulated manufacturing height map; obtaining the point spread function (PSF) corresponding to the wavelength of each sector based on the simulated manufacturing height map to form a PSF set corresponding to the current iteration; convolving the PSF set with the pre-simulated original hyperspectral image, and combining the spectral response and noise of the RGB sensor to generate the encoded image corresponding to the current iteration; inputting the encoded image and the PSF set into a hyperspectral image reconstruction model so that the hyperspectral image reconstruction model outputs the reconstructed hyperspectral image corresponding to the current iteration; determining the total loss of the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, the learnable depth, and the hyperspectral image reconstruction model based on the total loss. After the iteration is completed, the latest full-precision height map is used as the modeling result data of the DPE element, and the latest hyperspectral image reconstruction model is used as the target hyperspectral image reconstruction model. The modeling result data of the DPE element and the target hyperspectral image reconstruction model are then output.

2. The diffraction partitioning coding element modeling method according to claim 1, characterized in that, The process of dividing the design plane of the diffractive optical element into multiple equiangular sectors in polar coordinates, generating a diffraction partition coding (DPE) matrix, and then rotating and transforming this DPE matrix to generate an initial DPE height map includes: Acquire imaging band requirement data, which includes spectral range and number of bands; The design plane of the diffractive optical element is placed in a polar coordinate system, and the design plane is divided into multiple equiangular sectors according to the number of wavebands; Based on the spectral range, the wavelength corresponding to each band of each sector is determined, and initial height distribution parameters are generated for each sector to obtain an initialized DPE matrix, wherein the DPE matrix is ​​used to define the initial height distribution parameters of each sector in the radial direction. By rotating and transforming the DPE matrix, an initial DPE height map in the Cartesian coordinate system corresponding to the current iteration is generated, and the design plane is used as the height map plane.

3. The diffraction partition coding element modeling method according to claim 1, characterized in that, The process of quantizing and processing manufacturing errors on the full-precision heightmap corresponding to the current iteration based on the learnable depth of the current iteration to obtain the simulated manufacturing heightmap includes: Based on the learnable depth corresponding to the current iteration round, the full-precision height map corresponding to the current iteration round is scaled to obtain the scaled height map; wherein, the learnable depth corresponding to the first iteration round is a predefined parameter, and the value of the learnable depth is between π and 4π; The step height is determined based on the learnable depth corresponding to the current iteration round and the preset number of quantization steps; The scaled height map is uniformly quantized based on the step height to obtain a quantized height map with quantization noise. Add manufacturing errors to the quantized height map to simulate errors generated during the lithography and etching processes, and obtain the simulated manufacturing height map corresponding to the current iteration.

4. The diffraction partitioning coding element modeling method according to claim 1, characterized in that, The step of obtaining the point spread function (PSF) corresponding to the wavelength of each sector based on the simulated height map to form the PSF set corresponding to the current iteration includes: According to Fresnel's theorem, the incident light field is modeled as a spherical wave to obtain the corresponding incident light field; Based on the height map after the simulation manufacturing, the phase delay of the wavelength corresponding to each sector is determined. Based on the phase delay of each wavelength, the output light field obtained after the incident light field passes through the simulated height map is determined: The light field reaching the RGB sensor plane is determined based on the emitted light field and the distance between the RGB sensor plane and the aperture. The square of the light field reaching the sensor plane is determined to obtain the point spread function (PSF) for each wavelength corresponding to the current iteration round. The PSFs for each wavelength constitute the PSF set for the current iteration round.

5. The diffraction partitioning coding element modeling method according to claim 1, characterized in that, The step of convolving the PSF set with the pre-simulated original hyperspectral image and combining it with the spectral response and noise of the RGB sensor to generate the encoded image corresponding to the current iteration round includes: Based on the PSF set corresponding to the current iteration round and the original hyperspectral image obtained in advance based on the true value simulation, the hyperspectral image corresponding to each band of each sector is convolved with the PSF to obtain the convolved image. The convolutional image is integrated onto the three channels of the RGB sensor, and the spectral response and noise of the RGB sensor are combined to obtain the encoded image corresponding to the current iteration.

6. The diffraction partitioning coding element modeling method according to claim 1, characterized in that, The step of inputting the encoded image and PSF set into the hyperspectral image reconstruction model, so that the hyperspectral image reconstruction model outputs the reconstructed hyperspectral image corresponding to the current iteration, includes: The encoded image and the PSF set corresponding to the current iteration are input into the feature extraction module of the hyperspectral image reconstruction model, so that the feature extraction module extracts the preprocessed feature map and inputs the preprocessed feature map into the reconstruction network of the hyperspectral image reconstruction model, so that the reconstruction network reconstructs the preprocessed feature map to output the reconstructed hyperspectral image corresponding to the current iteration. The feature extraction module includes: The first branch is used to input the encoded image into the multi-feature spectral attention network MFSA-Net in the reconstruction network, the first deconvolution unit in the second branch, and the second deconvolution unit in the third branch, respectively. The second branch is used to map the PSF set to the spectral feature space based on the first RGB mapping unit, and to deconvolve the PSF set mapped to the spectral feature space and the encoded image based on the first deconvolution unit, so as to output the spatial features corresponding to the encoded image. The third branch is used to extract the restoration space features corresponding to the encoded image based on multiple convolution units, and to map the restoration space features to the spectral feature space based on the second RGB mapping unit. Then, based on the second deconvolution unit, the restoration space features mapped to the spectral feature space and the encoded image are deconvolved to output the spectral reconstruction features corresponding to the encoded image. A dual-branch feature fusion unit is used to fuse the spatial features and the spectral reconstruction features to obtain a preprocessed feature map, and input the preprocessed feature map into the Res-Unet in the reconstruction network; The reconstructed network includes: Multi-head Spectral Attention Network (MFSA-Net), Res-Unet, and an addition operation module; The multi-head spectral attention network MFSA-Net includes: The first convolutional unit is used to perform preliminary feature extraction and channel dimension transformation on the encoded image to obtain a first feature map, and the first feature map is input into the multi-level embedding MLE unit and the feature fusion unit respectively. The multi-level embedding (MLE) unit is used to achieve multi-scale feature embedding of the first feature map through multi-level convolution and upsampling operations to obtain the second feature map. The first spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the second feature map through the spectral attention mechanism to obtain the third feature map, and input the third feature map into the second multi-channel residual fusion (MCRF) unit and the first downsampling unit respectively. The first downsampling unit is used to reduce the spatial resolution of the third feature map and increase the channel dimension to obtain the fourth feature map; The second spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the fourth feature image through the spectral attention mechanism to obtain the fifth feature image, and input the fifth feature image into the first MCRF unit and the second downsampling unit respectively. The second downsampling unit is used to reduce the spatial resolution of the fifth feature map and increase the channel dimension to obtain the sixth feature map; The third spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the sixth feature map through the spectral attention mechanism to obtain the seventh feature map. The first upsampling unit is used to restore the spatial resolution of the seventh feature map and reduce the channel dimension to obtain the eighth feature map; The first multi-channel residual fusion MCRF unit is used to fuse the fifth feature map and the eighth feature map to obtain the ninth feature map; The second convolutional unit is used to extract features from the ninth feature map to obtain the tenth feature map; The fourth spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the tenth feature map through the spectral attention mechanism to obtain the eleventh feature map; The second upsampling unit is used to restore the eleventh feature map to the original input resolution to obtain the twelfth feature map; The second multi-channel residual fusion MCRF unit is used to fuse the third feature map and the twelfth feature map to obtain the thirteenth feature map; The third convolutional unit is used to extract features from the thirteenth feature map to obtain the fourteenth feature map; The fifth spectral attention (SA) unit is used to extract the spectral and spatial features corresponding to the fourteenth feature map through a spectral attention mechanism to obtain the fifteenth feature map; A feature fusion unit is used to perform residual connection on the first feature map and the fifteenth feature map to obtain the sixteenth feature map; The fourth convolutional unit is used to map the sixteenth feature map to the target spectral dimension to obtain a reconstructed hyperspectral map, and input the reconstructed hyperspectral map into the summation operation module; The Res-Unet is used to reconstruct the hyperspectral data cube corresponding to the preprocessed feature map, and the hyperspectral data cube is input into the summation operation module; the summation operation module is used to perform addition and convolution on the reconstructed hyperspectral map and the hyperspectral data cube to output the reconstructed hyperspectral image corresponding to the current iteration round; The first, second, third, fourth, and fifth spectral attention SA units are each equipped with a multi-head spectral attention (MHSA) unit and a multi-scale feedforward network (MS FFN) unit.

7. The diffraction partition coding element modeling method according to claim 5, characterized in that, The step of determining the total loss for the current iteration based on the PSF set, the reconstructed hyperspectral image, and the original hyperspectral image, and updating the full-precision height map, learnable depth, and hyperspectral image reconstruction model based on the total loss, includes: Based on the PSF set corresponding to the current iteration round, the average absolute difference value of the point spread function PSF corresponding to the wavelength of each sector is obtained by weighted summation, and the average absolute difference value is mapped through the activation function to obtain the spectral difference regularization loss corresponding to the current iteration round. Based on the PSF set corresponding to the current iteration round, obtain the corresponding Fisher information matrix and determine the sum of all elements in the Fisher information matrix, and determine the Fisher information regularization loss corresponding to the current iteration round based on the sum of all elements in the Fisher information matrix. In addition, the mean absolute error between the reconstructed hyperspectral image and the original hyperspectral image is determined, and the mean absolute error is used as the reconstruction loss of the reconstruction network in the hyperspectral image reconstruction model; The total loss for the current iteration round is determined based on the spectral difference regularization loss, the Fisher information regularization loss, and the reconstruction loss corresponding to the current iteration round. The total loss corresponding to the current iteration is backpropagated to update the weights of the reconstruction network in the full-precision height map, learnable depth, and hyperspectral image reconstruction model based on the total loss, thereby obtaining the latest full-precision height map, learnable depth, and hyperspectral image reconstruction model in sequence.

8. A diffraction partitioning encoding element, characterized in that, include: An optical substrate; as well as Micro-nanostructures formed on the optical substrate, wherein the surface of the micro-nanostructures has a height distribution based on polar coordinate partitioning; The polar coordinate partitioning divides the optical aperture into multiple equiangular sectors, and the height distribution of each sector is formed by rotating around the polar coordinate center. The height distribution is defined by the modeling result data of the DPE element. The DPE element modeling result data is obtained by the diffraction partition coding element modeling method as described in any one of claims 1 to 7, such that the total depth of the height distribution is in the range of 0.768 μm to 3.072 μm.

9. An encoded image acquisition device, characterized in that, include: RGB sensor and the diffraction partition coding element as described in claim 8; The diffraction partition coding element is installed at a preset distance in front of the RGB sensor, and the diffraction partition coding element is configured to perform phase modulation on the incident light field from the target scene. The preset distance includes: 50mm; The RGB sensor is configured to receive a light signal that has been phase-modulated by the diffraction partition coding element and convert the light signal into a coded image of the target scene.

10. A snapshot hyperspectral imaging method, characterized in that, include: The encoded image of the target scene acquired by the encoded image acquisition device as described in claim 9 is obtained, and optical simulation is performed on the modeling result data of the DPE element to obtain the PSF set corresponding to the diffraction partition encoding element. The encoded image of the target scene and the PSF set are input into the target hyperspectral image reconstruction model so that the target hyperspectral image reconstruction model outputs the reconstructed hyperspectral image of the target scene.