Hyperspectral image processing device and image processing method
Patent Information
- Application Number
- JP2024527322
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-10
- Filing Date
- 2022-11-09
- Publication Date
- 2025-10-31
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 277,741, filed November 10, 2021, the entire contents of which are incorporated herein by reference.
[0002] Disclosed exemplary embodiments of the present invention relate to hyperspectral imaging, and in particular to an apparatus and method for hyperspectral imaging using a metasurface encoder. [Background technology]
[0003] Hyperspectral imaging is of great interest in many fields, including civil, environmental, aviation, military, and biological, to evaluate spectral features that enable identification and remote sensing of complex materials. Terrestrial hyperspectral imaging enables automated classification for food inspection, surgery, biology, dentistry, and medical diagnosis. Similarly, airborne and undersea hyperspectral imaging is now breaking new ground in agriculture and marine biology through drone aerial imagery for the taxonomic classification of fauna, and for precision agriculture and resource and mineral exploration and inspection. However, the current state of the art in hyperspectral imaging still faces problems of expensive setup costs, time-consuming data post-processing, slow data acquisition, and the need for macroscopic optical and mechanical components. A single hyperspectral image from a high-resolution camera typically requires gigabytes of storage space, making it extremely difficult to perform real-time video analysis using today's computer vision techniques.
[0004] Computational hyperspectral image reconstruction from a single RGB image is one technique to overcome some of the challenges mentioned above. Hyperspectral cameras based on integrated diffractive optical elements have been proposed, while others have utilized deep neural networks to design the spectral reconstruction filters. While these approaches help address the speed issue, they still fail to address the issues of complexity, cost, and slow data processing. Other obstacles are the use of rudimentary filter responses that are not optimized beyond primitive thin-film interference patterns, and the lack of integration structures that can take advantage of the modern footprint of CCD / CMOS sensors. Summary of the Invention [Means for solving the problem]
[0005] The following Summary of the Invention is intended to introduce the reader to various aspects of the Detailed Description, but is not intended to define or limit any of the invention.
[0006] In at least one broad aspect, a hyperspectral imaging device is provided, the imaging device comprising: an encoder layer including an i×j array of coding sub-arrays, each coding sub-array including m×n arrays of spectral encoders having a respective one of a plurality of transmission characteristics, each of the plurality of transmission characteristics being selected to encode a hyperspectral frequency range in k-dimensional space, where k is m×n; an image processing layer comprising an i×j array of detection subarrays aligned with the i×j array of encoding subarrays of the encoder layer, each detection subarray including an m×n array of photodetectors, each photodetector arranged to detect a respective transmission response of a respective spectral encoder in response to broadband light, outputting an i×j array of pixel responses, each pixel response including a pixel vector of m×n transmission responses; and a processor configured to decode the i×j array of pixel responses into a corresponding i×j array of pixel spectra to generate an output image that encompasses a hyperspectral frequency range.
[0007] In some cases, each spectral encoder is a flat optical device.
[0008] In some cases, the planar optical devices include respective patterned nanostructures selected to produce respective transmission characteristics.
[0009] In some cases, each of the plurality of transfer characteristics is linear. In some cases, each of the plurality of transfer characteristics is non-linear.
[0010] In some cases, multiple respective transmission characteristics of each spectral encoder in each coding subarray are selected by iteratively minimizing a loss function while optimizing the transmission characteristics for the application.
[0011] In some cases, the transmission characteristics of each of the multiple respective spectral encoders in each encoding subarray are selected by determining the k principal components that encode the eigenvectors with minimal loss for the application.
[0012] In some cases, the k principal components are determined by performing a singular value recomposition.
[0013] In some cases, the processor uses a linear projector to decode each pixel response of an i×j array of pixel responses.
[0014] In another broad aspect, a method of hyperspectral imaging is provided, the method comprising: providing an encoder layer comprising an i×j array of coding sub-arrays, each coding sub-array comprising an m×n array of spectral encoders having a respective one of a plurality of transfer characteristics, each of the plurality of transfer characteristics being selected to encode a hyperspectral frequency range in k-dimensional space, where k is m×n; providing an image processing layer comprising an i×j array of detector subarrays aligned with an i×j array of encoding subarrays of the encoder layer, each detector subarray including an m×n array of photodetectors; exposing the encoder layer to light to obtain the light; detecting at each photodetector a respective transmission response of each spectral encoder in response to the broadband light; outputting an i×j array of pixel responses from the image processing layer, each pixel response comprising a pixel vector of m×n transmission responses; The method includes decoding the i×j array of pixel responses into a corresponding i×j array of pixel spectra to generate an output image that encompasses the hyperspectral frequency range.
[0015] In some cases, each spectral encoder is a flat optical device.
[0016] In some cases, the planar optical device includes respective patterned nanostructures selected to produce respective transmission characteristics.
[0017] In some cases, each of the plurality of transfer characteristics is linear. In some cases, each of the plurality of transfer characteristics is non-linear.
[0018] In some cases, each of a plurality of transmission characteristics for each spectral encoder in each coding subarray is selected by iteratively minimizing a loss function while optimizing the transmission characteristics for the application.
[0019] In some cases, each of the multiple transmission characteristics of each spectral encoder in each encoding subarray is selected by determining the k principal components that encode the eigenvectors with minimal loss for the application.
[0020] In some cases, the k principal components are determined by performing a singular value decomposition.
[0021] In some cases, the processor uses a linear projector to decode each pixel response of an i×j array of pixel responses. [Brief description of the drawings]
[0022] The drawings included herein are for the purpose of illustrating various examples of the articles, methods, apparatus, and systems herein and are not intended to limit the scope of the teachings in any way. [Figure 1A] FIG. 1 is a plan view of an imaging device according to at least one embodiment. [Figure 1B] FIG. 1B is a front view of the image processing device in FIG. 1A. [Figure 1C] FIG. 1B is an exploded perspective view of the image processing device in FIG. 1A. [Figure 1D] 1 is a scanning electron micrograph of an encoded subarray according to at least one embodiment. [Diagram 2] FIG. 1 is a flowchart diagram of an exemplary method for hyperspectral imaging in accordance with at least one embodiment. [Figure 3A] FIG. 1 is a flowchart diagram of a hyperspectral image processing method according to at least one embodiment. [Figure 3B] 1 is a chart illustrating an example power density spectrum measured and reconstructed in accordance with at least one embodiment. [Figure 4A] FIG. 1 is a schematic diagram illustrating a coupled-mode photonic network as a feedback loop with a skip connection. [Figure 4B] FIG. 2 is a schematic diagram illustrating a trainable coupled resonant layer. [Figure 4C] 1 is a micrograph illustrating a parametric geometry generated using learned differentiable projections, in accordance with at least one embodiment. [Figure 5A] 1 is a pie chart showing the distribution of object classes in the FVgNET dataset. [Figure 5B] 1 is a bar graph showing the distribution of object classes in the FVgNET dataset. [Figure 5C] 1 is a raster image illustrating a semantic segmentation (region classification, a method of labeling each pixel of an image) mask determined in accordance with at least one embodiment. [Figure 6A] 1 is a scanning electron microscope (SEM) image of an example of a fabricated coded subarray. [Figure 6B] FIG. 6B is a simulated raster image of a scene from the FVgNET dataset as perceived through the encoder of the coding subarray of FIG. [Figure 6C] FIG. 6C is a table of images showing a qualitative comparison between the hyperspectral reconstructions of the scene in FIG. 6B. [Figure 6D] 1 is a line graph table showing a quantitative comparison. [Figure 7] 1 is a table showing a comparison between spectral semantic segmentation and RGB-based semantic segmentation. [Figure 8A] FIG. 1 is a schematic diagram illustrating an example network model of a differentiable hybrid inverse problem design predictor in accordance with at least one embodiment. [Figure 8B] FIG. 8B is a schematic diagram showing fully connected blocks of the example network model of FIG. 8A. [Figure 8C] A set of three qualitative comparisons between the trained and ground truth (the actual data used to train and test the output of an AI model) spectral responses for example metasurfaces. [Figure 9]1 is a table of images used to augment the dataset. [Figure 10A] 1 is a pair of reflectance plots for real grapes and synthetic grapes. [Figure 10B] FIG. 10A is an image of an actual grape, the reflectance of which is plotted. [Figure 10C] FIG. 10A is an image of artificial grapes whose reflectance is plotted. [Figure 11] 1 is a table of images showing a comparison between RGB and spectral information models for semantic segmentation of fruits. [Figure 12] 1 is a series of images showing reconstructions of image spectra at different wavelengths. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0023] Hyperspectral imaging has attracted considerable attention for identifying spectral signatures for image classification and automatic pattern recognition in computer vision. Existing snapshot hyperspectral imaging implementations rely on bulky, non-integrated and expensive optical elements, including lenses, spectrometers and filters. These macroscopic components, together with the large data sizes associated with these systems (some in the gigabyte range), do not typically enable high-speed data processing, such as high-resolution video, in real time. Such macroscopic components, along with the large data sizes (some in the gigabyte range) associated with these systems, typically do not allow for high speed data processing, such as high resolution video in real time.
[0024] The described embodiments generally provide an integrated architecture for hyperspectral imaging devices that is CMOS compatible and replaces bulk optics with nanoscale planar optical metasurfaces that can use their spatial geometry to encode wavelengths of light to generate a desired transmission response. Examples of metasurfaces are described, for example, in U.S. Patent Application No. 62 / 799,324, entitled "POLARIZER BEAM SPLITTER FOR FLAT OPTICAL DEVICES," and U.S. Patent Application No. 2022 / 0091318A1, entitled "OPTICAL PROCESSING DEVICE BASED ON MULTILAYER NANO ELEMENTS." In some cases, metasurfaces can be inverse problem designed using machine learning techniques to retain substantially complete and reconfigurable information in the transmission response for a given application. Unlike traditional RGB narrowband color filters, metasurfaces can have various transmission characteristics that are not limited to a single band, and therefore can successfully reconstruct broadband information. Furthermore, metasurfaces can be integrated with various basic optical components for different applications.
[0025] The described embodiments do not require specialized spectrometers, but can instead utilize conventional monochrome image sensors or cameras, thus opening up the possibility of real-time, high-resolution hyperspectral imaging with reduced complexity and cost, and the performance of the image processor is fast enough to support real-time image and / or video acquisition. The described embodiments generally employ model-driven optimization to interface a physical metasurface layer with state-of-the-art visual computing approaches based on end-to-end training. The described embodiments leverage this technology to compress high-dimensional spectral data into a low-dimensional space via well-defined projectors (see, e.g., FIG. 1C and FIG. 1D, further described herein) designed with end-to-end learning based on large-scale hyperspectral datasets. Inverse problem design software utilizing artificial intelligence (AI) that can be used to design metasurface projectors is described, for example, in U.S. Patent Application No. 62 / 799,324; Getman et al., "Broadband Vector-Based Ultrathin Optical Elements with Experimental Efficiency up to 99% in the Visible Region via Universal Approximators," Light: Science & Applications, 10(1):1-14, Mar. 2021, and Makarenko et al., "Robust and Scalable Flat Optical Elements on Flexible Substrates via Evolving Neuronal Networks," Advanced Intelligent Systems, Aug. 2021, pp. 2100105. These nanostructures are patterned to encode the broadband information carried by the incident spectrum into a barcode consisting of a discrete pattern of intensity signals (see, e.g., FIG. 1D, described further herein).The framework, which takes into account the physical model, determines the optimal projector response through various learning schemes designed based on the desired application.
[0026] Because conventional RGB cameras project the entire visible spectrum onto filters with only the three primary colors, conventional hyperspectral reconstruction typically involves backprojection from a low-dimensional RGB image into a densely sampled hyperspectral image (HSI).Metamerism is the effect whereby different spectral power distributions result in similar activation levels of the visual sensor. This effect removes important hyperspectral information, making it difficult to distinguish between different objects, but hyperspectral reconstruction is an approach used to partially recover such lost information. Such spectral projections are similar to autoencoders in the sense that they downsample the input to a lower dimensional space. In some cases, with the right algorithm to efficiently search this space, it may be possible to retrieve enough data to reconstruct the initial input.
[0027] The sparse coding method (a method of representing a signal as a linear sum of a small number of bases) statically discovers a set of basis vectors from a priori known hyperspectral image (HSI) dataset. The K-SVD algorithm is used to create overcomplete HSI and RGB dictionaries. The hyperspectral image (HSI) is reconstructed by decomposing the input image into a linear combination of basis vectors, which are then transferred to the hyperspectral dictionary. One limitation of sparse coding methods is the matrix decomposition algorithms they apply, which are sensitive to outliers and lead to poor performance. However, the capabilities of these methods have been extended by deep learning, especially supervised learning, to train architectures such as UNet to predict hyperspectral images (HSIs) from a single RGB image. For example, a radial basis function network was trained to convert white-balanced RGB values to reflectance spectra. Similarly, a two-stage reconstruction approach consisting of an interpolation-based upsampling method for RGB images has been proposed. The proposed end-to-end learning recovers the true hyperspectral image (HSI) from the upsampled image. Another approach uses different RGB cameras to obtain non-overlapping spectral information and reconstruct the hyperspectral image (HSI). These approaches reconstruct the spectral information from highly nonlinear predictive models, which are limited by their supervised learning structure. The models constrain the data downsampling to a non-optimal RGB image by applying a color generation function to the HSI or a generic RGB camera. In contrast, the described embodiments avoid all the problems of sparse coding and deep learning reconstruction methods by performing spectral downsampling with an optimally designed metasurface encoder or projector.
[0028] Conventional camera optical projectors mimic human color vision based on three primary colors. However, the bandwidth range of human vision may not be sufficient or adequate for all real-world purposes. Thus, the described embodiments extend the RGB camera concept from three channels to sampling of arbitrarily low-dimensional reflectance spectra and employ various variations of optimization routines to converge from a number of initial candidates to a set of optimal projectors. The selected projector thereby provides a multi-channel reconstruction of the hyperspectral image (HSI). It was also demonstrated that the 1 × 1 convolution operation achieves a similar functionality to an optical projector while processing multispectral data frames. The network is like an autoencoder, where the input hyperspectral image (HSI) is downsampled and then reconstructed by a decoder network.
[0029] For the inverse design of metasurface projectors, optimizing the best-fitting filters is a dimensionality reduction problem, which involves finding the principal component directions that encode the eigenvectors that exhibit the minimum loss. The results are generated from computations or experimental measurements of thin-film filters and represent a rough approximation of the exact principal components. In hyperspectral imaging, these components generally exhibit an irregular frequency-dependent pattern, which consists of a complex distribution of sharp and broad resonances. Conventional metasurface design approaches typically rely on libraries of precomputed metasurface responses and polynomial fitting to further generalize the relationship between design parameters and device performance. However, in at least some of the described embodiments, metasurface optical projectors can be designed using a hybrid inverse problem design approach that combines classical optimization and deep learning. In some additional embodiments, this hybrid inverse problem design approach can be further extended by adding differentiability, physics model regularization, and complex decoder projectors to tackle different computer vision tasks and perform thousands of parameter optimizations through a supervised end-to-end learning process.
[0030] 1A-1D, a hyperspectral imaging device is shown in accordance with at least one embodiment. The device 100 includes an image processing subsystem 101 coupled to a processor 190. The image processing subsystem 101 includes an encoder layer 110 aligned with an image processing layer 120. FIG. 1A is a top view of the imaging device 100. FIG. 1B is a front view of the imaging device with the encoder layer 110 stacked on top of the image processing layer 120. FIG. 1C is an exploded perspective view of the imaging subsystem. FIG. 1D is a scanning electron microscope photograph of an encoding sub-array 112 of an exemplary encoder layer 110.
[0031] The encoder layer 110 comprises an array of i×j encoding subarrays 112, each of which is composed of an m×n array of spectral encoders 114, or projectors, having respective transmission characteristics. The spectral encoder is a plain optical system that, in at least some embodiments, is formed from patterned nanostructures designed to generate respective transmission characteristics. In particular, a plurality of respective transmission characteristics are selected to encode hyperspectral frequency ranges in k-dimensional space, where k is m×n. In at least some embodiments, the transmission characteristics are linear for use with a linear operator. However, in some alternative embodiments, one or more of the transmission characteristics may be nonlinear for use with an appropriate nonlinear operator.
[0032] In at least some embodiments, the transmission characteristics for each encoder in a subarray are selected by iteratively minimizing a loss function while optimizing the transmission characteristics for the application.
[0033] In some other embodiments, the transmission characteristics for each encoder in the subarray are selected by determining the k principal components that encode the eigenvectors with minimal loss for the application, which may be determined by performing a singular value decomposition.
[0034] In general, each encoding subarray 112 of the encoder layer 110 is aligned with a respective detector subarray 122 of the image processing layer 120. Each encoder 114 is then aligned with a respective photodetector 124 such that there is a one-to-one correspondence between each encoder 114 of the encoder layer 110 and a respective photodetector 124 of the image processing layer 120. Together, each encoder of the encoding subarray generates a "barcode" that is detected by a corresponding photodetector of the detector subarray to generate an output "pixel." The exact size of the encoding and detector subarrays may vary depending on the application. In some embodiments, the encoding and detector subarrays (and "barcodes") have a 3x3 size. In other embodiments, the size may be different, e.g., 2x2, 4x3, 3x4, etc. Although this description provides a rectangular example of the subarrays, the subarrays are not limited to a rectangular geometry.
[0035] As mentioned above, the encoder layer 110 acts as an optical linear spectral encoder that compresses the input high-dimensional HSI β into a low-dimensional multi-spectral image tensor of transmission response through the respective transmission characteristics of the spectral encoder. The encoder layer 110 functions as an optical linear spectral encoder and converts the input high-dimensional HSIβ into a low-dimensional multispectral image tensor of transmission response via the respective transmission characteristics of the spectral encoder.
number
[0036] In at least one embodiment, the encoder layer is fabricated using a square piece of fused silica glass with a width of 15 mm and a thickness of 0.5 mm as a substrate. A thin layer of amorphous silicon is deposited on the glass by plasma-enhanced deposition, the thickness of which is controlled on each sample to fit the design requirements. In addition, 200 nm of a first resist (e.g., ZEP-520A from ZEON) and 40 nm of a second resist (e.g., AR-PC 5090 from ALLRESIST) are spin-coated and patterned in the shape of the nanostructures using an electron beam lithography system with a 100 kV acceleration voltage. The second resist is then removed by immersing each sample in deionized water for 60 seconds. The devices are created by immersing them in a solvent (e.g., ZED-50 from ZEON) for 90 seconds and rinsing in isopropyl alcohol for 60 seconds. In addition, 22 nm of chromium is deposited using electron beam evaporation to create a hard mask and perform lift-off, followed by ultrasonic agitation for 1 minute. The unprotected silicon is then removed using reactive ion etching by immersing the device in an etchant (e.g., TechniEtch Cr01 from Microchemicals) for 30 seconds to remove the metal mask and rinsing with deionized water to yield the final device.
[0037] Other processes can also be used to fabricate the device. For example, it is possible to use only one resist, to vary the thickness of the resist (e.g., between 20 nm and 1000 nm). Electron beam lithography systems with different acceleration voltages (e.g., 50 kV) can be used. Solvents may be replaced with equivalents. Furthermore, the metal mask can be omitted if a reverse version of the pattern is exposed in the resist, or if a negative polarity resist is used and the etching is fully optimized.
[0038] In some cases, UV lithography can be used with sufficiently high resolution to be suitable for mass production, in other cases nanoimprint lithography may be used, or silicon structures may be grown inside holes in patterned resist.
[0039] 6A is a scanning electron microscope (SEM) image of an example encoding fabricated sub-array 600 detailing the nanoscale structure of each of the nine encoders. In the example shown, each encoder in the 3×3 sub-array is arranged to occupy a square area 2.4 μm wide. This is a typical size of photodetectors found in modern digital image sensors, allowing integration with the image processing layer 120. In the example, the optical response of each encoder is characterized using linearly polarized light from 400 nm to 1000 nm wavelength.
[0040] 1A-1D , the image processing layer 120 has an i×j array of detection subarrays 122 aligned with the i×j array of encoding subarrays 112 of the encoder layer 110, with each detection subarray 122 including an m×n array of photodetectors 124 arranged to detect a respective transmission response of a respective spectral encoder 114 in response to broadband light (e.g., encompassing visible light, near infrared, maximum mid-infrared, or ultraviolet, or any combination thereof). The image processing layer 120 outputs the i×j array of pixel responses to the processor 190, with each pixel response including a pixel vector of m×n transmission responses.
[0041] The processor 190 performs hyperspectral image reconstruction to derive the transmit response tensor based on the application specific decoder mapping.
number
[0042] In at least some embodiments described herein, the encoder layer is optical and generally acquires and encodes data at the speed of light, and thus the data acquisition speed is primarily limited by the frame rate of the sensor (e.g., 30 frames per second (FPS)) and the processing speed. For real-time classification / segmentation tasks, the remaining layers of the network introduce a delay between the real-time processing of the hyperspectral image and the output of the task. One approach to achieve real-time processing is to use shallow networks implemented on a graphics processing unit (GPU). In one embodiment, the specifications of the dataset used for training were matched, and the system was designed to operate from 400 nm to 700 nm with a spectral resolution of 10 nm and a spatial resolution of 512 x 512. In general, a spectral resolution of better than 2 nm can be achieved, covering the wavelength range from 400 nm to 700 nm. With currently commercially available high resolution imaging sensors (e.g., with a resolution of 12 megapixels or more), a hyperspectral imaging device with a resolution of 2 megapixels or more and an acquisition speed approaching 1 Tb / s can be achieved.
[0043] 2, a flow chart diagram of an exemplary method of hyperspectral image processing is shown. Method 200 may be performed by an encoder layer, an image processing layer, and a processor, such as those of device 100. As described, the encoder layer has an i×j array of encoding subarrays, each encoding subarray composed of an m×n array of spectral encoders having a plurality of respective transmission characteristics, the plurality of respective transmission characteristics being selected to encode a hyperspectral frequency range in k-dimensional space, where k is m×n. The image processing layer has an i×j array of detection subarrays aligned with the i×j array of encoding subarrays of the encoder layer, each detection subarray including an m×n array of photodetectors.
[0044] Method 200 begins in step 210 by exposing an encoder layer, such as encoder layer 110 of device 100, to broadband light from a hyperspectral scene, and the encoder layer encodes the light according to the transmission characteristics of the encoder in each encoding sub-layer, as described herein, to generate a plurality of transmission responses.
[0045] In step 220, each photodetector of an image processing layer, such as image processing layer 120 of device 100, detects a respective transmission response of a respective spectral encoder in response to the broadband light. The image processing layer then outputs an i×j array of pixel responses from the image processing layer, each pixel response including a pixel vector of m×n transmission responses.
[0046] In step 225, a processor, such as processor 190 of device 100, decodes the i×j array of pixel responses into an i×j array of corresponding pixel spectra to generate an output hyperspectral image encompassing the hyperspectral frequency range.
[0047] Optionally, in step 240, the processor may perform semantic segmentation based on the output hyperspectral image, as described further herein.
[0048] Hyperspectral image reconstruction serves to reconstruct an input hyperspectral image (HSI) or its tensor with minimal loss. The loss is the root mean square error (RMSE) of the HSI.
number
number
[0049] At the encoder layer, the transfer function of the sub-micron nanostructure array can approximate any arbitrarily defined continuous function. The described embodiments use this universal approximation capability to design and implement linear-spectral encoder hardware optimized for application-specific image processing tasks related to hyperspectral information.
[0050] 3A, a data workflow for a hyperspectral image processing method according to at least one embodiment is shown. The workflow includes a general linear encoder operator
number
[0051] A linear dimensionality reduction operator Λ is obtained that finds a new equivalent encoded representation of the tensor β. The hyperspectral tensor of the image dataset is flattened into a matrix B where each column contains the power density spectrum of a set of pixels. We then apply a linear encoding Λ+ to obtain an approximation of the matrix B via a set of linear projectors Λ(ω) that map the spectral coordinates βij to a set of scalar coefficients Sijk for each pixel.
number
[0052] The spectral information contained in the spectral coordinates βij(ω) is embedded in an equivalent “barcode” Sijk consisting of several components. To implement the Λ-encoder layer in hardware, two different approaches can be used.
[0053] In one approach, if the user's task does not impose additional constraints, e.g., spectral reconstruction, the encoder can simply model the physical metasurface response.
number
number
[0054] Alternatively, for tasks that may impose further constraints, such as hyperspectral semantic segmentation, a learnable backbone can be used that uses the differentiable hybrid inverse problem design approach described above to create a differentiable physics model that is trained in an end-to-end approach. The differentiable hybrid inverse problem design approach designs the metasurface shape through an iterative process that minimizes the loss function Lseg by simultaneously optimizing the projector response Λ and the vector L containing all parameters that define the metasurface.
number
[0055] As described above, a single image processing sub-array, or “pixel” response combines the transmission responses from multiple encoders or metasurface projectors (i.e., of the coding sub-arrays) within a two-dimensional sub-array of encoders (or “sub-pixels”), which are replicated in space to form the encoder layer. Each coding subarray encodes the reflectance spectrum resulting from a scene as a "barcode."
number
number
[0056] In the PCA hybrid inverse problem design approach, the linear encoder Λ is obtained by an unsupervised learning technique using principal component analysis (PCA). Principal component analysis (PCA) is
number
number
number
[0057] Equation (2) gives the closest linear approximation of B in the least-squares sense. The decoder D is a linear projector
number
number
[0058] The particular linear operator selected is tailored to the particular application.
[0059] In some embodiments, a linear encoder other than PCA may be used, such as, for example, JPEG compression.
[0060] In yet other embodiments, nonlinear encoders can be used if the metasurface is made from a material with nonlinear transmission properties.
[0061] In the differentiable hybrid inverse problem design approach, the decoder operator D is the input tensor
number
number
number
number
number
[0062] In the inverse problem design of the projector, the encoder ε = H, and H(ω) represents the output transfer function of the metasurface response, which is obtained from the solution of the following set of coupled-mode equations:
number
number
number
number
number
number
number
number
[0063] This approach is based on the time-domain coupled-mode theory (TDCMT), which uses a set of rigorous coupled-mode equations equivalent to Maxwell's equations. The principle of the coupled-mode approach is to divide the geometric space Ω of light propagation into the resonator space Ωr and the external space Ωe. The external space is assumed to contain no sources or charges. Under this formulation, the set of Maxwell's equations can be expressed as 1 / X inverse matrix X -1 This is converted into equation (3). Power conservation is governed by the matrix σ
number
[0064] Equation (3) expresses the dynamics of the system as three independent matrices, namely the coupling matrix
number
number
[0065] The input-output transfer function obtained by solving equation (3)
number
number
number
[0066] 4A-4C show an example of a coupled mode network as a physical model of a differentiable metasurface. FIG. 4A is a schematic diagram showing the resonance and propagation effects as light passes through a metasurface encoder. As shown in FIG. 4A, the physical model derived from the coupled mode theory shares a common architecture with skip-coupled neural networks. FIG. 4B is a schematic diagram showing a trainable coupled resonant layer. FIG. 4C is a scanning electron microscope (SEM) image of the trained resonant layer. The optimized geometry forms the function of the metasurface encoder described above.
[0067] In some embodiments, a supervised optimization process can be used to project the resonator quantities in equation (3) onto the metasurface input parameters L using a differentiable hybrid inverse problem design approach. A deep neural network is trained to learn the relationship between the metasurface input parameters L and the resonator variables in equation (3). Following the same approach described here, the network is trained on a supervised spectral prediction task using an array of silicon boxes with simulated transmission / reflection responses.
[0068] FIG. 8A shows a model of an example of a differentiable hybrid inverse problem design predictor 800, which has a first branch 805 with a continuous input 808 and several fully connected (FC) blocks 810a and 810b connected in series. Each FC block 810 (shown in FIG. 8B) is composed of a multi-layer perceptron 812 (MLP) of different sizes, a batch normalization layer 814, and a dropout 816. To process categorical input variables (e.g., period, thickness) separately, a second branch 845 is provided with a categorical input 840, a linear embedding layer 850, and an FC block 810c connected in series. The main purpose of the branch 845 is to balance the weights of categorical and continuous variables in the model. Both the continuous branch 805 and the category branch 845 are then concatenated in block 860 and fed to a read branch 890 which is comprised of multiple FC blocks 810 d , 810 e and a non-linear read section 880 .
[0069] In one example, a training data set can be used that includes over 600,000 simulation results of a pure silicon-on-glass structure under a total electromagnetic / scattered field (TFSF) simulation, where each simulation has periodic boundary conditions with one of three different periods (250 nm, 500 nm, or 750 nm) and one of ten discrete thicknesses from 50 nm to 300 nm in 25 nm steps. Each structure consists of a random combination (up to 5) of cubic resonators. The dataset is split into a test part and a training part, which account for 20% and 80% of the total, respectively, with 10% of the training set used as a validation set.
[0070] For the training part, the Adam optimization technique can be used, for example with a learning rate of 1×10 −5 and a step learning rate scheduler with hyperparameters of step size=50 and γ=0.1. To achieve the desired system response in either transmission or reflection, a sigmoid activation function is applied to the top layer of the FCN. This function maps the output spectrum to the range [0,1], which aids in convergence at the beginning of the learning stage. Due to the use of periodic boundary conditions, random translations and rotations may be used for data augmentation.
[0071] Using this approach, a validation mean squared error of 0.008 is achieved. Figure 8C provides a qualitative comparison of the learned spectral responses to the ground truth (the actual data used to train and test the output of the AI model) spectral responses.
[0072] The described embodiments may be trained and validated using a variety of datasets. In some embodiments, three publicly available datasets: Namely, the CAVE dataset (available at https: / / www1.cs.columbia.edu / CAVE / databases / multispectral / ), which consists of 32 indoor images covering 400nm to 700nm, and the Harvard and KAUST datasets (available at http: / / vision.seas.harvard.edu.hyperspec and https: / / repository.kaust.edu.sa / handle / 10754 / 670368, respectively), which contain both indoor and outdoor scenes, respectively, with spectral bands covering 420nm to 720nm and 400nm to 700nm, respectively, and contain 75 and 409 images, respectively. An additional hyperspectral dataset, FVgNET, is also available (available at https: / / github.com / makamoa / hyplex). FVgNET consists of 317 scenes showing both natural and man-made fruits and vegetables, photographed indoors under controlled lighting conditions, covering the range from 400nm to 1000nm. Approximately 40% of the scenes consist of one row of objects located in the focal plane of the camera. The remaining scenes show two rows of objects, with the focal plane located between them. The white reference panel is approximately constant throughout the dataset to facilitate normalization. The hyperspectral images have a spatial resolution of 512x512 pixels and 204 spectral bands. An RGB image is also provided as seen through the lens of the RGB camera for each scene with the same spatial resolution. In some cases, the dataset can be extended with, for example, 20 additional images (examples of which are shown in Figure 9) to verify generalization capabilities. The resulting reconstruction error for these exemplary images is 2.54±2.72, a value consistent with the results obtained with the KAUST dataset used to train the encoder.
[0073] The FVgNET images were acquired using a setup consisting of a blank sheet of paper placed in an infinite curve, a configuration used in photography to separate the object from the background, illuminating the object with overhead white LED room lighting, a 150W halogen lamp (Thorlabs™ OSL2) with a glass diffuser, and a 100W tungsten bulb mounted within a diffuse reflector to provide good spectral coverage while minimizing the presence of shadows in the final image.
[0074] Referring now to Figures 5A and 5B, a chart showing the distribution of object classes in the FVgNET dataset is shown. For each class of object (e.g. apples, oranges, peppers), an approximately equal number of scenes are generated showing a) only natural objects, and b) only man-made objects. The dataset consists of 12 classes, represented in the images in proportion to their color diversity. Furthermore, 80% of the images are annotated with an additional segmentation mask. Each class has an approximately equal number of instances in the dataset, except for apples and peppers, because they have more color diversity. Semantic segmentation masks are incorporated into the dataset by processing the RGB images generated from the 204 spectral channels. The images are acquired in such a way as to avoid object intersections, allowing the automatic generation of masks for the areas occupied by each object. Each marked object was then annotated to identify each object class and whether they are natural or man-made.
[0075] Figure 5C shows the implementation of semantic segmentation masks (labels) on the images of the dataset. On the left is an RGB visualization of the hyperspectral image. On the right are the segmentation masks and labels for each object.
[0076] Referring now to Figure 6B, there is shown a scene from the FVgNET dataset, based on example data, as perceived through each of the encoders of a coding layer composed of coding sub-arrays 600. Figure 6C shows a qualitative comparison of the hyperspectral reconstruction of this scene based on both simulated and example barcodes, with the original. 6D shows a quantitative comparison between the original spectrum 610 according to a conventional approach, the reconstructed spectrum 650 according to an exemplary embodiment, and the reconstructed spectrum 690. 80% of the dataset was designated for decoder training, and the remainder was designated for validation purposes.
[0077] 7A-7C, a comparison between spectral and RGB-based semantic segmentation is shown. In the illustrated example, segmentation is performed between artificial and real fruits from scenes in the FVgNET dataset. The artificial and real fruits have similar RGB colors. However, they differ significantly in their reflectance spectra, as shown in FIGS. 10A-10C.
[0078] The performance of the described embodiment can be illustrated by training two classification networks for comparison purposes: one model uses the described encoder for labeling of semantic segmentation, and the second model uses the RGB channels. Both models use the same U-Net-like decoder and the same parameters (number of epochs, batch size, learning rate). The results are summarized in Figures 7A-7C, where Figure 7A shows a comparison between the spectral information model, the RGB-only model, and the segmentation masks generated from the ground truth, Figure 7B shows the confusion matrix for the RGB-only model, and Figure 7C shows the confusion matrix for the spectral information model. Each value in the confusion matrix represents the number of pixels classified as the row item in the segmentation mask of the column item.
[0079] While the mask quality is similar for both methods, the average Intersection over Union (mIoU) score for the spectral information model is significantly higher compared to the RGB model. The mIoU calculated on the theoretical and experimental responses of the encoder reaches 81% and 74%, respectively. Conversely, for the RGB model, the mIoU decreases to 68%. The confusion matrix of the RGB trained model shows that the RGB model struggles to predict the correct outcome for pairs of real and artificial fruits with similar colors (e.g., Figure 7B). Conversely, the spectral information model generates correct labels for most real and artificial fruit pairs (e.g., Fig. 7C), outperforming the RGB model in mIoU and F1. These results demonstrate that the compact barcodes generated by the described embodiments efficiently compress spectral features that convey important information about the imaged objects.
[0080] Now referring to FIG. 11, a comparison between the RGB model and the spectral information model for semantic segmentation of fruits is shown. The first row from the top shows 16 RGB photos of fruits. The second row from the top shows 16 segmentation masks generated based on hyperspectral images corresponding to the 16 RGB photos in the first row. The third row from the top shows 16 segmentation masks generated based on RGB images corresponding to the 16 RGB photos in the first row. The fourth row from the top shows 16 segmentation masks generated based on the ground truth (actual data) of the 16 RGB photos in the first row.
[0081] Referring to Figure 12, a series of images are provided showing image spectral reconstructions at different wavelengths. The first and second rows from the top left show the original and spectrally reconstructed images of a bicycle at seven different wavelengths. A 3x3 grid on the right is used to simulate how the scene is perceived through each of the nine encoders.
[0082] Similarly, the third and fourth rows show the original and spectrally reconstructed images for a display of fruit at seven different wavelengths, using the 3x3 grid on the right to simulate the view of the scene as perceived through each of the nine different encoders.
[0083] Similarly, the fifth and sixth rows show the original and spectrally reconstructed images for samples written at seven different wavelengths, using the 3x3 grid on the right to simulate the view of the scene as perceived through each of the nine different encoders.
[0084] Similarly, the seventh and eighth rows show the original and spectrally reconstructed images of a fruit arrangement at seven different wavelengths, using the 3x3 grid on the right to simulate the view of the scene as perceived through each of the nine different encoders.
[0085] Various devices or processes have been described to provide example embodiments of the claimed subject matter. Such exemplary embodiments described do not limit any claim, and any claim may cover a process or device different from those described. The claims are not limited to devices or processes having all the features of any one of the devices or processes described above, or to features common to several or all of the devices or processes described above. The devices or processes described above may not be embodiments of the exclusive rights conferred by the issuance of this patent application. The subject matter described above, to which no exclusive rights are conferred by the issuance of this patent application, may be the subject matter of other protective documents, for example, continuing patent applications, and the applicants, inventors, or patentees do not intend to relinquish, abandon, or dedicate to the public the subject matter by its disclosure in this document.
[0086] For simplicity and clarity of description, reference numerals may be repeated among the figures to indicate corresponding or similar elements. Additionally, numerous specific details are described to provide a thorough understanding of the subject matter described herein. However, it will be understood by those skilled in the art that the subject matter described herein may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the subject matter described herein.
[0087] As used herein, the term "and / or" is intended to represent an inclusive "or." That is, "X and / or Y" is intended to mean, for example, X or Y, or both. As a further example, "X, Y, and / or Z" is intended to mean X, Y, Z, or any combination thereof.
[0088] Terms of degree, such as "substantially," "about," and "approximately," as used herein, refer to a reasonable amount of deviation from the modified term such that the results are not significantly altered. These terms of degree may also be construed to include deviations from the modified term if such deviations do not negate the meaning of the term they modify.
[0089] Any recitation herein of numerical ranges by endpoints includes all numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.90, 4, and 5). It is also to be understood that all such numbers and fractions are presumed to be modified by the term "about," which refers to a variation of the referenced number by a certain amount where the result would not change appreciably.
[0090] Some elements herein may be identified by a part number consisting of a base number followed by an alphabetic or subscript numeric suffix (e.g., 112a, or 112-1). All elements having a common base may be referred to collectively or generically using the base number without a suffix (e.g., 112).
[0091] The systems and methods described herein may be implemented as a combination of hardware or software. In some cases, the systems and methods described herein may be implemented, at least in part, by using one or more computer programs running on one or more programmable devices that include at least one processing element and data storage element (including volatile and non-volatile memory, and / or storage elements). These systems may also have at least one input device (e.g., a push button keyboard, a mouse, a touch screen, etc.) and at least one output device (e.g., a display screen, a printer, a wireless radio, etc.), depending on the nature of the device. Furthermore, in some examples, one or more of the systems and methods described herein may be implemented within or as part of a distributed or cloud-based computing system having multiple computing components distributed across a computing network. For example, the distributed or cloud-based computing system may correspond to a private distributed or cloud-based computing cluster associated with an organization. Additionally or alternatively, the distributed or cloud-based computing system is a publicly accessible distributed or cloud-based computing cluster, such as a computing cluster maintained by Microsoft Azure™, Amazon Web Services™, Google Cloud™, or another third-party provider. In some examples, the distributed computing elements of the distributed or cloud-based computing system may be configured to perform one or more parallelized, fault-tolerant, distributed computing and analysis processes, such as processes provided by the Apache Spark™ distributed cluster computing framework. Further, in addition to the CPU described herein, the distributed computing elements may also include one or more graphics processing units (GPUs) capable of processing thousands of operations (e.g., vector operations) in a single clock cycle, and may also, or alternatively, include one or more tensor processing units (TPUs) capable of processing hundreds of thousands of operations (e.g., matrix operations) in a single clock cycle.
[0092] Some elements used to implement at least some of the systems, methods, and devices described herein may be implemented via software written in a high-level procedural language, such as an object-oriented programming language. Thus, program code may be written in any suitable programming language, such as, for example, Python™ or Java™. Alternatively, or in addition, some of these elements implemented via software may be written in assembly language, machine language, or firmware, as appropriate. In any case, the language may be a compiled or interpreted language.
[0093] At least some of these software programs may be stored on a storage medium (e.g., a computer readable medium, such as, but not limited to, a read-only memory, a magnetic disk, an optical disk, etc.) or a device that is readable by a general-purpose or special-purpose programmable device. The software program code, when read by the programmable device, configures the programmable device to operate in a new, specific and predefined manner to perform at least one of the methods described herein.
[0094] Furthermore, at least some of the programs associated with the systems and methods described herein can be distributed in a computer program product that includes a computer-readable medium bearing computer usable instructions for one or more processors. The medium may be provided in a variety of forms, including, but not limited to, one or more diskettes, compact discs, tapes, chips, and non-transitory forms such as magnetic and electronic storage devices. Alternatively, the medium may be transitory in nature, such as, but not limited to, wired transmissions, satellite transmissions, Internet transmissions (e.g., downloads), media, digital and analog signals, and the like. The computer usable instructions may also be in a variety of formats, including compiled and non-compiled code.
[0095] Although the above description provides one or more examples of a process or apparatus, it will be understood that other processes or apparatus are within the scope of the following claims.
[0096] To the extent that any amendment, characterization, or other assertion, whether prior or not, previously made with respect to any technology (this or any related patent application or patent (including parent, sibling or progeny)) may be construed as a disclaimer of the subject matter supported by the present disclosure of this application, applicant hereby cancels and revokes such disclaimer, and applicant hereby submits that the prior art previously considered in the related patent application or patent (including parent, sibling or progeny) must be revisited.
Claims
1. an encoder layer including an i×j array of coding sub-arrays, each coding sub-array including an m×n array of spectral encoders having a respective one of a plurality of transfer characteristics, each of the plurality of transfer characteristics selected to encode a hyperspectral frequency range in k-dimensional space, where k is m×n; an image processing layer comprising an i×j array of detection sub-arrays aligned with an i×j array of encoding sub-arrays of said encoder layer, each detection sub-array comprising an m×n array of photodetectors, each photodetector arranged to detect a respective transmission response of a respective spectral encoder in response to broadband light, and outputting an i×j array of pixel responses, each pixel response comprising a pixel vector of m×n transmission responses; a processor configured to decode the i×j array of pixel responses into an i×j array of corresponding pixel spectra to generate an output image that encompasses a hyperspectral frequency range.
2. The apparatus of claim 1 , wherein each spectral encoder is a planar optical device.
3. The apparatus of claim 2 , wherein the planar optical device includes respective patterned nanostructures selected to produce the respective transmission characteristics.
4. 4. The apparatus of claim 1, wherein each of the plurality of transfer characteristics is linear.
5. 4. The apparatus of claim 1, wherein each of the plurality of transmission characteristics is non-linear.
6. 6. The apparatus of claim 5, wherein the plurality of respective transmission characteristics of the respective spectral encoders in each coding subarray are selected by iteratively minimizing a loss function while optimizing the transmission characteristics for an application.
7. 4. The apparatus of claim 1, wherein the respective transmission characteristics of the respective spectral encoders in each encoding sub-array are selected by determining k principal components that encode eigenvectors with minimal loss for an application.
8. The apparatus of claim 7 , wherein the k principal components are determined by performing a singular value decomposition.
9. 4. The apparatus of claim 1, wherein the processor decodes each pixel response of an i x j array of image responses using a linear projector.
10. 1. A hyperspectral image processing method, comprising: providing an encoder layer comprising an i×j array of coding sub-arrays, each coding sub-array comprising an m×n array of spectral encoders having a respective one of a plurality of transfer characteristics, each of the plurality of transfer characteristics selected to encode a hyperspectral frequency range in k-dimensional space, where k is m×n; providing an image processing layer comprising an i×j array of detection subarrays aligned with an i×j array of encoding subarrays of the encoder layer, each detection subarray comprising an m×n array of photodetectors; exposing the encoder layer to capture light; detecting, at respective photodetectors, a respective transmission response of each spectral encoder in response to the broadband light; outputting an i×j array of pixel responses from the image processing layer, each pixel response comprising an m×n pixel vector of transmission responses; A method comprising decoding an i×j array of pixel responses into a corresponding i×j array of pixel spectra to produce an output image encompassing a hyperspectral frequency range.
11. The method of claim 10 , wherein each spectral encoder is a planar optical device.
12. The method of claim 11 , wherein the planar optical device includes respective patterned nanostructures selected to produce the respective transmission characteristics.
13. 13. The method of claim 10, wherein each of the plurality of transfer characteristics is linear.
14. 13. The method of claim 10, wherein each of the plurality of transmission characteristics is non-linear.
15. 15. The method of claim 14, wherein each of a plurality of transmission characteristics of the respective spectral encoder in each coding subarray is selected by iteratively minimizing a loss function while optimizing the transmission characteristics for an application.
16. 13. The method of claim 10, wherein the respective transmission characteristics of the respective spectral encoders in each encoding sub-array are selected by determining k principal components that encode eigenvectors with minimal loss for an application.
17. The method of claim 16 , wherein the k principal components are determined by performing a singular value decomposition.
18. 13. A method according to any one of claims 10 to 12, wherein the processor decodes each pixel response of the i x j array of pixel responses using a linear projector.