Optical tensor convolution computing system and method based on multi-imaging projection architecture

By implementing an optical tensor convolution computation system based on a multi-imaging projection architecture, parallel frequency domain multiplication of multi-channel image matrices and convolution kernel matrices is achieved. This solves the problems of computational power limitations and insufficient resource utilization in existing optical tensor convolution computation technologies, and provides a highly parallel, high-speed, and low-power optical tensor convolution computation solution suitable for fields such as deep learning and autonomous driving.

CN116258624BActive Publication Date: 2026-07-24SHANGHAI INST OF OPTICS & FINE MECHANICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI INST OF OPTICS & FINE MECHANICS CHINESE ACAD OF SCI
Filing Date
2023-01-17
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing optical computing architectures suffer from limitations in computing power, computational speed, and resource utilization when performing optical tensor convolution calculations. This is especially true in applications with small convolution kernels in deep learning and convolutional neural networks, where it is difficult to leverage the advantages of multi-imaging projection architectures.

Method used

An optical tensor convolution computation system based on a multi-imaging projection architecture is adopted. The system loads multi-channel image matrix information through a light source array module, generates sub-beams of different diffraction orders using an imaging projection module, performs frequency domain multiplication in a signal modulation module, and finally obtains the optical tensor convolution result in a detection module, thereby realizing parallel frequency domain multiplication of multi-channel image matrices and convolution kernel matrices.

Benefits of technology

It achieves highly parallel, high-speed, and low-power optical tensor convolution calculations, improving the system's computing power and accuracy. It can directly obtain the convolution result matrix on the detection module end face, supports large-scale parallelism and high-precision tensor convolution calculations, and is suitable for practical applications such as deep learning and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116258624B_ABST
    Figure CN116258624B_ABST
Patent Text Reader

Abstract

An optical tensor convolution calculation system and method based on a multi-imaging projection architecture, comprising: an optical tensor convolution component and an electronic control component, the optical tensor convolution component comprising: a light source array module, loading information of a multi-channel image matrix to an input optical signal to obtain an optical signal carrying multi-channel image information; an imaging projection module, generating different diffraction orders of the optical signal carrying multi-channel image information; a signal modulation module, obtaining frequency domain multiplication information; a detection module, obtaining an optical tensor convolution result; the electronic control component comprising: a data parallel loading module for preprocessing and loading a multi-channel image matrix and a multi-channel kernel matrix; a data parallel downloading module for subsequent processing of the optical signal detected by the detection module; and an automatic control module. The application is expected to realize an optical deep neural network and be applied in practical scenarios such as target recognition, autonomous driving and high-performance computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical computing technology, and more specifically, to an optical tensor convolution computing system and method based on a multi-imaging projection architecture. Background Technology

[0002] Various optical computing schemes have been proposed and developed. Among them, matrix-vector multipliers based on planar optical waveguide architecture (Priority 1: Nat. Photon. 11, 441 (2017), Priority 2: Nature 569, 208 (2019), Priority 3: Nature 589, 52 (2021), Priority 4: Nature 589, 44 (2021)) can only perform one-dimensional vector-matrix multiplication calculations. Their scalability is severely limited by chip manufacturing processes, thus this scheme has an insurmountable computing power ceiling. In addition, spatial diffractive optical neural networks based on diffractive optical elements (Priority 5: Science 361, 1004-1008 (2018)) use cascaded diffractive optical devices to realize spatial interconnection between different convolutional layers. In principle, this avoids the obstacle of achieving high-density interconnection in one-dimensional vectors and can make full use of optical spatial interconnection capabilities. However, this scheme requires iterative calculation of the amplitude and phase of the diffractive optical element according to the relevant formulas of diffractive optics to achieve the calculation of the specified matrix, which actually causes additional computational load. This limits the scope of application of this scheme to some extent. Finally, the convolution calculation system based on the 4f system (Prior Art 6: Optica 7, 1812-1819 (2020)) is difficult to achieve large-scale, high-precision convolution calculation due to the limitation of the Fourier transform relationship between the object plane and the image plane. The Shanghai Institute of Optics and Fine Mechanics of the Chinese Academy of Sciences proposed an optical convolution calculation system based on a multi-imaging projection architecture based on the shadow delivery method. By introducing the Damman grating, it is expected to achieve large-scale matrix-matrix optical convolution, which is a perfect optical realization of the mathematical convolution process. However, in the actual application of convolutional neural networks, small convolution kernels of 3×3, 5×5, and 7×7 are often used, which makes it difficult to give full play to the advantages of this scheme. Therefore, the Shanghai Institute of Optics and Fine Mechanics proposed a corresponding matrix rearrangement method to realize the convolution of multi-channel convolution kernels and multi-channel input images. However, this method requires space resources for the input image matrix and the convolution kernel matrix, which reduces the computational power of the system.

[0003] In summary, planar integrated optical waveguide solutions face significant computational limitations. Currently, planar integrated optical waveguide solutions primarily employ wavelength division multiplexing (WDM) with time delay to perform optical tensor convolution calculations. However, this method still requires unfolding the convolution kernel into a one-dimensional vector input into the weight matrix, transforming the mathematically complex translation and sliding process of convolution into a sequential time-series operation. This, to some extent, limits the speed of optical convolution calculations.

[0004] While spatial diffraction optics can theoretically achieve large-scale matrix computations, realizing three-dimensional subwavelength optical elements with precise control over complex electromagnetic fields, especially in the optical band, remains a significant challenge in practice. Optical computing schemes based on 4f optical systems process the input signal in the spectral plane and then utilize the convolution theorem to achieve optical convolution. However, this scheme is constrained by the relationship between the object plane and the spectral plane, limiting its computational power. Optical convolution schemes based on multi-imaging projection architectures use beam splitters to copy and shift the input matrix; however, their computational advantage cannot be fully realized in practical convolutional neural networks. It can be seen that 4f optical systems can reuse convolution kernel matrices in the spectral plane, while multi-imaging projection architectures can reuse convolution kernel matrices in the object plane. However, currently, there is still no optical computing architecture that can simultaneously reuse input image matrices and convolution kernel matrices in both the object and spectral planes. Therefore, no optical computing architecture based on spatial interconnection optics for optical tensor convolution computation has yet been proposed. Summary of the Invention

[0005] To address the aforementioned deficiencies and gaps in the existing technology, this invention provides an optical tensor convolution calculation system and method based on a multi-imaging projection architecture.

[0006] According to one aspect of the present invention, an optical tensor convolution computation system based on a multi-imaging projection architecture is provided, comprising: an optical tensor convolution component, wherein the optical tensor convolution component comprises, in sequence according to the optical path propagation direction: a light source array module, an imaging projection module, a signal modulation module, and a detection module; wherein:

[0007] The light source array module is used to load information from a multi-channel image matrix onto the input optical signal to obtain an optical signal carrying multi-channel image information.

[0008] The imaging projection module is used to generate sub-beams of different diffraction orders from the optical signal carrying multi-channel image information, wherein each diffraction order carries all the information of the multi-channel image. Then, an optical Fourier transform is performed on the sub-beams of different diffraction orders to obtain the spectral information of the multi-channel image matrix, and this information is output to the signal modulation module.

[0009] The signal modulation module is used to load the spectral information of the multi-channel convolution kernel matrix, and to perform a frequency domain multiplication operation on the spectral information of the multi-channel image matrix carried in the sub-beams of different diffraction orders and the spectral information of the multi-channel convolution kernel matrix to obtain the frequency domain multiplication information of the multi-channel image matrix and the multi-channel convolution kernel matrix, thereby obtaining an optical signal carrying the frequency domain multiplication information.

[0010] The detection module is used to perform an optical inverse Fourier transform on the optical signal carrying frequency domain multiplication information to obtain the optical tensor convolution result of the multi-channel image matrix and the multi-channel convolution kernel matrix.

[0011] Preferably, the imaging projection module includes multiple beam splitters that divide the optical signal carrying multi-channel image matrix information into several sub-beams. Then, an optical Fourier transform is performed on each sub-beam to obtain the spectral information of the multi-channel image matrix. Each sub-beam carries the spectral information of the multi-channel image matrix and is incident on different positions of the signal modulation module according to its respective diffraction angle, achieving parallel translation of the optical signal carrying the spectral information of the multi-channel image matrix. The spectral information of the multi-channel image matrix carried in the sub-beams incident on different positions of the signal modulation module is multiplied at the pixel level with the spectral information of the multi-channel convolution kernel matrix, achieving parallel frequency domain multiplication of the multi-channel image matrix and the multi-channel convolution kernel matrix.

[0012] Preferably, the parallel frequency domain multiplication operation between the multi-channel image matrix and the multi-channel convolution kernel matrix includes:

[0013] The multi-channel image matrix is ​​replicated using the multiple beam splitters in the imaging projection module;

[0014] The copied multi-channel image matrix is ​​subjected to optical Fourier transform to obtain its spectral information, and then projected onto different positions of the signal modulation module.

[0015] The spectral information of the convolution kernel matrix of different channels is loaded at different positions of the modulation module to realize the dot product operation between the elements in the spectral information of the multi-channel image matrix and the corresponding elements in the spectrum of the multi-channel convolution kernel matrix.

[0016] The light beams at different positions carry the dot product information of the spectrum of their respective convolution kernel matrix and the spectrum of the multi-channel image matrix;

[0017] Optical tensor convolution results of multi-channel image matrix and multi-channel convolution kernel matrix are obtained by performing optical inverse Fourier transform on the frequency domain multiplication information at all positions. The characteristic is that optical signals carrying the frequency domain multiplication information at different positions are incident on different positions of the detection surface of the detection module according to their respective diffraction tilt angles.

[0018] Based on the specific neural network structure, by adjusting the diffraction tilt angle of the multiple beam splitters and the feature distance between the imaging projection module and the light source array module, the photon convolution results of the convolution kernel matrix of different channels and the multi-channel image matrix are added together to obtain the final result after convolution of different neural network structures.

[0019] Preferably, the signal loading surface of the light source array module and the modulation plane of the signal modulation module satisfy the object-image conjugate relationship, and the modulation plane of the signal modulation module and the detection surface of the detection module also satisfy the object-image conjugate relationship.

[0020] Preferably, by adjusting the diffraction tilt angles of the multiple beam splitters and the characteristic distance between the imaging projection module and the light source array module, the diffraction angles of each diffraction order of the imaging projection module and the corresponding positions of different channel convolution kernels in the signal modulation module are aligned and matched, thereby realizing the dot product operation of the spectral information of the multi-channel image matrix and the multi-channel convolution kernel matrix.

[0021] Preferably, the system further includes any one or more of the following:

[0022] - The light source array module adopts a light-emitting element array, a fiber array, or a spatial light modulator;

[0023] - The imaging projection module includes a Fourier transform lens and one or more Damman gratings, wherein the Damman gratings are one-dimensional Damman gratings or two-dimensional Damman gratings;

[0024] - The imaging projection module uses a metasurface beam splitter or a metamaterial beam splitter;

[0025] - The signal modulation module employs a spatial light modulator or an optical mask;

[0026] - The detection module includes a Fourier transform lens and an optical receiver array.

[0027] Preferably, the system further includes an electronic control component: the electronic control component includes a data parallel loading module, a data parallel downloading module, and an automation control module;

[0028] The data parallel loading module is used to preprocess the multi-channel image matrix and the multi-channel convolution kernel matrix, and load the preprocessed information in parallel into the light source array module and the signal modulation module.

[0029] The data parallel download module is used to convert the optical signals detected by the detection module into electrical signals in parallel, and to perform subsequent processing on the electrical signals;

[0030] The automated control module is used to automatically adjust the optical tensor convolution component according to different algorithm structures, and change the projection translation step size of the imaging projection module to achieve the addition of optical tensor convolution results.

[0031] According to a second aspect of the present invention, an optical tensor convolution calculation method based on a multi-imaging projection architecture is provided, characterized in that it includes:

[0032] The multi-channel image matrix information is loaded onto the input optical signal to obtain an optical signal carrying the multi-channel image matrix information;

[0033] The optical signal carrying multi-channel image matrix information is used to generate sub-beams of different diffraction orders;

[0034] Optical Fourier transform is performed on all diffraction orders of the sub-beams to obtain their corresponding spectral information;

[0035] In the signal modulation module, the spectral information of the convolution kernel matrix corresponding to the position of each diffraction order sub-beam is loaded, and the spectral information of the multi-channel image matrix carried by the sub-beams of different diffraction orders is multiplied by the spectral information of the multi-channel convolution kernel matrix to obtain the optical signal carrying the frequency domain multiplication information of the two at all positions.

[0036] An optical inverse Fourier transform is performed on the optical signal carrying frequency domain multiplication information, and the optical tensor convolution result of the multi-channel image matrix and the convolution matrix is ​​obtained in the detection module.

[0037] Preferably, the sub-beams of different diffraction orders are transmitted to different positions of the signal modulation module at different diffraction angles, wherein the optical signal of each diffraction order carries all the information of the multi-channel image matrix;

[0038] Preferably, the light source array module is loaded with a multi-channel image matrix.

[0039] ,

[0040] in, The image matrix for the first channel. The image matrix of the m-th channel, It is an image matrix with n×(m−1)+1 channels. Let be the image matrix of the n×m-th channel, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively. O is the zero matrix between adjacent image matrices. The number of zero elements in the zero matrix must be sufficient to ensure that the convolution results of adjacent image matrices do not overlap.

[0041] Perform a Fourier transform on the multi-channel image matrix.

[0042] ,

[0043] Among them, v 11 v12 …v nm It is the spectral information of the multi-channel image matrix, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively.

[0044] The spectral information of the multi-channel image matrix is ​​replicated by the multiple beam splitters in the imaging projection module. Each sub-beam of each diffraction order carries the complete spectral information and propagates along its respective diffraction direction to different positions on the modulation surface of the signal modulation module.

[0045] ,

[0046] Among them, V 11 V 12 …V kp It is the spectral information of the multi-channel image matrix at different positions of the modulation surface of the signal modulation module, and the subscripts k and p are the row number and column number of the sub-beam, respectively.

[0047] Multiply the spectral information of the multi-channel image matrix with the spectral information of the convolution kernel matrix.

[0048] ,

[0049] Among them, C 11 It is the spectrum of the first channel convolution kernel matrix, C 12 It is the spectrum of the second channel convolution kernel matrix, C kp Y is the spectrum of the convolution kernel matrix of the p×(k−1)+1th channel. 11 Y is the dot product of the spectrum of the first channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. 12 Y is the dot product of the spectrum of the second channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. kp It is the dot product of the spectrum of the convolution kernel matrix of the p×(k−1)+1th channel and the spectrum information of the multi-channel image matrix.

[0050] Perform an inverse Fourier transform on the dot product of the spectrum of the multi-channel convolution kernel matrix and the spectrum of the multi-channel image matrix.

[0051] ,

[0052] Among them, y 11 It is the convolution result of the first channel convolution kernel matrix and the multi-channel image matrix, y 12 It is the convolution result of the second channel convolution kernel matrix and the multi-channel image matrix, y kp It is the convolution result of the p×(k−1)+1th channel convolution kernel matrix and the multi-channel image matrix.

[0053] Preferably, the loaded multi-channel image information, the loaded multi-channel convolution kernel information, and the obtained convolution result matrix information are all analog quantities. The convolution result matrix information is quantized into a digital result after digital processing to realize analog optical tensor convolution operation.

[0054] Preferably, the high-bit multi-channel convolution kernel matrix information and the high-bit multi-channel image matrix information to be processed are represented as multiple encoded low-bit matrices, respectively, to obtain encoded low-bit multi-channel convolution kernel matrix information and low-bit multi-channel image matrix information; the low-bit multi-channel convolution kernel matrix information and the low-bit multi-channel image matrix information are used as the loaded matrix information to obtain low-bit convolution result matrix information; the low-bit convolution result matrix information is decoded into high-bit convolution result matrix information to realize digital optical tensor convolution operation.

[0055] Preferably, the system can be optimized and developed based on an optical system of any form that satisfies the object-image conjugate relationship between the multi-channel image matrix X and the convolution result matrix Y;

[0056] Preferably, the optical tensor convolution calculation system based on a multi-imaging projection architecture can further improve computing power through novel optical communication technology with expanded capacity. The system is characterized in that the light source array module simultaneously loads information from two or more multi-channel image matrices X using two or more optical signals with different characteristics. The characteristics of the optical signals include, but are not limited to, wavelength, mode, and polarization. After passing through the imaging projection module and the signal modulation module, the optical signals carrying multiple multi-channel image matrix information with different characteristics perform dot multiplication operations with the spectral information of the multi-channel convolution kernel matrix C. Finally, the detection module detects and obtains the convolution result matrix Y of multiple multi-channel image matrices V and the multi-channel convolution kernel matrix C.

[0057] By adopting the above technical solution, the present invention has at least one of the following beneficial effects compared with the prior art:

[0058] The optical tensor convolution computation system and method based on a multi-imaging projection architecture provided by this invention can perform optical tensor convolution computations of multi-channel image matrices and multi-channel convolutional matrices in a highly parallel, high-speed, and low-power manner, and obtain the convolution result matrix directly on the end face of the detection module (detector). Based on the system and method provided by this invention, tensor convolutions of arbitrary bit matrices with large-scale parallelism and sufficiently high accuracy can be efficiently computed. Moreover, tensor convolution is universal, and the obtained computation results are very easy to port to any other computing platform. By developing a high-speed spatial light modulator with higher contrast, optimizing a dedicated projection imaging system, and configuring a dedicated dot matrix light source, an optical tensor convolution processor with higher computing power and lower energy consumption compared to an electronic computer can be constructed. In addition, due to the characteristics of the imaging system itself, by cascading multiple 4f systems and employing additional multiplexing degrees of freedom, the computing power of the system can be multiplied, which is expected to construct a hybrid optoelectronic high-performance computing center or data center based on an optical tensor convolutional unit with a multi-imaging projection architecture.

[0059] The optical tensor convolution computation system and method based on a multi-imaging projection architecture provided by this invention significantly improves the pixel utilization of the spatial light modulator in the optical tensor convolution computation system, reduces the requirements for the dynamic detection range of the detection module, and can further improve the computational accuracy of the optical tensor convolution system. Compared with electronic AI accelerators and other optical computing solutions, the system and method provided by this invention can realize general-purpose, arbitrary-base digital optical tensor convolution computation, and has the characteristics of high speed, low power consumption, high parallelism, high tolerance, large scale, and reconfigurability. The system and method provided by this invention lays the research foundation for digital photonic tensor computation, and is expected to further develop digital optical computing systems based on matrix transformation, decomposition, and other operations. Its computing power far exceeds that of NVIDIA GPUs, and it has significant application value and good economic benefits in deep learning and other fields involving massive matrix operations.

[0060] This invention provides an optical tensor convolution computation system and method based on a multi-imaging projection architecture. This system offers truly large-scale parallelism and sufficiently high precision. By loading a multi-channel image matrix and a multi-channel convolution kernel matrix into the input module, a large-scale, high-precision optical tensor convolution computation result matrix can be directly obtained in the detection module after the optical signal passes through the system only once. The introduction of an imaging projection module enables the replication and processing of the spectral information of the multi-channel input image matrix, significantly improving the system's computing power. Furthermore, by adjusting the feature distance between the multi-channel image matrix and the imaging projection module and perfectly matching it with the diffraction angle of the beam splitter, various complex neural network structures can be implemented. This invention provides the first optical computing architecture that simultaneously reuses matrices on the object plane and spectral plane in a spatially interconnected optical system. Compared to previous optical computing architectures, this invention overcomes the limitations of spatially interconnected optical computing schemes when implementing small convolution kernels, allowing for full utilization of system resources to achieve higher computing power. Therefore, this invention can be truly applied in practical deep learning models, especially in real-world applications such as target recognition and autonomous driving, where convolutional neural networks are the core technology. Attached Figure Description

[0061] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0062] Figure 1 This is a schematic diagram illustrating the working principle of an optical tensor convolution calculation system based on a multi-imaging projection architecture in a preferred embodiment of the present invention. In this diagram, 101 is the light source array module, 102 is the beam splitter in the imaging projection module, 103 is the Fourier transform in the imaging module, 104 is the signal modulation module, 105 is the inverse Fourier transform in the detection module, and 106 is the detection module.

[0063] Figure 2 This is a schematic diagram of the architecture of an optical tensor convolution calculation system based on a multi-imaging projection architecture according to an embodiment of the present invention. In this diagram, 200 is the optical tensor convolution calculation part, 300 is the electronic control component, 201 is the light source array module, 202 is the imaging projection module, 2021 is the beam splitter in the imaging projection module, 2022 is the Fourier transform in the imaging module, 203 is the signal modulation module, 204 is the detection module, 2401 is the inverse Fourier transform in the detection module, 301 is the system initialization, 302 is the automatic control module, 303 is the parallel signal loading module, 3031 is the preprocessing, 3032 is the parallelization, 304 is the parallel signal download module, 3041 is the photoelectric conversion, and 3042 is the post-processing.

[0064] Figure 3This is a schematic diagram of an optical tensor convolution calculation system based on a multi-imaging projection architecture in a specific application example of the present invention. The optical elements are arranged in the following order: 401 is a VCSEL light source array, 402 is a Damman grating, 403 is a first Fourier transform lens, 404 is a spatial light modulator, 405 is a second Fourier transform lens, and 406 is a CMOS camera.

[0065] Figure 4 This is a flowchart of an optical convolution calculation method based on a multi-imaging projection architecture in one embodiment of the present invention. Detailed Implementation

[0066] The embodiments of the present invention are described in detail below: These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention.

[0067] A preferred embodiment of the present invention provides an optical tensor convolution computation system based on a multi-imaging projection architecture, comprising: an optical tensor convolution component, the optical tensor convolution component including: a light source array module, an imaging projection module disposed at the rear end of the light source array module, a signal modulation module disposed at the rear end of the imaging projection module, and a detection module disposed at the rear end of the signal modulation module; wherein:

[0068] The light source array module is used to load information from a multi-channel image matrix onto the input optical signal to obtain an optical signal carrying multi-channel image information.

[0069] The imaging projection module generates sub-beams of different diffraction orders from the optical signal carrying multi-channel image information. Each diffraction order carries all the information of the multi-channel image. Then, optical Fourier transform is performed on the sub-beams of different diffraction orders to obtain the spectral information of the multi-channel image matrix, which is then output to the signal modulation module.

[0070] The signal modulation module is used to load the spectral information of the multi-channel convolution kernel matrix and perform frequency domain multiplication operation on the spectral information of the multi-channel image matrix carried in the sub-beams of different diffraction orders and the spectral information of the multi-channel convolution kernel matrix to obtain the frequency domain multiplication information of the multi-channel image matrix and the multi-channel convolution kernel matrix, and then obtain the optical signal carrying the frequency domain multiplication information.

[0071] The detection module is used to perform optical inverse Fourier transform on the optical signal carrying frequency domain multiplication information to obtain the optical tensor convolution result of the multi-channel image matrix and the multi-channel convolution kernel matrix.

[0072] In a preferred embodiment, the imaging projection module includes multiple beam splitters that divide the optical signal carrying multi-channel image matrix information into several sub-beams. Then, the sub-beams undergo optical Fourier transform to obtain the spectral information of the multi-channel image matrix. Each sub-beam carries the spectral information of the multi-channel image matrix and is incident on different positions of the signal modulation module according to its respective diffraction angle, achieving parallel translation of the optical signal carrying the spectral information of the multi-channel image matrix. The spectral information of the multi-channel image matrix carried in the sub-beams incident on different positions of the signal modulation module is multiplied at the pixel level with the spectral information of the multi-channel convolution kernel matrix, achieving parallel frequency domain multiplication of the multi-channel image matrix and the multi-channel convolution kernel matrix.

[0073] As a preferred embodiment, the parallel frequency domain multiplication operation of the multi-channel image matrix and the multi-channel convolution kernel matrix includes:

[0074] The replication of a multi-channel image matrix is ​​achieved through multiple beam splitters in the imaging projection module;

[0075] The optical Fourier transform of the copied multi-channel image matrix is ​​performed to obtain its spectral information, which is then projected onto different positions of the signal modulation module.

[0076] The spectral information of the convolution kernel matrix of different channels is loaded at different positions of the modulation module to realize the dot product operation between the elements in the spectral information of the multi-channel image matrix and the corresponding elements in the spectrum of the multi-channel convolution kernel matrix.

[0077] The light beams at different locations carry dot product information of the spectrum of their respective convolution kernel matrix and the spectrum of the multi-channel image matrix;

[0078] Optical inverse Fourier transform is performed on the frequency domain multiplication information at all positions to obtain the optical tensor convolution result of the multi-channel image matrix and the multi-channel convolution kernel matrix. The characteristic is that the optical signals carrying the frequency domain multiplication information at different positions are incident on different positions of the detection surface of the detection module according to their respective diffraction tilt angles.

[0079] In this preferred embodiment, based on the specific neural network structure, by adjusting the diffraction tilt angle of multiple beam splitters and the feature distance between the imaging projection module and the light source array module, the photon convolution results of the convolution kernel matrix of different channels and the multi-channel image matrix are added together to obtain the final result after convolution of different neural network structures.

[0080] In a preferred embodiment, the signal loading surface of the light source array module and the modulation plane of the signal modulation module satisfy the object-image conjugate relationship, and the modulation plane of the signal modulation module and the detection surface of the detection module also satisfy the object-image conjugate relationship.

[0081] In a preferred embodiment, by adjusting the diffraction tilt angles of multiple beam splitters and the characteristic distance between the imaging projection module and the light source array module, the diffraction angles of each diffraction order of the imaging projection module and the corresponding positions of different channel convolution kernels in the signal modulation module are aligned and matched, thereby realizing the dot product operation of the spectral information of the multi-channel image matrix and the multi-channel convolution kernel matrix.

[0082] As a preferred embodiment, the optical tensor convolution calculation system further includes any one or more of the following:

[0083] As a preferred embodiment, the light source array module adopts a light-emitting element array, a fiber array, or a spatial light modulator;

[0084] As a preferred embodiment, the imaging projection module includes a Fourier transform lens and one or more Damman gratings, wherein the Damman gratings are one-dimensional Damman gratings or two-dimensional Damman gratings.

[0085] As a preferred embodiment, the imaging projection module employs a metasurface beam splitter or a metamaterial beam splitter.

[0086] As a preferred embodiment, the signal modulation module employs a spatial light modulator or an optical mask;

[0087] In a preferred embodiment, the detection module includes a Fourier transform lens and an optical receiver array.

[0088] In a preferred embodiment, the system further includes an electronic control component: the electronic control component includes a data parallel loading module, a data parallel downloading module, and an automation control module;

[0089] The parallel data loading module is used to preprocess the multi-channel image matrix and multi-channel convolution kernel matrix, and then load the preprocessed information into the light source array module and the signal modulation module in parallel.

[0090] The parallel data download module is used to convert the optical signals detected by the detection module into electrical signals in parallel, and to perform subsequent processing on the electrical signals;

[0091] An automated control module is used to automatically adjust the optical tensor convolution component according to different algorithm structures, and change the projection translation step size of the imaging projection module to achieve the addition of optical tensor convolution results.

[0092] As a preferred embodiment, the present invention also provides an optical tensor convolution calculation method based on a multi-imaging projection architecture, characterized in that it includes:

[0093] The input optical signal is loaded with multi-channel image matrix information to obtain an optical signal carrying multi-channel image matrix information;

[0094] Sub-beams of different diffraction orders are generated from optical signals carrying multi-channel image matrix information;

[0095] Optical Fourier transform is performed on all diffraction orders of the sub-beams to obtain their corresponding spectral information;

[0096] In the signal modulation module, the spectral information of the convolution kernel matrix corresponding to each diffraction order sub-beam is loaded, and the spectral information of the multi-channel image matrix carried by the sub-beams of different diffraction orders is multiplied by the spectral information of the multi-channel convolution kernel matrix to obtain the optical signal carrying the frequency domain multiplication information of the two at all positions.

[0097] An optical inverse Fourier transform is performed on the optical signal carrying frequency domain multiplication information, and the multi-channel image matrix and the optical tensor convolution result of the convolution matrix are obtained in the detection module.

[0098] In a preferred embodiment, sub-beams of different diffraction orders are transmitted to different positions of the signal modulation module at different diffraction angles, wherein the optical signal of each diffraction order carries all the information of the multi-channel image matrix.

[0099] In a preferred embodiment, a multi-channel image matrix is ​​loaded in the light source array module.

[0100] ,

[0101] in, The image matrix for the first channel. The image matrix of the m-th channel, It is an image matrix with n×(m−1)+1 channels. Let be the image matrix of the n×m-th channel, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively. O is the zero matrix between adjacent image matrices. The number of zero elements in the zero matrix must be sufficient to ensure that the convolution results of adjacent image matrices do not overlap.

[0102] Perform Fourier transform on the multi-channel image matrix.

[0103] ,

[0104] Among them, v 11 v 12 …v nm It is the spectral information of the multi-channel image matrix, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively.

[0105] The spectral information of the multi-channel image matrix is ​​replicated by multiple beam splitters in the imaging projection module. Each sub-beam of each diffraction order carries the complete spectral information and propagates along its respective diffraction direction to different positions on the modulation surface of the signal modulation module.

[0106] ,

[0107] Among them, V 11 V 12 …V kp It is the spectral information of the multi-channel image matrix at different positions of the modulation surface of the signal modulation module, and the subscripts k and p are the row number and column number of the sub-beam, respectively.

[0108] Multiply the spectral information of the multi-channel image matrix with the spectral information of the convolution kernel matrix.

[0109] ,

[0110] Among them, C 11 It is the spectrum of the first channel convolution kernel matrix, C 12 It is the spectrum of the second channel convolution kernel matrix, C kp Y is the spectrum of the convolution kernel matrix of the p×(k−1)+1th channel. 11 Y is the dot product of the spectrum of the first channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. 12 Y is the dot product of the spectrum of the second channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. kp It is the dot product of the spectrum of the convolution kernel matrix of the p×(k−1)+1th channel and the spectrum information of the multi-channel image matrix.

[0111] Perform an inverse Fourier transform on the dot product of the spectrum of the multi-channel convolution kernel matrix and the spectrum information of the multi-channel image matrix.

[0112] ,

[0113] Among them, y 11 It is the convolution result of the first channel convolution kernel matrix and the multi-channel image matrix, y 12 It is the convolution result of the second channel convolution kernel matrix and the multi-channel image matrix, y kp It is the convolution result of the p×(k−1)+1th channel convolution kernel matrix and the multi-channel image matrix.

[0114] In a preferred embodiment, the loaded multi-channel image information, the loaded multi-channel convolution kernel information, and the obtained convolution result matrix information are all analog quantities. The convolution result matrix information is quantized into a digital result after digital processing to realize analog optical tensor convolution operation.

[0115] In a preferred embodiment, the high-bit multi-channel convolution kernel matrix information and the high-bit multi-channel image matrix information to be processed are represented as multiple encoded low-bit matrices, respectively, to obtain encoded low-bit multi-channel convolution kernel matrix information and low-bit multi-channel image matrix information; the low-bit multi-channel convolution kernel matrix information and the low-bit multi-channel image matrix information are used as loaded matrix information to obtain low-bit convolution result matrix information; the low-bit convolution result matrix information is decoded into high-bit convolution result matrix information to realize digital optical tensor convolution operation.

[0116] As a preferred embodiment, the light source array module can be an array of light-emitting elements, including but not limited to LEDs, LDs, fiber arrays, vertical cavity surface-mount semiconductor laser arrays (VCSELs), or spatial light modulators (SLMs), including but not limited to liquid crystal spatial light modulators (LCSLMs), digital micromirror arrays (DMDs), microelectromechanical systems (MEMS), fiber arrays, and optical waveguide arrays.

[0117] As a preferred embodiment, the large-scale beam splitter can be a one-dimensional Damman grating with a beam splitting ratio of 1×3 to 1×128, or a two-dimensional Damman grating with a beam splitting ratio of 3×3 to 128×128 or larger. By combining Damman gratings, higher diffraction efficiency and larger-scale beam splitting effect can be achieved. Other beam splitting elements can also be used, including but not limited to diffractive optical elements (DOE), low-dimensional functional materials, metasurfaces, and metamaterials.

[0118] As a preferred embodiment, the signal modulation module includes, but is not limited to, various spatial light modulators (SLMs), such as liquid crystal spatial light modulators (LDSLMs), digital micromirror arrays (DMDs), microelectromechanical systems (MEMS), fiber arrays, and optical waveguide arrays.

[0119] In a preferred embodiment, the detection module includes a Fourier transform lens and an optical receiver array. The lens is used to perform the inverse Fourier transform of the spectral dot product information, and the optical receiver array is used to detect the optical tensor convolution signal, including but not limited to CMOS, CCD and various photodetector arrays.

[0120] The optical tensor convolution calculation system based on the multi-imaging projection architecture provided in this preferred embodiment can be optimized and developed based on any form of optical system that satisfies the object-image conjugate relationship between the multi-channel image matrix and the optical tensor convolution result matrix. It can be a transmissive optical system, a refractive optical system, or a reflective optical system. It can be a paraxial optical system, an off-axis optical system, or a planar space optical waveguide system. It includes, but is not limited to, various deformations and combinations of single or multiple cascaded 4f canonical signal systems, microscopic systems, telescope systems, projection systems, and the above optical systems.

[0121] The optical tensor convolution computing system based on a multi-imaging projection architecture provided in this preferred embodiment can further improve computing power through novel optical communication technologies that expand capacity, including but not limited to wavelength division multiplexing, mode division multiplexing, and polarization multiplexing.

[0122] Compared to currently proposed technical solutions, the optical tensor convolution computation system and method based on a multi-imaging projection architecture provided in the above embodiments of the present invention can simultaneously reuse multi-channel input image matrices and multi-channel convolution kernel matrices in both the object plane and the spectral plane. This allows for full utilization of pixels in both the object plane and the spectral plane of the optical system to achieve complex convolutional neural networks. The present invention is expected to enable the construction of practical optical tensor convolution computation systems in real-world application scenarios, demonstrating significant application value and good economic benefits in fields such as target recognition and autonomous driving.

[0123] The technical solutions provided by the above embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0124] This invention implements optical tensor convolution operations based on the principle of Optical Multiple Imaging Projection (OMica). Figure 1This is a schematic diagram of an optical tensor convolution calculation system, including a light source array module 101, a beam splitter 102 in the imaging projection module, a Fourier transform 103 in the imaging module, a signal modulation module 104, an inverse Fourier transform 105 in the detection module, and a detection module 106. The light source array module 101 loads information from the multi-channel image matrix, the signal modulation module 104 loads the spectral information of the multi-channel convolution kernel matrix, and the detection module 106 detects the information of the optical tensor convolution result matrix. The beam splitter 102 in the imaging projection module performs replication and translation of the multi-channel image matrix, and the Fourier transform 103 in the imaging projection module performs Fourier transforms on sub-beams to obtain the spectral information of the multi-channel image matrix. Furthermore, the Fourier transform 103 in the imaging projection module and the inverse Fourier transform in the detection module form a Fourier transform pair. In a preferred embodiment, the Fourier transform relationship can be implemented by a 4f optical system. The plane containing the multi-channel image matrix and the plane containing the optical tensor convolution result matrix satisfy the object-image conjugate relationship. The plane containing the multi-channel convolution kernel matrix, the plane containing the multi-channel image matrix, and the optical tensor convolution result matrix respectively satisfy the Fourier transform relationship. After adding imaging projection modules 102 and 103, a perfect match is achieved by changing the characteristic distance d between the beam splitter 102 and the light source array module 101 in the imaging projection module and the diffraction angle of the beam splitter 102 in the imaging projection module. Each diffraction order of the beam splitter 102 in the imaging projection module carries all the information of the convolution kernel matrix B, and its spectral information is obtained through the Fourier transform 103 in the imaging projection module. Then, sub-beams with different spectral components are incident on different positions of the modulation plane in the signal modulation module 104 and perform dot product operations with the convolution kernel matrices of each channel loaded at different positions. Finally, the frequency domain dot product information is inversely transformed by the inverse Fourier transform 105 in the modulation module, and the optical tensor convolution result 106 is output by the detector array in the detection module. The optical tensor convolution matrix calculation system based on a multi-imaging projection architecture proposed in this invention has the potential to realize a novel optical computing system with ultra-high computing power and ultra-low power consumption.

[0125] The schematic diagram of the optical tensor convolution calculation system based on a multi-imaging projection architecture provided in the above embodiments of the present invention is as follows: Figure 2As shown. The architecture includes an optical tensor convolution calculation section 200, an electronic control component 300, a light source array module 201, an imaging projection module 202, a beam splitter 2021 in the imaging projection module, a Fourier transform 2022 in the imaging module, a signal modulation module 203, a detection module 204, and an inverse Fourier transform 2401 in the detection module. The light source array module 210 is used to load information from a multi-channel image matrix. The light source array module 210 can be an array of light-emitting elements, including but not limited to LEDs, LDs, fiber optic arrays, and VCSELs, or a spatial light modulator (SLM), including but not limited to liquid crystal spatial light modulators, digital micromirror arrays (DMDs), and microelectromechanical systems (MEMS). The system includes fiber optic arrays and waveguide arrays. A signal modulation module 203 modulates the optical signal after passing through the imaging projection module 202. The signal modulation module 203 loads the spectral information of a multi-channel convolution kernel matrix. The signal modulation module 203 includes, but is not limited to, various spatial light modulators (SLMs), such as liquid crystal spatial light modulators, digital micromirror arrays (DMDs), microelectromechanical systems (MEMS), fiber optic arrays, and waveguide arrays. A detection module 204 performs convergence summation and detection of the convolution results. Sub-beams of different diffraction orders are obtained through the beam splitter 2021 in the imaging projection module 202, and the spectral information of the sub-beams is obtained through the Fourier transform 2022 in the imaging projection module 202. The signal modulation module 203 performs a dot product of the spectrum of the multi-channel image matrix and the spectrum of the multi-channel convolution kernel matrix. Finally, the inverse Fourier transform 2401 in the detection module 204 yields the optical tensor convolution result matrix.

[0126] The electronic control component 300 mainly includes: system initialization 301, automatic control module 302, parallel signal loading module 303, preprocessing 3031, parallelization 3032, parallel signal downloading module 304, photoelectric conversion 3041, and post-processing 3042. System initialization 301 initializes and sets the parameters of each system. Automatic control module 302, based on the results of system initialization 301 and the specific neural network architecture, sets the relevant parameters for the imaging module of the optical tensor convolution component. It achieves the overlap, addition, or separation of the optical tensor convolution result matrix by adjusting the feature distance between the beam splitter 2021 and the light source array module 201 in the imaging projection and by replacing the beam splitter. The parallel signal loading module 303 loads the multi-channel image matrix and multi-channel convolution kernel matrix required for a single convolution process into the corresponding light source array module 201 and signal modulation module 203. Before loading the signal, it performs data preprocessing 3031 and parallelization 3032. The parallel signal download module 304 downloads the optical tensor convolution result matrix from the detection module. First, the optical signal is converted into an electrical signal by photoelectric conversion 3401, and then post-processing 3402 is performed to obtain the final convolution result. The final convolution result is transmitted to the parallel signal loading module 303 according to the control signal of the automatic control module 302 to enter the next convolution process or directly output as the final convolution result.

[0127] A preferred specific application illustration of the optical tensor convolution calculation system based on a multi-imaging projection architecture provided in the above embodiments of the present invention is shown below. Figure 3 As shown. The optical components, in order, are: VCSEL light source array 401, Damman grating 402, first Fourier transform lens 403, spatial light modulator 404, second Fourier transform lens 405, and CMOS camera 406.

[0128] The light source array 401 is a VCSEL light source array with a wavelength of 450nm. Its beam arrangement ratio is 526×526, and the spacing between the beams is 40 μm. The light source array 401 modulates the light source array by controlling the switch of the input current and completes the loading of multi-channel image matrix information.

[0129] The Damman grating 402 has a beam splitting ratio of 1:10, a period of 1 mm, and an adjacent order diffraction angle of 0.45 mrad.

[0130] A Damman grating 402 is placed at a distance *d* behind the light source array 401. It uniformly splits the optical signal, loaded with multi-channel image matrix information from the light source array 401, into a 10×10 array of parallel light. This parallel light then undergoes an optical Fourier transform via a first Fourier transform lens 403, yielding the spectral information of the 10×10 array of parallel light on the rear focal plane of the first Fourier transform lens 403. Next, a spatial light modulator 404 is placed on the rear focal plane of the first Fourier transform lens 403. The spectral information of the 10×10 array of parallel light is incident on different positions of the spatial light modulator 404, and is multiplied by the spectral information of the convolution kernel matrices of different channels loaded on it.

[0131] The back focal plane of the first Fourier transform lens 403 is the front focal plane of the spatial modulator 404, and the back focal plane of the second Fourier transform lens 405 is the plane where the CMOS camera 406 is located. The first Fourier transform lens 403 has a focal length of 300mm and an aperture of 50mm, while the second Fourier transform lens 405 has a focal length of 400mm and an aperture of 50mm.

[0132] Spatial light modulator 404 is selected as a liquid crystal spatial light modulator with a resolution of 1920×1080 and a single pixel size of 8 μm;

[0133] A single coded element of the spatial light modulator 404 is represented by 54×54 pixels, corresponding to a matrix unit size of 432μm;

[0134] The spatial light modulator 404 loads the multi-channel convolution kernel matrix by adjusting the transmittance at different positions on the control panel.

[0135] The detection module uses a CMOS camera 406 with a resolution of 1920×1200 and a single pixel size of 4.8μm. The CMOS detector receives the final optical tensor convolution result matrix.

[0136] Based on the working principle of the optical tensor convolution system based on the multi-imaging projection architecture, the photon tensor process is described using a four-channel 3×3 input image matrix and a four-channel 3×3 matrix as examples:

[0137] First, the four-channel 3×3 input image matrix is ​​represented as follows:

[0138] Its Fourier transform result is:

[0139] The four channel convolution kernel matrices are as follows:

[0140] , , , ,

[0141] The convolution result of the four-channel image matrix and the four-channel convolution kernel matrix is:

[0142] ,

[0143] By adjusting the diffraction angle of the beam splitter in the imaging projection module and the distance between the beam splitter and the light source array module, the photon convolution results of the four channels can be summed.

[0144] ,

[0145] In summary, compared with traditional optical convolution schemes or existing optical computing schemes, the optical tensor convolution computing system based on a multi-imaging projection architecture proposed in this invention achieves large-scale, high-precision, and highly parallel optical tensor convolution computation. By introducing an imaging projection module and effectively utilizing the diffraction effect of beam splitters, such as Dammann gratings, high-throughput optical tensor convolution computation is achieved. Furthermore, higher-order Dammann gratings can be completely filtered by the aperture without affecting the computation results. By optimizing and combining Dammann gratings with different splitting ratios, efficient beam splitting of larger matrices with almost no energy loss can be achieved. Therefore, the physical size of the matrix elements can be significantly reduced. Finally, based on a suitable numerical encoding algorithm, by optimizing the optical system and using higher-contrast SLMs (including MEMS and DMD), higher precision and larger-scale convolution computation can be achieved.

[0146] Furthermore, if high-dimensional multiplexing methods, including those involving polarization, wavelength, and mode, are used to improve optical communication capacity, at least 10⁻⁶ dimensionality multiplexing can be achieved. 2 Up to 10 3 The system achieves a 100-fold increase in computing power. By using high-performance devices (such as large modulators with higher update frequencies, detectors or detector arrays with wider dynamic ranges and higher sampling frequencies) to progressively improve the computational power and energy efficiency of convolution, the proposed optical convolutional computing architecture can continuously improve computing power and accuracy, and has good scalability, opening a new door to realizing large-scale, high-precision optical convolutional computing.

[0147] Figure 4 This is a flowchart of an optical tensor convolution calculation method based on a multi-imaging projection architecture, provided as an embodiment of the present invention.

[0148] like Figure 4 As shown, the optical tensor convolution calculation method based on a multi-imaging projection architecture provided in this embodiment may include the following steps:

[0149] S100: Load multi-channel image matrix information onto the input optical signal to obtain an optical signal carrying multi-channel image information;

[0150] S200 replicates and translates optical signals carrying multi-channel image information to generate sub-beams of different diffraction orders;

[0151] S300 performs Fourier transform on sub-beams of different diffraction orders;

[0152] S400 loads the spectral information of the multi-channel convolution kernel matrix and performs a dot product operation between the sub-beams of different diffraction orders and the convolution kernel matrix of the corresponding channel.

[0153] S500 performs an inverse Fourier transform on the spectral dot product information to obtain the optical tensor convolution result matrix.

[0154] In a preferred embodiment, in S200, sub-beams of different diffraction orders are transmitted to different positions of the signal modulation module at different angles, wherein the optical signal of each diffraction order carries all the information of the multi-channel image matrix spectrum.

[0155] In a preferred embodiment, a multi-channel image matrix is ​​loaded in the light source array module.

[0156] ,

[0157] in, The image matrix for the first channel. The image matrix of the m-th channel, It is an image matrix with n×(m−1)+1 channels. Let be the image matrix of the n×m-th channel, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively. O is the zero matrix between adjacent image matrices. The number of zero elements in the zero matrix must be sufficient to ensure that the convolution results of adjacent image matrices do not overlap.

[0158] Perform Fourier transform on the multi-channel image matrix.

[0159] ,

[0160] Among them, v 11 v 12 …v nm It is the spectral information of the multi-channel image matrix, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively.

[0161] The spectral information of the multi-channel image matrix is ​​replicated by multiple beam splitters in the imaging projection module. Each sub-beam of each diffraction order carries the complete spectral information and propagates along its respective diffraction direction to different positions on the modulation surface of the signal modulation module.

[0162] ,

[0163] Among them, V 11 V 12 …V kp It is the spectral information of the multi-channel image matrix at different positions of the modulation surface of the signal modulation module, and the subscripts k and p are the row number and column number of the sub-beam, respectively.

[0164] Multiply the spectral information of the multi-channel image matrix with the spectral information of the convolution kernel matrix.

[0165] ,

[0166] Among them, C 11 It is the spectrum of the first channel convolution kernel matrix, C 12 It is the spectrum of the second channel convolution kernel matrix, C kp Y is the spectrum of the convolution kernel matrix of the p×(k−1)+1th channel. 11 Y is the dot product of the spectrum of the first channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. 12 Y is the dot product of the spectrum of the second channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. kp It is the dot product of the spectrum of the p×(k−1)+1 channel convolution kernel matrix and the spectrum information of the multi-channel image matrix.

[0167] Perform an inverse Fourier transform on the dot product of the spectrum of the multi-channel convolution kernel matrix and the spectrum information of the multi-channel image matrix.

[0168] ,

[0169] Among them, y 11 It is the convolution result of the first channel convolution kernel matrix and the multi-channel image matrix, y 12 It is the convolution result of the second channel convolution kernel matrix and the multi-channel image matrix, y kp It is the convolution result of the p×(k−1)+1th channel convolution kernel matrix and the multi-channel image matrix.

[0170] In a preferred embodiment, the loaded multi-channel image information, the loaded multi-channel convolution kernel information, and the obtained convolution result matrix information are all analog quantities. The convolution result matrix information is quantized into a digital result after digital processing to realize analog optical tensor convolution operation.

[0171] In a preferred embodiment, the high-bit multi-channel convolution kernel matrix information and the high-bit multi-channel image matrix information to be processed are represented as multiple encoded low-bit matrices, respectively, to obtain encoded low-bit multi-channel convolution kernel matrix information and low-bit multi-channel image matrix information; the low-bit multi-channel convolution kernel matrix information and the low-bit multi-channel image matrix information are used as loaded matrix information to obtain low-bit convolution result matrix information; the low-bit convolution result matrix information is decoded into high-bit convolution result matrix information to realize digital optical tensor convolution operation.

[0172] It should be noted that the steps in the method provided by the present invention can be implemented using corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the system to implement the steps and flow of the method. That is, the embodiments in the system can be understood as preferred examples of the method, and will not be elaborated here.

[0173] The optical convolution calculation system and method based on a multi-imaging projection architecture provided in the above embodiments of the present invention achieve high parallelism and high precision convolution calculation of the convolution kernel matrix B and the input matrix A. Using the technical solution provided in the above embodiments of the present invention, by simply loading the convolution kernel matrix B and the input matrix A onto the corresponding light source array module and signal modulation module, a large-scale, high-precision convolution operation result matrix C can be directly obtained on the detection module after the optical signal passes through the optical convolution calculation system in a single pass. The optical convolution calculation system and method based on a multi-imaging projection architecture proposed in the above embodiments of the present invention can realize a new technical route for large-scale, high-precision, and fully parallel optical computing, providing a general and efficient solution to meet the needs of artificial intelligence, neural network image processing, and other tasks for massive convolution operations.

[0174] The optical convolution calculation system and method based on a multi-imaging projection architecture disclosed above represent only one specific embodiment of the present invention and should not be construed as limiting the scope of protection of the present invention. It should be noted that those skilled in the art can make several non-inventive modifications and improvements to the specific implementation details and representative devices proposed in this patent without departing from the basic idea of ​​the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

[0175] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples, without contradiction. Having described the invention in such detail and with reference to its preferred embodiments, it will be apparent that modifications and variations are possible without departing from the scope of the invention as defined in the appended claims.

[0176] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. An optical tensor convolution computation system based on a multi-imaging projection architecture, characterized in that, include: An optical tensor convolution component, comprising, in sequence according to the optical path propagation direction: a light source array module, an imaging projection module, a signal modulation module, and a detection module; wherein: The light source array module is used to load information from a multi-channel image matrix onto the input optical signal to obtain an optical signal carrying multi-channel image information. The imaging projection module is used to generate sub-beams of different diffraction orders from the optical signal carrying multi-channel image information, wherein each diffraction order sub-beam carries all the information of the multi-channel image; then, optical Fourier transform is performed on the sub-beams of different diffraction orders to obtain the spectral information of the multi-channel image matrix and output to the signal modulation module. The signal modulation module is used to load the spectral information of the multi-channel convolution kernel matrix, and to perform a frequency domain multiplication operation on the spectral information of the multi-channel image matrix carried in the sub-beams of different diffraction orders and the spectral information of the multi-channel convolution kernel matrix to obtain the frequency domain multiplication information of the multi-channel image matrix and the multi-channel convolution kernel matrix, thereby obtaining an optical signal carrying the frequency domain multiplication information. The detection module is used to perform an optical inverse Fourier transform on the optical signal carrying frequency domain multiplication information to obtain the optical tensor convolution result of the multi-channel image matrix and the multi-channel convolution kernel matrix; The imaging projection module includes multiple beam splitters that divide the optical signal carrying multi-channel image matrix information into several sub-beams. Then, an optical Fourier transform is performed on each sub-beam to obtain the spectral information of the multi-channel image matrix. Each sub-beam carries the spectral information of the multi-channel image matrix and is incident on different positions of the signal modulation module according to its respective diffraction angle, achieving parallel translation of the optical signal carrying the spectral information of the multi-channel image matrix. The spectral information of the multi-channel image matrix carried in the sub-beams incident on different positions of the signal modulation module is multiplied at the pixel level with the spectral information of the multi-channel convolution kernel matrix, achieving parallel frequency domain multiplication between the multi-channel image matrix and the multi-channel convolution kernel matrix. The parallel frequency domain multiplication operation between the multi-channel image matrix and the multi-channel convolution kernel matrix includes: The multi-channel image matrix is ​​replicated using the multiple beam splitters in the imaging projection module; The copied multi-channel image matrix is ​​subjected to optical Fourier transform to obtain its spectral information, and then projected onto different positions of the signal modulation module. The spectral information of the convolution kernel matrix of different channels is loaded at different positions of the modulation module to realize the dot product operation between the elements in the spectral information of the multi-channel image matrix and the corresponding elements in the spectrum of the multi-channel convolution kernel matrix. The light beams at different positions carry the dot product information of the spectrum of their respective convolution kernel matrix and the spectrum of the multi-channel image matrix; Optical tensor convolution results of multi-channel image matrix and multi-channel convolution kernel matrix are obtained by performing optical inverse Fourier transform on the frequency domain multiplication information at all positions. The characteristic is that optical signals carrying the frequency domain multiplication information at different positions are incident on different positions of the detection surface of the detection module according to their respective diffraction tilt angles. Based on the specific neural network structure, by adjusting the diffraction tilt angle of the multiple beam splitters and the feature distance between the imaging projection module and the light source array module, the optical convolution results of the convolution kernel matrix of different channels and the multi-channel image matrix are added together to obtain the final result after convolution of different neural network structures.

2. The optical tensor convolution calculation system based on a multi-imaging projection architecture according to claim 1, characterized in that, The signal loading plane of the light source array module and the modulation plane of the signal modulation module satisfy a Fourier transform relationship, and the modulation plane of the signal modulation module and the detection plane of the detection module also satisfy a Fourier transform relationship.

3. The optical tensor convolution calculation system based on a multi-imaging projection architecture according to claim 1, characterized in that, By adjusting the diffraction tilt angles of multiple beam splitters and the characteristic distance between the imaging projection module and the light source array module, the diffraction angles of each diffraction order of the imaging projection module and the corresponding positions of different channel convolution kernels in the signal modulation module are aligned and matched, thereby realizing the dot product operation of the spectral information of the multi-channel image matrix and the multi-channel convolution kernel matrix.

4. The optical tensor convolution calculation system based on a multi-imaging projection architecture according to claim 1, characterized in that, The system also includes any one or more of the following: - The light source array module adopts a light-emitting element array, a fiber array, or a spatial light modulator; - The imaging projection module includes a Fourier transform lens and one or more Damman gratings, wherein the Damman gratings are one-dimensional Damman gratings or two-dimensional Damman gratings; - The imaging projection module uses a metasurface beam splitter or a metamaterial beam splitter; - The signal modulation module employs a spatial light modulator or an optical mask; - The detection module includes a Fourier transform lens and an optical receiver array.

5. The optical tensor convolution computation system based on a multi-imaging projection architecture according to any one of claims 1-4, characterized in that, Also includes: Electronic control components: The electronic control components include a data parallel loading module, a data parallel downloading module, and an automation control module; The data parallel loading module is used to preprocess the multi-channel image matrix and the multi-channel convolution kernel matrix, and load the preprocessed information in parallel into the light source array module and the signal modulation module. The data parallel download module is used to convert the optical signals detected by the detection module into electrical signals in parallel, and to perform subsequent processing on the electrical signals; The automated control module is used to automatically adjust the optical tensor convolution component according to different algorithm structures, and change the projection translation step size of the imaging projection module to achieve the addition of optical tensor convolution results.

6. A method for calculating optical tensor convolution based on a multi-imaging projection architecture, implemented using the system described in claim 1, characterized in that, include: The input optical signal is loaded with multi-channel image matrix information to obtain an optical signal carrying multi-channel image matrix information; Sub-beams of different diffraction orders are generated from the optical signal carrying multi-channel image matrix information; Optical Fourier transform is performed on all diffraction orders of the sub-beams to obtain their corresponding spectral information; In the signal modulation module, the spectral information of the convolution kernel matrix corresponding to each diffraction order sub-beam is loaded, and the spectral information of the multi-channel image matrix carried by the sub-beams of different diffraction orders is multiplied by the spectral information of the multi-channel convolution kernel matrix to obtain the optical signal carrying the frequency domain multiplication information of the two at all positions. An optical inverse Fourier transform is performed on the optical signal carrying frequency domain multiplication information, and the optical tensor convolution result of the multi-channel image matrix and the convolution kernel matrix is ​​obtained in the detection module.

7. The optical tensor convolution calculation method based on a multi-imaging projection architecture according to claim 6, characterized in that, The sub-beams of different diffraction orders are transmitted to different positions of the signal modulation module at different diffraction angles, wherein the optical signal of each diffraction order carries all the information of the multi-channel image matrix.

8. The optical tensor convolution calculation method based on a multi-imaging projection architecture according to claim 6, characterized in that, Load a multi-channel image matrix into the light source array module. , in, The image matrix for the first channel. The image matrix of the m-th channel, It is an image matrix with n×(m−1)+1 channels. It is the image matrix of the n×m channel, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively. O is the zero matrix between adjacent image matrices; the number of zero elements in the zero matrix is ​​required to ensure that the result of convolution of adjacent image matrices does not overlap. Perform a Fourier transform on the multi-channel image matrix. , Among them, v 11 v 12 …v nm It is the spectral information of the multi-channel image matrix, where n and m are the number of rows and columns of each element in the multi-channel image matrix, respectively; The spectral information of the multi-channel image matrix is ​​replicated by multiple beam splitters in the imaging projection module. Each sub-beam of each diffraction order carries the complete spectral information and propagates along its respective diffraction direction to different positions on the modulation surface of the signal modulation module. , Among them, V 11 V 12 …V kp It is the spectral information of the multi-channel image matrix at different positions of the modulation surface of the signal modulation module, and the subscripts k and p are the row number and column number of the sub-beam, respectively; Multiply the spectral information of the multi-channel image matrix with the spectral information of the convolution kernel matrix. , Among them, C 11 It is the spectrum of the first channel convolution kernel matrix, C 12 It is the spectrum of the second channel convolution kernel matrix, C kp Y is the spectrum of the convolution kernel matrix of the p×(k−1)+1th channel; 11 Y is the dot product of the spectrum of the first channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. 12 Y is the dot product of the spectrum of the second channel convolution kernel matrix and the spectrum information of the multi-channel image matrix. kp It is the dot product of the spectrum of the convolution kernel matrix of the p×(k−1)+1th channel and the spectrum information of the multi-channel image matrix; Perform an inverse Fourier transform on the dot product of the spectrum of the multi-channel convolution kernel matrix and the spectrum of the multi-channel image matrix. , Among them, y 11 It is the convolution result of the first channel convolution kernel matrix and the multi-channel image matrix, y 12 It is the convolution result of the second channel convolution kernel matrix and the multi-channel image matrix, y kp It is the convolution result of the p×(k−1)+1th channel convolution kernel matrix and the multi-channel image matrix.

9. The optical tensor convolution calculation method based on a multi-imaging projection architecture according to any one of claims 6-8, characterized in that, The loaded multi-channel image information, the loaded multi-channel convolution kernel information, and the obtained convolution result matrix information are all analog quantities. The convolution result matrix information is quantized into a digital result after digital processing to realize analog optical tensor convolution operation.

10. The optical tensor convolution calculation method based on a multi-imaging projection architecture according to any one of claims 6-8, characterized in that, The high-bit multi-channel convolution kernel matrix information and the high-bit multi-channel image matrix information to be processed are represented as multiple encoded low-bit matrices, respectively, to obtain encoded low-bit multi-channel convolution kernel matrix information and low-bit multi-channel image matrix information; the low-bit multi-channel convolution kernel matrix information and the low-bit multi-channel image matrix information are respectively used as loaded matrix information to obtain low-bit convolution result matrix information; the low-bit convolution result matrix information is decoded into high-bit convolution result matrix information to realize digital optical tensor convolution operation.

11. The optical tensor convolution calculation method based on a multi-imaging projection architecture according to any one of claims 6-8, characterized in that, The optical tensor convolution calculation system based on the multi-imaging projection architecture is optimized and developed based on an optical system of arbitrary form that satisfies the object-image conjugate relationship between the multi-channel image matrix X and the convolution result matrix Y. The optical tensor convolution computing system based on a multi-imaging projection architecture can further improve computing power through novel optical communication technologies that expand capacity. Its characteristic is that the light source array module simultaneously loads information from two or more multi-channel image matrices X using two or more optical signals with different characteristics. The characteristics of the optical signals include, but are not limited to, wavelength, mode, and polarization. After passing through the imaging projection module and the signal modulation module, the various optical signals carrying information from multiple multi-channel image matrices perform dot-multiplication operations with the spectral information of the multi-channel convolution kernel matrix C. Finally, the detection module detects and obtains the convolution result matrix Y of multiple multi-channel image matrices V and the multi-channel convolution kernel matrix C.