Optical computing system and method for implementing optical neural network operation

By fusing convolution-batch normalization-activation operations in an optical computing system and using femtosecond lasers to control a two-dimensional material array, the problem of fixed transfer functions in traditional optical devices is solved, enabling low-power, high-efficiency neural network computing.

CN122334374APending Publication Date: 2026-07-03HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-03-23
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Once traditional optical devices are fabricated, their transfer functions are fixed, making it difficult to adapt to the dynamic adjustment requirements of different network layers for activation function shape or normalization parameters, resulting in low computational efficiency and high energy consumption.

Method used

An optical computing system consisting of a pump laser, an optical path time delay unit, a digital micromirror chip, a probe laser, a dichroic mirror, a two-dimensional material array, and a photodiode is used to fuse convolution, batch normalization, and nonlinear activation functions through femtosecond laser modulation. The convolution-batch normalization-activation operation is completed on a single hardware platform by utilizing the optical properties of two-dimensional materials.

Benefits of technology

It achieves low-power, high-efficiency neural network computing, reduces data transfer overhead by 60%, and has a computing speed 3-4 orders of magnitude faster than electronic GPUs, while reducing power consumption by more than 50%, adapting to the dynamic adjustment needs of different network layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122334374A_ABST
    Figure CN122334374A_ABST
Patent Text Reader

Abstract

This application belongs to the field of optoelectronic computing and artificial intelligence hardware acceleration technology, specifically disclosing an optical computing system and a method for implementing optical neural network operations. A digital micromirror chip is used to disperse a first femtosecond laser into beams of equal intensity but different angles. A dichroic mirror is used to combine a third femtosecond laser with the split beams. A two-dimensional material array performs operations on the CONV-BN layer of a convolutional neural network. A photodiode is used to collect the change in total transmitted light intensity after passing through the two-dimensional material array and perform photoelectric nonlinear conversion to obtain the relative transmittance change, realizing the nonlinear ReLU activation function. This application reduces inter-operator data transfer overhead by 60%, significantly reduces latency caused by the memory wall effect, and uses a DMD chip to encode the input feature map into an optical signal array. A single frame of input can process the entire feature map in parallel, achieving high parallelism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of optoelectronic computing and artificial intelligence hardware acceleration technology, and more specifically, relates to an optical computing system and a method for implementing optical neural network operations. Background Technology

[0002] Convolutional Neural Networks (CNNs) are a core technology in computer vision, but their deployment on traditional electronic hardware (such as GPUs and TPUs) faces severe challenges: redundant computational steps and data transport bottlenecks. Traditional CNN inference requires the sequential execution of three independent steps: "convolution operation - nonlinear activation - batch normalization," which are repeated multiple times. After each step, the intermediate results must be written to memory and read out before the next step, resulting in a large number of data transport operations. Data transport energy consumption accounts for 60%-70% of the total energy consumption of CNNs, severely limiting energy efficiency.

[0003] Optical computing, leveraging the ultra-wide bandwidth, low loss, and parallel propagation characteristics of optical transmission, can equivalently map large-scale matrix operations in neural network training / inference to optical linear transmission and interferometry, potentially breaking through the bottlenecks of traditional electronic architectures in terms of energy efficiency and computational efficiency. However, existing optical computing methods typically utilize Mach-Zehnder interferometer (MZI) arrays to achieve linear convolution, while the activation function still relies on backend electronic circuits or independent nonlinear optical modules. This discrete optical-electrical-optical or linear-nonlinear architecture increases system latency and hardware complexity, failing to fundamentally reduce communication overhead between operators. Furthermore, once traditional optical devices are fabricated, their transfer functions are fixed, making it difficult to adapt to the dynamic adjustment requirements of activation function shapes or normalization parameters for different network layers. Therefore, there is an urgent need for an operator fusion technology that can simultaneously achieve linear weighting, nonlinear mapping, and parameter normalization at the microscopic scale using a single physical mechanism, to completely eliminate intermediate data transfer and significantly reduce power consumption. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this application is to provide an optical computing system and a method for implementing optical neural network operations, aiming to solve the problem that once traditional optical devices are fabricated, their transfer functions are fixed and it is difficult to adapt to the dynamic adjustment requirements of different network layers for activation function shapes or normalization parameters.

[0005] The first aspect of this application relates to an optical computing system for implementing the CONV-BN layer function in a convolutional neural network; comprising: a pump laser, an optical path time delay unit and a digital micromirror chip arranged in sequence, a probe laser, a dichroic mirror, a two-dimensional material array and a photodiode arranged in sequence along the light output direction of the dichroic mirror; The pump laser and probe laser are used to emit the first femtosecond laser and the third femtosecond laser, respectively, and their power levels serve as the first codes of the input signals. Second encoding The optical time delay device is used to create a time interval between the excitation of the first femtosecond laser and the third femtosecond laser. The digital micromirror chip is used to disperse the first femtosecond laser beam into... Beam splitting laser, of which 1~ Each laser beam is used to encode a single pixel in the input image signal, and each pixel encoding corresponds to an input window in the convolution operation. ; × The size of the input image signal; A dichroic mirror is used to combine a third femtosecond laser beam with a split laser beam; the two-dimensional material array contains... One computing unit, used to perform Operations; among which, convolution kernel operation ; For input window The constructed matrix; The quantity is known; A photodiode is used to collect the change in total transmitted light intensity after the split laser and the third femtosecond laser pass through a two-dimensional material array. The photoelectric nonlinear conversion is performed to obtain the relative transmittance change, thereby realizing the function of the nonlinear ReLU activation function.

[0006] In some embodiments, the shape of the operator fusion function is dynamically reconstructed by adjusting the pulse power of the first and third femtosecond lasers, the time interval between their excitation, and the encoding of the digital micromirror chip.

[0007] In some implementations, the optical computing system also includes a frequency multiplier disposed between the pump laser and the optical path time delay unit to reduce the wavelength of the first femtosecond laser by half to serve as the second femtosecond laser.

[0008] In some implementations, the wavelength of the first femtosecond laser is shorter than the wavelength of the third femtosecond laser; the wavelengths of the first and third femtosecond lasers are selected based on the band gap of the two-dimensional material array, according to the following formula; ; in, It is the laser wavelength. It is the band gap of a two-dimensional material.

[0009] In some embodiments, the band gap of the two-dimensional material is controlled between 1 eV and 6 eV, and the two-dimensional material includes chalcogenides. ReS and SnS, selenides , , , SnSe and Phosphates BP, SiP, At least one of CrOCl and Te. In some embodiments, the wavelength of the first femtosecond laser is between 400 nm and 800 nm, and the wavelength of the third femtosecond laser is between 800 nm and 150 nm.

[0010] In some implementations, the change in relative transmittance in the photodiode The relationship with the nonlinear ReLU activation function is as follows: = K Relu( ); ; in, This represents the lowest photosensitivity threshold of a photodiode. Photoelectric conversion responsivity; K This is a correction factor.

[0011] The second aspect of this application relates to a method for implementing optical neural network operations based on an optical computing system, specifically including the following steps: The excitation time interval between the first femtosecond laser and the third femtosecond laser is taken as... It emits the first femtosecond laser and the second femtosecond laser; Disperse the first femtosecond laser into Beam splitting laser, of which 1~ Each laser beam is used to encode a single pixel in the input image signal, and each pixel encoding corresponds to an input window in the convolution operation. ; × The size of the input image signal; The third femtosecond laser and the split laser beam are combined and input into a two-dimensional material array to perform... Operations; Two-dimensional material arrays include Each computational unit, convolution kernel operation ; For input window The constructed matrix; The quantity is known; The change in total transmitted light intensity after the split-beam laser and the third femtosecond laser pass through the two-dimensional material array is collected. The photoelectric nonlinear conversion is performed to obtain the relative transmittance change, thereby realizing the function of the nonlinear ReLU activation function.

[0012] In some implementations, the shape of the operator fusion function is dynamically reconstructed by adjusting the pulse power of the first and third femtosecond lasers, the time interval between their excitation, and the encoding of the digital micromirror chip.

[0013] In some implementations, the relative transmittance change The relationship with the nonlinear ReLU activation function is as follows: = K Relu( ); ; in, This represents the lowest photosensitivity threshold of a photodiode. Photoelectric conversion responsivity; K This is a correction factor.

[0014] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art: This application fuses the traditional serial three-step convolution-activation-normalization calculation into a single physical process of light propagation in two-dimensional materials. Intermediate results do not need to be written to memory, reducing inter-operator data transfer overhead by 60% and significantly mitigating latency caused by the memory wall effect. Utilizing the femtosecond-level carrier relaxation time of two-dimensional materials, the time required for a single operator fusion calculation is only in the picosecond to femtosecond range, theoretically 3-4 orders of magnitude faster than electronic GPUs. By using a DMD chip to encode the input feature map into an optical signal array, the entire feature map can be processed in parallel for a single frame of input, achieving high parallelism.

[0015] By eliminating high-frequency data read / write operations and additional electro-optic / photoelectric conversion circuits, and by using femtosecond lasers to control two-dimensional materials with extremely low energy consumption (on the order of femtojoules), this application reduces power consumption by more than 50% under the same computing power compared to traditional electron accelerators. Compared to traditional discrete optical computing methods, the power consumption reduction effect is significant.

[0016] Based on two-dimensional materials with atomic-level thickness, the computational unit area is related to the beam spot size. The unit computational area is small and easily compatible with existing silicon-based photonic integrated circuits, providing a new approach to realizing on-chip ultra-compact AI accelerators. Attached Figure Description

[0017] Figure 1 This is an optical computing system architecture diagram provided in the embodiments of this application.

[0018] Figure 2 This is a diagram showing the effect of Pump laser power, Probe laser power, and relaxation time on the relative transmittance of the material when Pump laser power is fixed, according to an embodiment of this application.

[0019] Figure 3 This is a diagram showing the effect of fixed Probe laser power, Pump laser power, and relaxation time on the relative transmittance of the material, provided in an embodiment of this application.

[0020] Figure 4 This is a weight relationship diagram provided in the embodiments of this application.

[0021] Figure 5 This is the operator fusion mapping diagram provided in the embodiments of this application.

[0022] Figure 6 This is the neural network calculation result provided in the embodiments of this application.

[0023] In all the accompanying drawings, the same reference numerals are used to denote the same elements or structures, wherein: 10 is the first femtosecond laser; 11 is the frequency doubler; 12 is the optical path time delayer; 13 is the digital micromirror chip; 20 is the third femtosecond laser; 14 is the beam splitter laser; 30 is the dichroic mirror; 31 is the beam combiner; 32 is the two-dimensional material array; 40 is the photodiode. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0025] In this application, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A existing alone, A and B existing simultaneously, and B existing alone. In this application, the symbol " / " indicates that the related objects are in an "or" relationship, for example, A / B means A or B.

[0026] In this application, the terms “first” and “second” are used to distinguish different objects, rather than to describe a specific order of objects.

[0027] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0028] CONV-BN (Convolution-Batch Normalization) layers are a commonly used and effective combination of techniques in deep neural networks. Their core function is to extract effective feature information from images through convolution operations, and then use batch normalization to stabilize the input data distribution of the intermediate layers, thereby significantly accelerating the training process and improving the stability of deep neural networks. However, existing CNN inference processes require the sequential execution of three independent steps: "convolution operation - nonlinear activation - batch normalization," which are repeated multiple times. After each step, the intermediate results must be written to memory and read out again before the next step begins. This frequent data transfer operation leads to significant energy consumption, accounting for 60% to 70% of the total energy consumption of CNNs, severely limiting the overall energy efficiency. To address this bottleneck, this application proposes using optical computing to replace traditional electronic computing. It fully leverages the ultra-wide bandwidth, low loss, and parallel propagation characteristics of optical transmission, mapping large-scale matrix operations in neural network training and inference to equivalent optical linear transmission and interference processes, thereby achieving low-power, high-speed neural network computing at the hardware level.

[0029] like Figure 1 As shown, the optical computing system provided in this application includes key components such as a first optical path subsystem, a second optical path subsystem, a dichroic mirror, a two-dimensional material array, and a photodiode. The first optical path subsystem, arranged sequentially along the optical path propagation direction, includes a pump laser, a frequency doubler, an optical path time delay unit, and a digital micromirror device (DMD). During system operation, the linear polarization angles of both the first and second femtosecond lasers are fixed to preset angles. Through precise optical path control and the linear optical response of the two-dimensional material array, the three core operations—CONV convolution, BN batch normalization, and ReLU activation—are directly performed on a single hardware platform. This avoids the repetitive data transfer and distributed computation inherent in traditional architectures, achieving true in-situ computation within the optical domain.

[0030] In practical applications, firstly, a suitable dataset is selected based on task requirements, and an appropriate CNN (Convolutional Neural Network) architecture is chosen for network training. Then, layer fusion optimization is performed to adapt to the characteristics of optical computing hardware. Next, the weight parameters obtained from training are checked, and necessary compensation and calibration are performed according to the optical parameters of the specific hardware. Subsequently, the core computing unit array (i.e., a two-dimensional material array) of the large-scale optical computing system is prepared. Finally, the input signal and the compensated weight data are synchronously loaded onto the computing unit array through a DMD chip, and large-scale matrix operations are completed using the parallel propagation characteristics of light to achieve high-speed, low-power optical neural network inference.

[0031] First, after obtaining the parameters of the trained neural network, extract the parts that conform to convolution, batch normalization, and ReLU activation. The convolutional layer parameters include: Convolution kernel: Convolution bias ; BN layer parameters include: linear weights Linear bias running average Operating variance ; Operator fusion uses mathematical transformations to merge the parameters of the BN layer into the preceding convolutional layer. Since the weights do not need to be changed during inference, the BN layer is equivalent to a fixed linear transformation. Substituting this formula into the output formula of the convolutional layer yields a set of linear convolutional fusion weights and fusion biases. The final result is that the original convolutional layer and BN layer are mathematically equivalent to a completely new convolutional layer.

[0032] Convolutional fusion weights: ; Fusion bias: ; in, These are the convolutional fusion weights; For fusion bias; Subsequently, this application provides an optical computing system, including a first optical path subsystem, a second optical path subsystem, a dichroic mirror, a two-dimensional material array, a relative transmittance detector, and a photodiode; The first optical path subsystem includes a Pump laser, a frequency multiplier 11, an optical path time delay unit 12, and a digital micromirror chip 13 arranged in sequence. The pump laser emits a first femtosecond laser beam with a wavelength of α nm. A frequency multiplier converts the first femtosecond laser beam into a second femtosecond laser beam with a wavelength reduced to 0.5 α nm. This second femtosecond laser beam is used to excite the two-dimensional material array into a non-equilibrium excited state, serving as the signal input. The power of the second femtosecond laser beam is used as the first encoding of the input signal. The optical path time delay unit 12 is used for linear scaling of the convolution fusion weights; the optical path time delay unit 12 is used to realize the excitation time interval between the second femtosecond laser and the third femtosecond laser 20. The second optical subsystem includes a probe laser for generating a third femtosecond laser beam 20 with a wavelength of bnm, which directly illuminates a dichroic mirror 30; the power of the third femtosecond laser serves as the second encoding of the input signal. This is used for linear adjustment of the fusion bias; Within the chip, a digital micromirror device (DMD) chip 13 is used to achieve parallel input of input image information. The DMD disperses the second femtosecond laser into... 14 beams of laser light with the same beam intensity but different angles are used by the DMD to... A laser beam is aimed at a two-dimensional material array to encode the input image signal. x Corresponding input window Where r is the laser beam required to encode one input pixel. × The size of the input image signal; m n It refers to the size of the convolution kernel, generally speaking. m = n .

[0033] Dichroic mirror 30 is used to connect the third femtosecond laser with... The beam splitter laser 31 is combined and enters the computing unit through a microlens; it is then transferred to a thin layer of transparent quartz glass by mechanical peeling and etched to form a... A two-dimensional material array of computing units; the two-dimensional material array contains Each computational unit processes one pixel from the input image signal. The two-dimensional material array 32 can be viewed as a combination of several computational units, where the computational unit is a unit linear calculator; The purpose of a single-unit linear calculator is to calculate ,in For input, For output; due to the characteristics of the hardware, the relationship between its input and output can be described by the following model: ; in, , Through measurement and calibration of the actual computing unit, X is determined to be the input signal, which is the encoding of the image input signal by the microlens array; the current computational encoding in the system... The linear scaling factor is , Encoded as Encoded as The output is , where is the light intensity of the pump laser and probe laser after passing through the two-dimensional material array; Next, we will define the two-dimensional convolution operation: Dynamic window: Convolution kernel (its elements are) In a larger input matrix Swipe up.

[0034] Local window: For each position (i, j) in the output matrix, there is a submatrix (or window) in the input matrix X. . From China and Israel Starting from a point, the submatrix obtained by sampling with a step size s: The dimensions are, It is the data in the (i,j)th convolution window. Row and column indices inside the submatrix window ; It should be noted here that the size of the convolution kernel is assumed to be... m × n (i.e., the size of the window), then the window It's just one m OK n A matrix of columns; ; in, The original input matrix, For the input window matrix; i and j Let be the coordinates of the window's starting point, and let be the coordinates of the window in the original matrix. The starting row and column in is the step size, and is the sampling interval; Core operation: The result of convolution is the transformation of the convolution kernel... Its corresponding input window Perform element-wise multiplication, then sum all the products, and finally add a bias. ; ; A convolution operation Σ( ) + Includes m n This involves performing a large summation of multiple multiplications. Since an optical unit can only perform this task once, this large summation task is distributed to multiple parallel optical computing unit arrays to complete simultaneously.

[0035] Therefore, each time the system performs... Array computation.

[0036] Instead of directly calculating the sum, it assigns each computational unit in the array to calculate one term in the convolution. .

[0037] At the same time, it will increase the overall bias. Distributed equally m n The bias assigned to each multiplication is... / ( m n ).

[0038] Thus, when all m n After all the multiplication operations are completed, the sum of their outputs is exactly equal to the original convolution result: Photodiode 40 is used to collect the change in total transmitted light intensity after the split laser and the third femtosecond laser pass through the two-dimensional material array. The photoelectric nonlinear conversion is performed to obtain the relative transmittance change, thereby realizing the function of the nonlinear ReLU activation function.

[0039] Studies have shown that the change in relative transmittance is highly linearly related to power, and the operation of the BN layer in a neural network can be realized based on a reconfigurable optical computing unit. Furthermore, the nonlinear photoelectric conversion property of a photodiode can be used to realize the function of a nonlinear ReLU activation function.

[0040] = K Relu( ); ; in, K This is a correction factor.

[0041] Therefore, the optical computing system provided in this application ultimately achieves... The mapping relationship can effectively improve the efficiency of computational neuromorphism.

[0042] In some specific implementations, the shape of the operator fusion function is dynamically reconstructed by adjusting the pulse power and delay time of the two femtosecond laser beams in real time, as well as the encoding of the microlens array of the DMD chip. This adapts to the requirements of convolutional neural network layers of different scales, enabling parallel processing of large-scale operator fusion operations on a two-dimensional material array. The adjustment of the femtosecond laser pulse power is used to regulate... The accuracy; In some implementations, the wavelength of the first femtosecond laser is shorter than that of the third femtosecond laser, and the wavelength is selected based on the band gap of the two-dimensional material array. In some specific implementations, a suitable two-dimensional material is selected as the computing medium for femtosecond light. Optionally, the band gap range of the two-dimensional material is controlled between 1 eV and 6 eV, and the two-dimensional material includes chalcogenides. ReS and SnS, selenides , , , SnSe and Phosphates BP, SiP, At least one of CrOCl and Te.

[0043] Furthermore, the wavelength of the first femtosecond laser is between 400nm and 800nm, and the wavelength of the third femtosecond laser is between 800nm ​​and 1500nm. Example A method for implementing optical neural network operations based on an optical computing system specifically includes the following steps: Step 1: Select ResNet21 to train on the CIFAR10 dataset and complete layer fusion, check the layer fusion results and provide the layer fusion convolution function curve; Here, the images from the CIFAR10 dataset are compressed to 32. 32 pixel units; Step 2: Fit the layer fusion convolution function corresponding to the calculation results to obtain the implementation of layer fusion on the optical computing system. function; Step 3: Input signals are fed into a two-dimensional material array via a DMD chip and large-scale parallel computation is performed.

[0044] The computational scenario selected in this application requires the two-dimensional material array to have a good laser threshold and be easily fabricated on a large scale. Specifically, the two-dimensional CrOCl material exhibits a significant ultrafast optical response in the visible to near-infrared band; it has a wide bandgap, stable physical structure, can be stored for extended periods in high-temperature environments, and possesses a high laser damage threshold; furthermore, the interlayer forces are predominantly van der Waals forces, with extremely low binding energy (approximately 0.3~0.4 e). V / This makes the interlayer bonding force in the Z-axis direction (perpendicular to the layer plane) much weaker than the covalent / metallic bonds in the plane (XY plane), making it easier to peel off large-sized thin single-crystal materials, which is beneficial for the preparation of large-scale material arrays.

[0045] Step three specifically includes the following steps: Step 3.1: Fabrication of two-dimensional material arrays Single-layer or few-layer CrOCl nanosheets with a thickness of 2 nm to 12 nm were exfoliated from bulk two-dimensional material crystals using a mechanical exfoliation method, with length and width dimensions controlled to approximately [missing information]. ; The exfoliated CrOCl nanosheets were then precisely transferred to a transparent quartz glass substrate with a thickness of 0.08 mm to 0.2 mm using a dry transfer technique to serve as computing array units. The units were then laser-etched to a side length of 1 mm. A square, etched in total 32 A two-dimensional CrOCl material array is composed of 32 = 1024 computational units; Step S3.2: Construction of the optical computing system and specific mapping and encoding process A first optical path subsystem is constructed, and a first femtosecond laser with a wavelength of 800nm ​​is set. The wavelength of the first femtosecond laser is reduced to 400nm by a frequency doubler 11, serving as the pump laser to excite the two-dimensional material array into a non-equilibrium excited state. The signal input power is also specified. Encoding the input signal; detecting changes in material transmittance and reading the calculation results; The excitation time interval between the pump laser (a second femtosecond laser with a wavelength of 400 nm) and the probe laser (a third femtosecond laser with a wavelength of 1100 nm) is achieved by using an optical time delay device. =20fs, sampling range is -800 fs ~ 800fs, corresponding weight adjustment.

[0046] Subsequently, a digital micromirror device (DMD) chip, a precision optical semiconductor switch composed of millions of miniature movable mirrors, disperses the second femtosecond laser beam into 16,384 beams of varying intensities; the DMD chip targets 32... The input image of size 32 pixels is encoded. Considering that DMD can only perform 01 encoding, one pixel module is encoded by every 16 microlenses, which is extended to a 4-bit weighted encoding; that is, the beam focused by every 16 microlenses is focused on one CrOCl computing unit. A second optical path subsystem was constructed, in which a third femtosecond laser beam with a wavelength of 1100nm was set up to directly illuminate a dichroic mirror. The power level was used as... ; The dichroic mirror focuses 16,385 beams of light onto a single path, which then passes through an array of CrOCl material on a transparent quartz glass. The laser light passing through the two-dimensional material array is collected by a photodiode. By fixing the pump laser power at 20mW and varying the probe laser power (0-20mW) and relaxation time, the test data obtained are as follows: Figure 2 As shown; By fixing the probe laser power at 20mW and varying the pump laser power from 0 to 20mW and the relaxation time, the test data obtained are as follows: Figure 3 As shown; The linear component is extracted as the updated conductance state for the weights. The weights can be effectively distinguished into 8 weight states, such as... Figure 4 As shown; where, It is △τ and The function, .

[0047] The specific model mapping relationship is as follows: ; ;in, Let be the delay time between the two beams; here, the pump optical power is used to represent the linear scaling factor of the input signal. It represents the weights of the convolution kernel and is used to encode the convolution kernel weights; For probe optical power; ; In the actual execution of the algorithm, the convolution kernel used is... The size is 3 3, The values ​​of i and j range from 1 to 32.

[0048] It is a constant, and its value is found by addressing in the three-dimensional mapping relationship obtained through measurement after the device is manufactured; The actual calculations performed in the hardware system are as follows: First, in the preparatory stage, the computer calculates all the weights and converts the standard weights into hardware weights through addressing conversion encoding. Retrieve the fixed value of the current layer and Then take out each round ; Each calculation will calculate the current calculation Encoded as In advance After encoding, compensation and correction are performed to make... It falls directly within the available range, requiring no further scaling; the output is already set. ; Finally, when calculating the activation function for the network output, the properties of a photodiode (PD) are utilized. During convolution, the PD remains in a photosensitive state, and the calculation results are automatically accumulated in the PD's well. If the result exceeds the well's capacity, it is considered saturated and has no impact on inference (similar to how PyTorch handles overflow). ReLU is achieved through zero-point offset and the PD's minimum light cutoff. To achieve positive and negative multiplication, the zero point is offset by setting the PD's minimum photosensitive value to achieve the reverse offset of the zero point. During this process, all values ​​below the minimum photosensitive value are recorded as 0, thus achieving ReLU activation.

[0049] = K Relu( ); ; in, This represents the lowest photosensitivity threshold of a photodiode. Photoelectric conversion responsivity; K This is a correction factor.

[0050] Considering the sufficiently high linearity, the pixel weights of the 16-bit input signal were obtained through interpolation. Simulation mapping yielded the operator fusion relationship, such as... Figure 5 As shown; The emitted light signal is received by a photodiode and converted into an electrical signal, which is then directly sent to the next-level network or output as a classification result. Since there is no intermediate storage, the system latency mainly depends on the light flight time and the detector signal acquisition response time (<1). ).

[0051] Performance verification: The core operator fusion component was deployed on a spatial light system. After 350 epochs of training and hardware-in-the-loop testing, the system achieved a recognition accuracy of 92.4% on the CIFAR-10 dataset. Figure 6 As shown, the total power consumption of the optical path core is on the order of hundreds of fJ, and the energy consumption of a single operator fusion calculation multiplication is as low as aJ. In comparison of calculation steps, traditional GPUs need to perform 3 kernel startups and 3 memory read / write operations; this application only requires 1 optical transmission process, reducing the number of calculation steps by 66.7%.

[0052] This application achieves, for the first time, all-optical fusion of CNN core operators at the physical level through deep coupling of femtosecond lasers and two-dimensional materials. This fundamentally reduces the problems of redundant computation and high power consumption caused by the discreteness of operators in traditional computing architectures, and provides a powerful solution for the next generation of ultra-low power and ultra-high speed AI hardware.

[0053] In summary, this application has the following advantages compared with the prior art: This application fuses the traditional serial three-step convolution-activation-normalization calculation into a single physical process of light propagation in two-dimensional materials. Intermediate results do not need to be written to memory, reducing inter-operator data transfer overhead by 60% and significantly mitigating latency caused by the memory wall effect. Utilizing the femtosecond-level carrier relaxation time of two-dimensional materials, the time required for a single operator fusion calculation is only in the picosecond to femtosecond range, theoretically 3-4 orders of magnitude faster than electronic GPUs. By using a DMD chip to encode the input feature map into an optical signal array, the entire feature map can be processed in parallel for a single frame of input, achieving high parallelism.

[0054] By eliminating high-frequency data read / write operations and additional electro-optic / photoelectric conversion circuits, and by using femtosecond lasers to control two-dimensional materials with extremely low energy consumption (on the order of femtojoules), this application reduces power consumption by more than 50% under the same computing power compared to traditional electron accelerators. Compared to traditional discrete optical computing methods, the power consumption reduction effect is significant.

[0055] Based on two-dimensional materials with atomic-level thickness, the computational unit area is related to the beam spot size. The unit computational area is small and easily compatible with existing silicon-based photonic integrated circuits, providing a new approach to realizing on-chip ultra-compact AI accelerators.

[0056] It should be understood that expressions such as “comprising” and “may include” used in this application indicate the existence of the disclosed functions, operations, or constituent elements, and do not limit one or more additional functions, operations, and constituent elements. In this application, terms such as “comprising” and / or “having” are to be interpreted as indicating a particular characteristic, number, operation, constituent element, component, or combination thereof, but not to exclude the existence or possibility of adding one or more other characteristics, numbers, operations, constituent elements, components, or combinations thereof.

[0057] Furthermore, in this application, the expression "and / or" includes any and all combinations of the associated listed words. For example, the expression "A and / or B" may include A, may include B, or may include both A and B.

[0058] In the description of the embodiments of this application, it should be noted that, unless otherwise explicitly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a detachable connection or a non-detachable connection; it can be a direct connection or an indirect connection through an intermediate medium.

[0059] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An optical computing system, characterized in that, Used to implement the CONV-BN layer function in a convolutional neural network; including: a pump laser, an optical path time delay unit and a digital micromirror chip arranged in sequence, a probe laser, a dichroic mirror, a two-dimensional material array and a photodiode arranged in sequence along the light output direction of the dichroic mirror; The pump laser and probe laser are used to emit the first femtosecond laser and the third femtosecond laser, respectively, and their power levels serve as the first codes of the input signals. Second encoding The optical time delay device is used to create a time interval between the excitation of the first femtosecond laser and the third femtosecond laser. The digital micromirror chip is used to disperse the first femtosecond laser beam into... Beam splitting laser, of which 1~ Each laser beam is used to encode a single pixel in the input image signal, and each pixel encoding corresponds to an input window in the convolution operation. ; × The size of the input image signal; A dichroic mirror is used to combine a third femtosecond laser beam with a split laser beam; the two-dimensional material array contains... One computing unit is used to perform Operations; among which, convolution kernel operation ; For input window The constructed matrix; The quantity is known; A photodiode is used to collect the change in total transmitted light intensity after the split laser and the third femtosecond laser pass through a two-dimensional material array. The photoelectric nonlinear conversion is performed to obtain the relative transmittance change, thereby realizing the function of the nonlinear ReLU activation function.

2. The optical computing system according to claim 1, characterized in that, The shape of the operator fusion function is dynamically reconstructed by adjusting the pulse power of the first and third femtosecond lasers, the time interval between their excitation, and the encoding of the digital micromirror chip.

3. The optical computing system according to claim 1, characterized in that, Relative transmittance change in photodiode The relationship with the nonlinear ReLU activation function is as follows: = K Re( ); ; in, This represents the lowest photosensitivity threshold of the photodiode. Photoelectric conversion responsivity; K This is a correction factor.

4. The optical computing system according to claim 1, characterized in that, It also includes a frequency multiplier, which is placed between the pump laser and the optical path time delay unit to reduce the wavelength of the first femtosecond laser to half, so as to serve as the second femtosecond laser.

5. The optical computing system according to any one of claims 1 to 3, characterized in that, The wavelength of the first femtosecond laser is shorter than that of the third femtosecond laser; the wavelengths of the first and third femtosecond lasers are selected based on the band gap of the two-dimensional material array.

6. The optical computing system according to claim 5, characterized in that, The band gap range of two-dimensional materials is controlled between 1 eV and 6 eV. Two-dimensional materials include chalcogenides. ReS and SnS, selenides , , , SnSe and Phosphates BP, SiP, At least one of CrOCl and Te.

7. The optical computing system according to claim 6, characterized in that, The wavelength of the first femtosecond laser is between 400nm and 800nm, while the wavelength of the third femtosecond laser is between 800nm ​​and 150nm.

8. A method for implementing optical neural network operations based on the optical computing system of claim 1, characterized in that, Includes the following steps: The excitation time interval between the first femtosecond laser and the third femtosecond laser is taken as It emits the first femtosecond laser and the second femtosecond laser; Disperse the first femtosecond laser into Beam splitting laser, of which 1~ Each laser beam is used to encode a single pixel in the input image signal, and each pixel encoding corresponds to an input window in the convolution operation. ; × The size of the input image signal; The third femtosecond laser and the split laser beam are combined and input into a two-dimensional material array to perform... Operations; Two-dimensional material arrays include Each computational unit, convolution kernel operation ; For input window The constructed matrix; The quantity is known; The change in total transmitted light intensity after the split-beam laser and the third femtosecond laser pass through the two-dimensional material array is collected. The photoelectric nonlinear conversion is performed to obtain the relative transmittance change, thereby realizing the function of the nonlinear ReLU activation function.

9. The method for implementing optical neural network operations according to claim 8, characterized in that, The shape of the operator fusion function is dynamically reconstructed by adjusting the pulse power of the first and third femtosecond lasers, the time interval between their excitation, and the encoding of the digital micromirror chip.

10. The method for implementing optical neural network operations according to claim 8, characterized in that, Relative transmittance change The relationship with the nonlinear ReLU activation function is as follows: = K Re( ); ; in, This represents the lowest photosensitivity threshold of a photodiode. Photoelectric conversion responsivity; K This is a correction factor.