An optical convolutional neural network structure based on thin film lithium niobate and silicon hybrid
Patent Information
- Application Number
- CN202311424796.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2043-10-31
AI Technical Summary
[0004]针对现有技术的以上缺陷或改进需求,本发明的目的在于提供一种基于薄膜铌酸锂和硅混合的光学卷积神经网络结构,旨在解决现有的集成光学卷积架构中存在的计算冗余,空间资源耗费大,运算速率低的问题
1、本发明提供的一种基于薄膜铌酸锂和硅混合的光学卷积神经网络结构采用高速频率流的形式进行卷积计算,相比传统方案采用将卷积计算转化为矩阵乘法的方式带来的冗余,本方案充分利用现有波长资源,并且避免了在电子计算机上完成复杂的傅里叶变换。
Smart Images

Figure CN117436486B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of optical computing technology, and more specifically, relates to an optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon. Background Technology
[0002] With the advent of the era of big data and artificial intelligence, the demand for computing power and energy consumption is growing exponentially. However, traditional integrated circuit design is approaching the limits of Moore's Law; the size and power consumption of electronic transistors can no longer be reduced in the traditional way. This means that the ever-increasing demand for computing power cannot be met, and the heat generated by electronic transistors is becoming increasingly difficult to manage. To address this challenge, new computing architectures are emerging. Photonic devices have attracted much attention due to their excellent performance characteristics, including high bandwidth, low crosstalk, and low power consumption. These characteristics enable photonic devices to be widely used in communication and computing fields. Convolutional computing plays a crucial role in big data processing and artificial intelligence algorithms, requiring enormous computing resources. Optical devices are widely used in convolutional computing applications, such as optical convolutional neural networks and image processing, due to their multi-dimensional resources and parallel transmission capabilities. Integrated optical convolutional computing will become an important acceleration engine for future CPUs.
[0003] Currently, optical convolutional neural networks (CNNs) are mainly classified into four categories based on their convolution implementation principles: matrix multiplication, delay accumulation, spatial Fourier transform, and frequency domain convolution. The core idea of the first and second categories is to utilize various integrated photonic devices to perform dot multiplication. Both of these schemes require continuous repetition and shifting of input information, and this redundant computation significantly increases memory usage and access. Furthermore, as the scale of convolution computation increases, the number of devices required for these schemes also increases, which is detrimental to chip compactness. The core idea of the third and fourth categories is to utilize Fourier transform. However, the third category, which uses spatial information to load the input vector, faces severe scalability problems when the input scale increases. The fourth category currently mainly faces problems such as the need for complex electrical Fourier transform operations for input data loading and low bandwidth. Summary of the Invention
[0004] In view of the above-mentioned defects or improvement needs of the existing technology, the purpose of this invention is to provide an optical convolutional neural network structure based on thin-film lithium niobate and silicon hybrid, which aims to solve the problems of computational redundancy, large space resource consumption and low computing speed in the existing integrated optical convolutional architecture.
[0005] To achieve the above objectives, the present invention provides an optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon. The optical convolutional neural network structure includes an input information loading layer, a convolutional layer, a pooling layer, a nonlinear layer, and a fully connected layer.
[0006] The input information loading layer comprises two parts: an optical frequency comb source and an optical frequency comb spectral shaping layer. The optical frequency comb source can be generated by a distributed feedback (DFB) laser producing a single-wavelength laser, which is then modulated by a thin-film lithium niobate microring modulator driven by a single-frequency microwave. Alternatively, it can be generated by a combination of a phase modulator (PM) and an intensity modulator (IM) of thin-film lithium niobate. It can also be generated by microrings of various other materials (such as...). , Alternatively, multiple wavelength light sources with equal frequency intervals can be used. Optical frequency comb shaping transforms the intensity of the generated optical frequency comb teeth into a neural network input vector. This can be accomplished by applying a voltage to a thin-film lithium niobate microring resonator array via a field-programmable gate array (FPGA) module through a digital-to-analog converter (DAC), where the intensity of each microring corresponds one-to-one with that of each comb tooth. Alternatively, it can be accomplished by other spectral shaping devices, such as wavelength division multiplexing (WDM) + IM + WDM.
[0007] The convolutional layer is constructed using a single thin-film lithium niobate phase modulator, driven by a high-speed radio frequency (RF) signal generated by an AWG (autogenous ring gear). The aforementioned optical frequency comb loaded with the input vector enters the thin-film lithium niobate phase modulator driven by the AWG signal for convolution calculation. The repetition frequency of this AWG signal must be the same as the repetition frequency of the signal driving the thin-film lithium niobate microring modulator, i.e., the same frequency interval as the input vector. To generate the desired target convolutional kernel, the AWG needs to be controlled by an FPGA.
[0008] The pooling layer is implemented using a flat-top filter with a certain bandwidth. This filter can be implemented using multiple micro-ring arrays. Each micro-ring can be a single ring, or multiple rings connected in series or parallel to form a flat-top filter with a certain bandwidth. This allows each micro-ring to sequentially filter out the wavelengths of multiple adjacent convolution results, achieving average pooling. The pooling layer can also be implemented using other flat-top filters with a certain bandwidth.
[0009] The nonlinear layer is achieved by an active silicon-based microring driven by a photodetector (PD). The PD is an on-chip silicon-germanium photodetector, and the active microring is a PN-doped microring modulator. Each output wavelength of the pooling layer enters a different PD, converting light intensity information into photocurrent. The photocurrent drives multiple active microrings to generate nonlinearity. The active microrings are cascaded on a waveguide, and their input light source is the same as that of the input information loading layer, including initial optical frequency comb generation and optical frequency comb shaping, so that the multiple wavelengths entering the active microring array have the same intensity. Each active microring loads a nonlinear response onto its corresponding wavelength.
[0010] The fully connected layer consists of a silicon-based microring array. The pooling result (N wavelengths) output from the pooling layer is divided into M parts using an on-chip beam splitter, and each part is fed into a weight library consisting of N cascaded microrings. Each microring controls the weight corresponding to a wavelength. The output intensity is summed using a PD to achieve M×N matrix vector multiplication. Alternatively, other incoherent wavelength intensity superposition schemes can be used to implement the fully connected layer, such as MZI or phase change material schemes.
[0011] Furthermore, this convolutional neural network utilizes online training to achieve network convergence. The network's output port is connected to an FPGA via an ADC, which controls the convolution vectors generated by the AWG (Automatic Gaussian Network), and also connects to a micro-ring array via a DAC, thereby controlling the input information loading and the matrix of the fully connected layers, thus forming a feedback system. A digital-to-analog converter (DAC) chip is used to output analog voltage and control the thermal phase shifter and modulator. An analog-to-digital converter (ADC) chip is used to convert the analog electrical signals detected by the photodetector into digital electrical signals, facilitating subsequent processing by the FPGA. Online training is performed using the gradient descent algorithm to adjust various parameters of the network.
[0012] Furthermore, this network can perform convolution calculations in the real number domain. The input signal is split into two paths, and each path is convolved with a thin-film lithium niobate PM. The results are then pooled and differentially analyzed by a balanced detector to complete the convolution calculation in the real number domain.
[0013] Furthermore, the convolutional layers, pooling layers, nonlinear layers, and fully connected layers of this network can be arbitrarily arranged and combined or arbitrarily cascaded and expanded. Due to the relay supplementation of the light source, this network can realize the expansion of the number of layers of deep convolutional neural networks, thereby realizing larger-scale and more complex convolutional neural networks.
[0014] The optical frequency comb light source and the optical frequency comb spectral shaping unit can be either thin-film lithium niobate devices integrated on a hybrid platform or off-chip discrete devices.
[0015] The thin-film lithium niobate modulator, along with other silicon-based microrings, PD detectors, etc., are integrated on a silicon-thin-film lithium niobate hybrid platform. Thin-film lithium niobate is used to realize optical frequency comb light source generation and convolutional layers; silicon is used to realize pooling, nonlinear, and fully connected layers.
[0016] The FPGA module includes an FPGA chip, a digital-to-analog (DAC) converter, and an analog-to-digital (ADC) converter chip. The FPGA chip controls the DAC chip and performs basic arithmetic operations and particle swarm optimization. The DAC chip outputs analog voltage and controls the thermal phase shifter and modulator. The ADC chip converts the analog electrical signals detected by the photodetector into digital electrical signals for subsequent processing by the FPGA.
[0017] The detector array can be an on-chip germanium-silicon PIN photodetector or an off-chip discrete group III-V PIN photodetector, and its function is to convert the output light intensity information into voltage information.
[0018] Furthermore, in practical applications, the particle swarm optimization algorithm is used to optimize the voltage value of the thermoelectric electrode to minimize the loss function. Its applications are not limited to optical neural networks.
[0019] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. The optical convolutional neural network structure based on thin-film lithium niobate and silicon hybrid provided by this invention uses high-speed frequency stream to perform convolution calculation. Compared with the redundancy caused by the traditional method of converting convolution calculation into matrix multiplication, this solution makes full use of existing wavelength resources and avoids performing complex Fourier transforms on electronic computers.
[0020] 2. The optical convolutional neural network structure based on thin-film lithium niobate and silicon hybrid provided by the present invention has a compact spatial scale and is not limited by the expansion of the convolution scale. Compared with the problems of large spatial resource consumption and poor scalability of existing solutions, the scale of its input vector depends only on the number of wavelengths of the optical frequency comb, and the scale of the convolution kernel depends only on the AWG sampling rate and driving power.
[0021] 3. The optical convolutional neural network structure based on thin-film lithium niobate and silicon hybrid provided by the present invention uses a hybrid lithium niobate platform and a silicon-based platform. Some modules can be monolithically integrated or hybrid integrated, or they can be used independently as separate modules, ensuring that each material is optimally selected according to its function. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of an optical convolutional neural network structure based on a mixture of thin-film lithium niobate and silicon, provided by an example of the present invention.
[0023] Figure 2 This is a schematic diagram of the implementation principle of a nonlinear layer in an optical convolutional neural network structure based on a hybrid thin-film lithium niobate chip and a silicon-based chip, provided by an example of the present invention. (a) The case where the input voltage of the active micro-ring modulator is less than the threshold voltage, and (b) The case where the input voltage of the active micro-ring modulator is greater than the threshold voltage.
[0024] Figure 3 This is a schematic diagram of a convolutional architecture based on a hybrid optical convolutional neural network structure of thin-film lithium niobate and silicon, provided by an example of the present invention.
[0025] Figure 4 This is a comparison diagram of several sets of convolution calculation experimental results and target results provided by the present invention based on a thin-film lithium niobate and silicon hybrid optical convolutional neural network structure. (a) and (e) are input vectors, (b) and (f) are comparison diagrams of convolution kernels obtained through experimental training and target convolution kernels, (c) and (g) are comparison diagrams of convolution results obtained in the experiment and target convolution results, and (d) and (h) are spectral diagrams of convolution results.
[0026] Figure 5 This is a schematic diagram of a real-domain convolutional architecture based on a hybrid optical convolutional neural network structure of thin-film lithium niobate and silicon, provided by an example of the present invention.
[0027] Figure 6 This is an experimental architecture diagram of an optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon, provided by an example of the present invention.
[0028] Figure 7 This invention provides the results of an optical convolutional neural network structure based on a hybrid thin-film lithium niobate and silicon on the MNIST handwritten digit set, and compares them with the results of simulation using the same architecture on a computer. (a) and (b) show the iterative curves and confusion matrices for online training to achieve binary classification, (e) and (f) show the iterative curves and confusion matrices for computer-based binary classification, (c) and (d) show the iterative curves and confusion matrices for online training to achieve four-class classification, and (g) and (h) show the iterative curves and confusion matrices for online training to achieve four-class classification. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0030] like Figure 1As shown, an example of the present invention discloses an optical convolutional neural network architecture based on a hybrid of thin-film lithium niobate chips and silicon-based chips, comprising an input information loading layer 1, a convolutional layer 2, a pooling layer 3, a nonlinear layer 4, and a fully connected layer 5. In the input information loading layer 1, a multi-wavelength optical frequency comb is generated by an optical frequency comb light source. The input information is loaded by a spectral shaping unit controlled by an FPGA via a DAC. In the convolutional layer 2, the optical frequency comb loaded with the input information completes convolution calculation in the frequency domain through a thin-film lithium niobate phase modulator. The convolution kernel is loaded into a PM by a signal generated by an AWG. In the pooling layer 3, the results of the convolution calculation are filtered in groups of several by a silicon-based passive microring to complete the average pooling operation. In the nonlinear layer 4, the pooling results are converted into an electrical signal by a PD detector to drive a silicon-based active microring. At the input end of the microring, a new optical frequency comb is input by another set of DFB lasers and thin-film lithium niobate microrings to complete the nonlinear operation. In the fully connected layer 5, the optical frequency comb obtained through nonlinear operations is input into the micro-ring array for matrix-vector multiplication. The weights of the micro-rings are controlled by the FPGA. Finally, PD probes are used at the download end of the micro-rings to complete the full connection. The final output is fed back to the FPGA via the ADC, and the network parameters are adjusted using the gradient descent algorithm to realize the function of online training of optical convolutional neural networks.
[0031] Figure 2 The implementation principle of nonlinear layer 4 in an optical convolutional neural network structure based on a hybrid of thin-film lithium niobate chip and silicon-based chip is demonstrated. This is the input voltage of the active micro-ring modulator (MRM) (driving a forward-biased PN junction). Consider a case where the resonant wavelength of the MRM... Initially related to the wavelength of the supplied light Alignment. For example... Figure 2 As shown in (a), when the input voltage of the MRM... Less than the threshold voltage When the potential (corresponding to the internal potential of the micro-ring PN junction) is reached, the PN junction remains in the off state, and no charge carriers are injected into the PN junction. Therefore, and Maintain alignment while outputting optical power It remains low because the supply light is filtered by the notch response of the MRM. For example... Figure 2 As shown in (b), when the drive current is large enough, such that Exceed At this time, the PN junction opens, and the injected charge carriers change the refractive index of the optical waveguide in the PN junction. As a result, The offset increases the output optical power.
[0032] Figure 3This paper demonstrates the principle of optical frequency-current convolution based on a high-speed lithium niobate modulator. In the experiment, we first generate single-wavelength light using a laser, which is then passed through thin-film lithium niobate (IM) and (PM) filters to generate an optical frequency comb. An optical waveshaper (WS) is used to weight the input wavelength power, thus loading the input information. The convolution is then performed using the thin-film lithium niobate (PM) filter, with the convolution kernel provided by an RF signal generated by an AWG. We assume the input light is... C in ( ω The modulation function (i.e., the convolution kernel) is... K ( ω The output light is C out ( ω Convolution calculation can be expressed as:
[0033] K ( ω ) is the phase modulation function h ( t The Fourier transform of ) is expressed as:
[0034] a m Let m be the m-th order amplitude of the modulated signal. J n Let be an nth-order Bessel function. In the frequency domain, this expression can be transformed into a convolutional form:
[0035] To obtain the target convolution kernel, we can directly input single-wavelength light into a thin-film lithium niobate PM. The output is then filtered and converted into an electrical vector by a PD before being input into an FPGA. The FPGA then uses a gradient descent algorithm to adjust the AWG (Automatic Gaussian Gear) until iterative convergence, and the output is the target convolution kernel. At this point, by adjusting the WS (Wire Bragg Controller) to load any input vector into the PM, arbitrary convolution calculations with this kernel can be performed. Figure 4 The experiment results for two sets of convolution calculations are shown. We first used the online training method described above to obtain... Figure 4 The convolution kernels shown in (b) and (f) are given, where the correlation coefficient is defined as: Then input two sets of input optical frequency combs (such as...). Figure 4 As shown in (a) and (e) in the figure), the corresponding two sets of convolution results are obtained (as shown in the figure). Figure 4 (as shown in (c) and (g)). Figure 4 (d) and (h) in the figure show that the two sets of convolution calculations were performed at frequency intervals of 4 GHz and 8 GHz, respectively.
[0036] Figure 5 This paper presents the schematic diagram of the real-domain computation of optical frequency-current convolution based on a high-speed thin-film lithium niobate modulator. Since real-domain convolution kernels are frequently required in image processing and other fields for operations such as edge extraction, we... Figure 3 A minor modification to the architecture enables it to perform convolution calculations in the real number domain. This is achieved through the data loading process... Figure 3 The architecture is consistent; in the convolution part, the input optical frequency comb is split into two paths, and convolution calculations are performed on the upper and lower lithium niobate PMs respectively. That is, convolution is performed with the positive and negative parts of the convolution kernel respectively. The results are filtered by micro-ring filters and then differentially analyzed by a balanced detector to complete the convolution in the real number domain. In the method of obtaining the real number domain convolution kernel through training, we do not need to care about the convolution kernel corresponding to the AWG signal loaded on each lithium niobate PM, but only need to ensure that the differential result is closer to the target real number domain convolution kernel.
[0037] Figure 6 The experimental architecture diagram of an optical convolutional neural network based on a hybrid of thin-film lithium niobate and silicon is shown. Here, the input information is loaded into layer 1 and convolutional layer 2. Figure 3 The structure is consistent with the previous one. Pooling layer 3 uses two microrings connected in parallel to realize a flat-top filter with a certain bandwidth to complete average pooling. The subsequent nonlinear units and fully connected units are completed in the electrical domain due to the limitations of the experimental devices. Here we mainly show its classification task on the MNIST handwritten digit dataset. We compress the 28×28 image to 5×5, and then convert it into a 25×1 vector and load it onto the input optical frequency comb. After convolution calculation by thin-film lithium niobate PM, the silicon-based microring filter performs pooling operation to obtain 8×1 vectors (2-class classification task) / 16×1 vectors (4-class classification task). After PD detection, nonlinear active function (NAF) module and fully connected (FC) module, it is fed back to FPGA. The FPGA then uses gradient descent algorithm to control AWG and WS to complete online training. 30 images are used as training set, and the iterative process is as follows. Figure 7 (a) (2-category task) and Figure 7 As shown in (c) (4-class classification task). After training, the confusion matrix obtained using 200 images as the test set is as follows. Figure 7 (b) (2-category task) and Figure 7 As shown in (d) (4-class classification task), the accuracies are 95.5% and 71.5%, respectively. This is comparable to the results of training and testing the same single-layer structure on a computer, with accuracies of 97.5% and 75.5%, respectively (e.g., ...). Figure 7(as shown in (e)-(h)). Judging from the application results of the optical neural network, the optical convolutional neural network based on a hybrid of thin-film lithium niobate and silicon proposed in this invention is entirely feasible.
[0038] This invention provides an optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon. This network exhibits unrestricted convolutional scaling and effectively avoids data redundancy inherent in traditional matrix multiplication by employing frequency-stream convolution. The input data fully utilizes existing wavelength resources and avoids the complex Fourier transforms required by electronic computers. Our fully on-chip hybrid integrated optical convolutional neural network allows for optimal selection of thin-film lithium niobate and silicon-based materials according to their functionalities, thereby achieving TOPS-level computing power and mW-level power consumption. This invention lays the foundation for future heterogeneous integrated optical computing architectures.
[0039] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon, characterized in that, It includes an input information loading layer, convolutional layers, pooling layers, non-linear layers, and fully connected layers; The input information loading layer is used to encode the input information onto the optical frequency comb teeth through spectral shaping; the convolutional layer is used to perform convolution calculation using a thin-film lithium niobate phase modulator; the pooling layer is used to perform average pooling on the convolution result; the nonlinear layer is used to perform nonlinear operation on the average pooling result; and the fully connected layer is used to perform matrix-vector multiplication. The convolutional layer includes a single thin-film lithium niobate phase modulator, which is driven by a high-speed radio frequency signal generated by an AWG. An optical frequency comb loaded with the input vector enters the thin-film lithium niobate phase modulator driven by the AWG signal for convolution calculation. The repetition frequency of this AWG signal must be the same as the repetition frequency of the driving signal of the input information loading layer, that is, the same as the frequency interval of the input vector. The pooling layer includes multiple microring arrays. Each microring array consists of a single microring or multiple microrings connected in series or in parallel to form a flat-top filter with a certain bandwidth, so that each microring array can sequentially filter out the wavelengths of multiple adjacent convolution results to achieve average pooling. The nonlinear layer includes an active silicon-based microring, which is driven by a photodetector. Each set of output wavelengths of the pooling layer enters a different photodetector to convert light intensity information into photocurrent. The photocurrent drives multiple active silicon-based microrings to generate nonlinearity, and each active microring loads the nonlinear response onto the corresponding wavelength. The fully connected layer includes M×N silicon-based microrings. The N wavelengths output by the pooling layer are divided into M parts. Each part enters a weight library consisting of N cascaded silicon-based microrings. Each microring controls the weight corresponding to a wavelength. The output uses a photodetector to sum the intensities, realizing M×N matrix vector multiplication.
2. The optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon according to claim 1, characterized in that, The input information loading layer includes an optical frequency comb light source and an optical frequency comb spectral shaping module. The optical frequency comb light source is used to generate a frequency comb, and the optical frequency comb shaping module is used to shape the intensity of the generated optical frequency comb teeth into the input vector of the neural network.
3. The optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon according to claim 1, characterized in that, The convolutional neural network converges through online training. The network's output port is connected to a field-programmable gate array (FPGA) via an ADC. The FPGA controls the convolution vectors generated by the AWG and is connected to a silicon-based microring array via a digital-to-analog converter (DAC), thereby controlling the loading of input information and the matrix of the fully connected layer, thus forming a feedback system.
4. The optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon according to claim 1, characterized in that, The input signal is split into two paths, and each path is convolved with a thin-film lithium niobate phase modulator. The results are then pooled and differentially analyzed by a balanced detector to complete the convolution calculation in the real domain.
5. The optical convolutional neural network structure based on a hybrid of thin-film lithium niobate and silicon according to claim 1, characterized in that, The convolutional layers, pooling layers, nonlinear layers, and fully connected layers of this network can be arbitrarily arranged, combined, or cascaded to expand the number of layers in a deep convolutional neural network, thereby enabling larger-scale and more complex convolutional neural networks.
Citation Information
Patent Citations
Photon convolutional neural network accelerator based on micro-ring resonator and nonvolatile phase change material
CN113657580A
Photonic neural network on silicon substrate based on tunable filter, and modulation method therefor
WO2022001002A1