Optoelectronic computer based on time-domain all-optical signal serial-parallel hybrid computation
Optoelectronic computers that use time-domain all-optical signal serial-parallel hybrid computing solve the problems of high energy consumption in traditional electronic computing systems and low integration and bandwidth loss in optical neural networks, achieving efficient neural network computing.
Patent Information
- Application Number
- CN202510613086.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Traditional electronic computing systems suffer from high energy consumption, latency due to frequent data transfer, and limited large-scale parallel processing capabilities in neural network computing. Optical neural network implementation schemes suffer from low integration and bandwidth loss.
The optoelectronic computer employing time-domain all-optical signal serial-parallel hybrid computing achieves parallel computation of matrix-vector multiplication and serial processing of nonlinear operations through electrical transceiver modules, light source modules, linear matrix-vector calculation modules, summation and separation modules, and nonlinear activation modules, thereby reducing hardware deployment pressure.
It improves computing speed, reduces device size and energy consumption, alleviates bandwidth loss, and enables high-speed, large-scale data computing.
Smart Images

Figure CN120447677B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular to an optoelectronic computer based on time-domain all-optical signal serial-parallel hybrid computing. BACKGROUND
[0002] In recent years, the rapid development of artificial intelligence and deep learning technology has put unprecedented demands on computing systems. Traditional electronic computing systems based on von Neumann architecture rely on the transmission of electronic signals and Boolean logic operations, and its inherent defects are increasingly evident in neural network computing. For example, the resistance characteristics of electronic devices result in a large amount of Joule heat during the operation process. For a typical GPU, the single-chip power consumption can reach more than 300W. As Moore's Law approaches the physical limit, the space for improving energy efficiency through process miniaturization has been saturated. At the same time, the frequent data transfer under the architecture of separation of storage and calculation causes serious delay. The weight matrix and activation value in the neural network need to be repeatedly migrated between the processor and the memory, resulting in up to 90% of the time and more than 60% of the energy consumption being consumed in the data communication link, rather than the multiply-add operation itself. Finally, the serial transmission characteristics of electrical interconnection limit the large-scale parallel processing capability. The typical PCIe 4.0 bus bandwidth is only 64GB / s, which is difficult to meet the demand for exponential growth of neural network model parameters (such as GPT-3 with 175 billion parameters).
[0003] Compared with the above, optical analog computing shows a breakthrough potential. Due to its ultra-low power consumption, integrated storage and calculation, and natural parallelism, it has attracted people's attention. For example, the wavelength / space / mode / time multiple degrees of freedom support ultra-wideband data channels, and a single optical pulse can complete NxN matrix operation (N>1000) synchronously. However, the current optical neural network implementation scheme still has some deficiencies. First, due to the low integration of optical devices, a large number of physical computing units will inevitably be caused to implement large-scale operation. Although time coding can improve the reuse rate of devices, it will also cause a certain bandwidth loss. SUMMARY
[0004] The present application aims to overcome the deficiencies of the prior art and provide an optoelectronic computer based on time-domain all-optical signal serial-parallel hybrid computing. According to the computing characteristics and signal coding methods of different data in the neural network, the present application performs serial-parallel hybrid computing in the time domain: for linear operations such as matrix-vector multiplication, parallel computing is performed in the time domain, which alleviates the bandwidth loss of pure serial time-domain computing methods and improves the computing speed of the computer; for summation and nonlinear operations, only one optical element is used to serially complete the summation operation and nonlinear activation operation on multiple multiplication vectors, thereby reducing the deployment pressure of the neural network on the hardware side.
[0005] The all-optical computer based on time-domain full-optical signal series-parallel hybrid computation provided by the application comprises an electrical transceiver module, a light source module, a linear matrix vector computation module, a summation separation module and a nonlinear activation module.
[0006] The electrical transceiver module is configured to control the weight matrix and the input vector to be emitted on the rectangular light waveform emitted by the light source module in parallel through electricity and to receive the signal output from the nonlinear activation module in series.
[0007] The light source module emits the isohypse rectangular light waveform with a certain time interval in series and sends the rectangular light waveform to the linear matrix vector computation module.
[0008] The linear matrix vector computation module performs parallel computation, accepts the m*n weight matrix w and the input vector x with a length of n input from the electrical transceiver module i , encodes the rectangular light waveform signal from the light source module and performs linear computation; the linear matrix vector computation module has m parallel computation units, and the m row multiplication vectors output by the m computation units are merged in parallel, and finally a parallel overlapping row multiplication vector set is generated in time.
[0009] The summation separation module receives the row multiplication vector set from the linear matrix vector computation module, completes the serial separation of the row multiplication vectors in the row multiplication vector set in time, and sums all the elements of each row multiplication vector to generate a multiplication accumulation vector.
[0010] The nonlinear activation module has a nonlinear activation function, which is configured to receive the multiplication accumulation vector from the summation module, perform nonlinear activation operation on all the elements of the multiplication accumulation vector in series, generate an output vector and collect the output vector by the electrical transceiver module.
[0011] According to the preferred scheme of the application, for the jth computation unit in the linear matrix vector computation module, the signal received from the electrical transceiver module is the jth row of the weight matrix and the complete input vector x i , and after computation, an n-element length row multiplication vector is generated, the elements of the row multiplication vector are arranged in time sequence, and each element has a duration t sp = Δt, so a row multiplication vector has a duration of n· Δt, and the jth computation unit generates a time delay of (j-1)· Δt, which is used to arrange the accumulation results of different row multiplication vectors in time in the summation separation module to distinguish them.
[0012] According to a preferred scheme of the present application, the electrical transceiver module comprises an arbitrary waveform generator (AWG), an oscilloscope, a photoelectric detector and a computer; the AWG receives a weight matrix and an input vector information from the computer and modulates them into electrical signals to the linear matrix vector calculation module; the photoelectric detector is used to detect the output of the nonlinear activation module and is converted into an electrical signal to be collected by the oscilloscope and returned to the computer to obtain the operation result.
[0013] According to a preferred scheme of the present application, the light source module comprises a mode-locked laser (MLL), a first dispersion fiber and a beam shaper; the MLL generates ideal Gaussian pulses with a short duration; the first dispersion fiber has a length L1 and a dispersion constant β2, which expands the ideal Gaussian pulse signal in the time domain by group velocity dispersion based on the frequency domain bandwidth of the optical pulse signal to obtain a broadened pulse signal; the beam shaper changes the broadened pulse signal into a flat rectangular light waveform in the time domain and enters the linear matrix vector operation module for modulation.
[0014] According to a preferred scheme of the present application, the linear matrix vector calculation module comprises an optical beam splitter, m linear calculation units and an optical beam combiner; the optical beam splitter is used to receive the rectangular light waveform from the light source module and broadcast it to the m linear calculation units; each linear calculation unit is composed of two electro-optical modulators and a time delay line connected in sequence; for the jth linear calculation unit, the first electro-optical modulator serially receives the jth row information of the weight matrix from the electrical transceiver module and modulates it onto the rectangular light waveform, each information having a duration Δt to obtain a time sequence signal; when the time sequence signal passes through the second electro-optical modulator, it receives the input vector information from the electrical transceiver module and sequentially modulates it onto the corresponding matrix elements to complete the multiplication operation of each element of a row of the matrix with each element of the vector, thereby obtaining a row multiplied vector; the row multiplied vector passes through a time delay line with a time delay capability of (j-1)·Δt and is sent to the optical beam combiner; the optical beam combiner receives the parallel overlapping row multiplied vector set obtained by combining the m row multiplied vectors and inputs it into the summation separation module.
[0015] According to a preferred scheme of the present application, the summation separation module comprises a second dispersion fiber; the second dispersion fiber has a dispersion coefficient -β2 and a length L2; the summation separation module sums each row multiplied vector in the row multiplied vector set to obtain a multiply-accumulate vector. After summation, the value of the multiply-accumulate vector is obtained, and the summation is based on compression in the time domain. After compression is completed, the values of the row multiplied vectors are naturally separated in time, thereby realizing the function of separation. After the summation of the m row multiplied vectors, m multiply-accumulate values are obtained, and the vector composed of the m values is called a multiply-accumulate vector.
[0016] According to a preferred scheme of the present application, the nonlinear activation module uses a semiconductor optical amplifier for nonlinear activation.
[0017] Compared with the prior art, the application encodes the information flow by timing, so that the information flow does not excessively depend on the physical carrier, and meanwhile, the serial and parallel technologies are reasonably used in different calculation units to reduce the bandwidth loss problem caused by the timing encoding.
[0018] The application uses parallel operation when performing matrix-vector element-by-element multiplication, and since the summation operation and the nonlinear operation do not depend on the input data, the data can be encoded in a serial manner. The operation of a fully connected neural network includes three steps of matrix-vector multiplication, multiplication result accumulation and nonlinear activation, and the application combines the serial and parallel calculation modes according to the data characteristics, so that high-speed large-data calculation can be realized with a small amount of physical device requirements, and all mathematical operation operations of any layer of a fully connected neural network can be realized. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 The application is based on the architecture flowchart of the optoelectronic computer of the time-domain all-optical signal serial-parallel hybrid calculation.
[0020] Figure 2 The application is a row multiplication vector delay component row multiplication vector set diagram.
[0021] Figure 3 The application is a summation separation module separating parallel signals and summing up.
[0022] Figure 4 The application is a mathematical function of a part of the physical structure implementation in the artificial neural network.
[0023] Figure 5 The application is a structure composition diagram of embodiment one. DETAILED DESCRIPTION
[0024] The application will be described in detail below with reference to the drawings.
[0025] As shown in Figure 1 , the application provides an optoelectronic computer based on time-domain all-optical signal serial-parallel hybrid calculation, which includes an electrical transceiver module, a light source module, a linear matrix vector calculation module, a summation separation module and a nonlinear activation module.
[0026] The electrical transceiver module is used to parallel control the weight matrix and the input vector to the rectangular light waveform emitted by the light source module in an electrical manner, and serially receive the signal output from the nonlinear activation module.
[0027] The light source module serially emits the isohypse rectangular light waveform with a certain time interval and continuously in time, and sends the rectangular light waveform to the linear matrix vector calculation module.
[0028] The linear matrix-vector calculation module performs parallel computation, accepting an m×n dimensional weight matrix w (m rows and n columns) and an input vector x of length n from the electrical input signal. i The signal is encoded into a rectangular light waveform signal from the light source module and linearly calculated. Specifically, the linear matrix-vector calculation module has m parallel calculation units. For the j-th (j<=m) calculation unit, the signal received from the electrical transceiver module is the j-th row of the weight matrix and the complete input vector x. i Then, perform element-wise multiplication of the two in the time domain, and generate a row multiplication vector of length n elements:
[0029] x n =(w j1 ·x i1 ,w j2 ·x i2 ,…,w jn ·x in (1)
[0030] Each element of the row multiplication vector is ordered by time sequence and has a duration Δt, so a row multiplication vector has a duration of n·Δt. Simultaneously, the j-th computation unit generates a delay of (j-1)·Δt to sort the accumulated results of different row multiplication vectors in time for differentiation in the summation and separation module. The m row multiplication vectors from the m computation units are merged in parallel, finally producing a parallel and overlapping set of row multiplication vectors x in time. o Specifically, such as Figure 2 As shown, Figure 2 Taking three computational units as an example, the process of assembling a row multiplication vector set from the time delays of the three row multiplication vectors is demonstrated. The time axis is divided into time frames of Δt units. For the row multiplication vector generated by the first computational unit, there is no time delay. For the row multiplication vector generated by the second computational unit, there is a time delay of Δt. Therefore, the first time frame contains the first term of the row multiplication vector generated by the first computational unit: w. 11 *x i1 The second time frame contains the second term of the row multiplication vector generated by the first computation unit and the first term of the row multiplication vector generated by the second computation unit. For ease of explanation, the set of row multiplication vectors is written in column vector form:
[0031]
[0032] Each element of this vector has a duration Δt and will be sent to the summation and separation module.
[0033] The summation and separation module receives the row multiplication vector set x from the linear matrix-vector calculation module. o and complete the work on xo element-wise serial separation and summation in time domain, resulting in a multiply-accumulate vector x s :
[0034]
[0035] x s Each element will have its own duration and be separated in time from other elements, and at this point, the summation separation is complete. Figure 3 A schematic diagram showing a 3-way parallel computation result in time domain, serial aliasing, and separation and summation of corresponding elements by a summation separation module. After passing through the summation separation module, the multiply-accumulate vector will be input to a nonlinear activation module.
[0036] The nonlinear activation module has a nonlinear activation function, which is used to receive the multiply-accumulate vector from the summation module and serially perform nonlinear activation operations on all elements of the multiply-accumulate vector, resulting in an output vector and collected by an electrical transceiver module.
[0037] The present application will be further described below in conjunction with specific cases:
[0038] Example 1
[0039] Example 1 is a serial-parallel computation optical neural network based on a dispersion optical fiber time stretching method, which can be implemented by all-optical analog computation Figure 4 A mathematical operation process of a fully connected neural network is shown.
[0040] In this embodiment, the electrical transceiver module is composed of an arbitrary waveform generator (AWG), an oscilloscope, a photodetector, and a personal computer. The AWG is used to receive weight matrix and vector information from the personal computer and modulate them in parallel as electrical signals to the linear matrix vector computation module.
[0041] In this embodiment, the light source module is composed of a mode-locker laser (MLL), a dispersion optical fiber 1, and a beam shaper. The MLL generates an ideal Gaussian pulse E1 with a short duration t1:
[0042]
[0043] where E0 is the pulse peak power, and Ω, m represents the sharpness and is inversely proportional to the sharpness. According to the dispersion Fourier transform, the corresponding frequency domain signal S1 of the pulse is as follows:
[0044]
[0045] where ω is the angular frequency of the pulse, and j is a complex term.
[0046] The dispersion fiber 1 has a length L1 and a dispersion constant β2, which can spread an ideal Gaussian pulse in time domain by group velocity dispersion (GVD) based on the frequency domain bandwidth of the optical pulse signal. Specifically, assuming that the spreaded pulse has a time duration t2 and a spectral bandwidth S2, then:
[0047]
[0048] By inverse Fourier transform The time domain signal duration of the spreaded pulse is:
[0049]
[0050] It is known that t2>t1, and the spreaded pulse signal becomes a flat rectangular optical waveform in time domain after passing through a waveform shaper, and enters a linear matrix vector operation module for modulation.
[0051] In this embodiment, to calculate an m×n dimensional weight matrix and an input vector with length n, the linear matrix vector calculation module is composed of an optical beam splitter, 2m electro-optic modulators, m optical time delay lines, and an optical combiner. The optical beam splitter is used to receive the rectangular optical waveform from the light source module and broadcast it equally to m linear calculation units. For the jth linear calculation unit, it is composed of two electro-optic modulators and a time delay line connected in sequence. The first electro-optic modulator serially receives the jth row information of the weight matrix from the AWG and modulates it onto the rectangular optical waveform, each information having a time duration Δt, to obtain a time sequence signal W tj :
[0052] W tj =(w j1 ,w j2 ,…,w jn ) (8)
[0053] When the time sequence signal passes through the second electro-optic modulator, it receives the input vector information from the AWG and sequentially modulates it onto the corresponding matrix elements, completing the matrix row and vector multiplication operation, to obtain the row multiplied vector as formula (1).
[0054] The m linear calculation units complete the calculation of m row multiplied vectors in parallel at the same time, greatly reducing the system calculation time. For the jth linear calculation unit, its row multiplied vector passes through a time delay line with a time delay capability of (j-1)·Δt, and is sent to the optical combiner. The optical combiner receives m row multiplied vectors, combines to obtain the parallel overlapping vector x o , which is input into the sum separation module.
[0055] The summation separation module is implemented by the dispersive fiber 2. Specifically, each element of the vector x o contains the intensity calculation information of a plurality of parallel computing units, although they are time-parallel coincident, since the previous calculation operation only modulates the intensity of the optical waveform and does not change the frequency spectrum of the signal, the frequency domain signal of any computing unit is independent, and then the inverse Fourier transform is performed, the dispersive fiber 2 has a dispersion coefficient -β2 and a length L2, and the total phase of the optical signal after summation separation compared with the pulse signal of the light source module after two dispersions is:
[0056]
[0057] The corresponding pulse duration is:
[0058]
[0059] As shown in equation (10), the time t3 < t2, so that the stretched pulse signal is compressed again. The accumulated compressed pulse reflects the result of the signal multiplied by the single-row weight matrix and the input xi. Therefore, the multiplication of the weight matrix and the input vector can be realized in parallel by modulating different row weights in a plurality of rectangular optical waveforms.
[0060] Thus, the multiply-accumulate vector x s is obtained and input to the nonlinear activation module, and a semiconductor optical amplifier is used for nonlinear activation x fs = fs(x s ), wherein fs is the nonlinear transmission function of the semiconductor optical amplifier. Finally, x fs is converted into an electrical signal by a photodetector and collected by an oscilloscope, and the operation result is obtained by returning to a computer.
[0061] From the operation speed, the time for performing a complete calculation of the neural network of the n input vectors is n·Δt. If the matrix-vector multiplication operation is directly performed in a serial manner, m serial calculations need to be performed on the n-dimensional vector, each calculation needs a duration of n·Δt, and the total calculation time is m·n·Δt. The present application performs the time-consuming matrix-vector multiplication operation in parallel by using the method that the spectrum of the waveform signal is unchanged, which is more time-saving than directly performing the matrix-vector multiplication operation in a serial manner.
[0062] From the device utilization rate, since the m row multiplication vectors have the same spectral physical parameters, the summation and nonlinear operation operations each use only one optical device to complete the corresponding functions, which greatly relieves the device size pressure.
[0063] The above embodiment one is only a preferred embodiment, and the purpose, technical scheme and advantages of the present application are further described in detail. It should be emphasized that the above-described embodiment does not limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. An opto-electronic computer based on time-domain all-optical signal serial-parallel hybrid computation, characterized by, The electrical transceiver module is used for regulating the weight matrix and the input vector to the rectangular light waveform emitted by the light source module in parallel through electricity, and receiving the signal output from the nonlinear activation module in series; the electrical transceiver module comprises an arbitrary waveform generator (AWG), an oscilloscope, a photodetector and a computer; the AWG receives the weight matrix and the input vector information from the computer and modulates them to the linear matrix vector calculation module in parallel in the form of electrical signals; the photodetector is used for detecting the output of the nonlinear activation module and converting it into an electrical signal which is collected by the oscilloscope and returned to the computer to obtain the operation result. The light source module emits the same rectangular light waveform which is continuous in time with a certain time interval in series, and sends the rectangular light waveform to the linear matrix vector calculation module. The linear matrix vector calculation module comprises an optical beam splitter, m linear calculation units and an optical beam combiner. The linear matrix vector calculation module performs parallel calculation, accepts an m*n weight matrix w and an input vector x with a length of n from an electrical transceiving module i , encodes the rectangular light waveform signal from the light source module and performs linear calculation; the linear matrix vector calculation module has m parallel calculation units, m row multiplication vectors output by the m parallel calculation units are merged in parallel, and finally a set of parallel overlapping row multiplication vectors in time is generated; The optical beam splitter is used for receiving the rectangular light waveform from the light source module and broadcasting it to the m linear calculation units in m equal parts. Each linear calculation unit is composed of two electro-optical modulators and a time delay line connected in sequence; for the jth linear calculation unit, the first electro-optical modulator serially receives the jth row information of the weight matrix emitted by the electrical transceiver module and modulates it to the rectangular light waveform, each information having a duration Δt to obtain a time sequence signal; when the time sequence signal passes through the second electro-optical modulator, the input vector information emitted by the electrical transceiver module is received and sequentially modulated to the corresponding matrix element to complete the multiplication operation of each element of a row of the matrix and each element of the vector, thereby obtaining a row multiplied vector; the row multiplied vector passes through a time delay line with a time delay capacity of (j-1)·Δt and is sent to the optical beam combiner. The optical beam combiner receives the parallel overlapping row multiplied vector set combined by the m row multiplied vectors and inputs it into the summation separation module. The summation separation module receives the row multiplied vector set of the linear matrix vector calculation module, completes the serial separation of the row multiplied vectors in the row multiplied vector set in the time domain, and sums all the elements of each row multiplied vector to generate a multiply-accumulate vector. The nonlinear activation module has a nonlinear activation function, which is used for receiving the multiply-accumulate vector from the summation module and performing nonlinear activation operation on all the elements of the multiply-accumulate vector in series to generate an output vector which is collected by the electrical transceiver module. The light source module comprises a mode-locked laser (MLL), a first dispersion fiber and a beam shaper.
2. The opto-electronic computer based on time-domain all-optical signal series-parallel hybrid computation of claim 1, wherein, For the jth computing unit in the linear matrix vector computation module, it receives the signal from the electrical transceiver module as the jth row of the weight matrix and the complete input vector x i , and after computation, it generates a row-by-vector with n elements, the elements of which are ordered in time series, and each element has a duration t sp = Δt, so a row-by-vector has a duration of n· Δt, and the jth computing unit generates a time delay of (j-1)· Δt, which is used to order the accumulated results of different row-by-vectors in time in the summation separation module for differentiation.
3. The opto-electronic computer based on time-domain all-optical signal series-parallel hybrid computation of claim 1, wherein, The MLL generates ideal Gaussian pulses with a short duration; The first dispersion fiber has a length L1 and a dispersion constant β2, which expands the ideal Gaussian pulse signal in the time domain based on the frequency domain bandwidth of the optical pulse signal through group velocity dispersion to obtain a broadened pulse signal; The beam shaper changes the broadened pulse signal into a rectangular light waveform which is flat in the time domain and enters the linear matrix vector operation module for modulation. 4. The opto-electronic computer based on time-domain all-optical signal series-parallel hybrid computation of claim 1, wherein, The summation separation module includes a second dispersion fiber; the second dispersion fiber has a dispersion coefficient -β2 and a length L2; and the multiplication and accumulation vector is obtained by time compression.
5. The optoelectronic computer based on time-domain all-optical signal series-parallel hybrid computation of claim 1, wherein, The nonlinear activation module uses a semiconductor optical amplifier for nonlinear activation.
Citation Information
Patent Citations
Parallel optical computing system for efficiently realizing large-scale matrix operation
CN110908428A
Photon-assisted optical series-to-parallel conversion system and optical communication apparatus using same
CN111245553A