Optical convolution accelerator and method of controlling the same

By designing an optical convolution accelerator, efficient matrix operations were achieved using optical elements. This solved the problem of a linear relationship between device size and computational scale in existing technologies, improved the system's scalability and flexibility, reduced power consumption and complexity, and achieved high-bandwidth, low-latency computing performance.

CN118607604BActive Publication Date: 2025-12-26SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410481479.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2025-12-26
Estimated Expiration
2044-04-22

AI Technical Summary

Technical Problem

Existing optical matrix operation schemes exhibit a linear or quadratic relationship between device size and operation scale, resulting in poor scalability, high power consumption, insufficient flexibility, and high design and manufacturing complexity, making it difficult to achieve matrix operations of arbitrary size and form.

Method used

An optical convolution accelerator, including a light source, intensity modulator, micro-ring modulator, and photodetector, is used to achieve the modulation and superposition of the convolution kernel and the information to be processed through time-division multiplexing. The convolution operation is performed by controlling the delay using the micro-ring structure. Vector dot product and matrix multiplication of arbitrary length can be achieved by adjusting the sequence length and coupling rate.

Benefits of technology

It improves the system's scalability and flexibility, reduces power consumption and complexity, achieves high-bandwidth and low-latency matrix operations, and improves computing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118607604B_ABST
    Figure CN118607604B_ABST
Patent Text Reader

Abstract

The application discloses an optical convolution accelerator and a control method thereof, comprising a light source, an intensity modulator, a micro-ring modulator and a photoelectric detector; the light source is used for generating an optical carrier signal; the intensity modulator is used for modulating the optical carrier signal according to a first linear sequence on a time domain to obtain an optical intensity signal; the first linear sequence carries information to be processed; the micro-ring modulator is used for modulating a second linear sequence of convolution kernel elements into the optical intensity signal in a time division multiplexing manner to obtain a modulation sub-signal, and superimposes modulation sub-signals of different code element periods through delay to obtain an output signal; the output signal carries a convolution operation result of the first linear sequence and the second linear sequence; and the photoelectric detector is used for detecting the output signal. The embodiment of the application can improve expandability and flexibility, reduce complexity and power consumption, and can be widely applied to the field of optical computing technology.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of optical computing, in particular to an optical convolution accelerator and a control method thereof. BACKGROUND

[0002] With the emergence of increasingly complex artificial intelligence service scenarios such as autonomous driving, chatGPT, edge computing, or image processing, the demand for the computing power of neuromorphic hardware is also growing. However, with the gradual invalidation of Moore's Law, the expansion of the computing power of electronic integrated circuits is facing practical challenges due to the limitations of transistor density, the increase in data storage, mobile energy consumption, and the intensification of non-ideal effects such as transistor device leakage. In contrast, photonic devices and systems have the advantages of simple structure, higher transmission rate, larger transmission bandwidth, and parallelism, making optical accelerators have unique advantages in the field of neuromorphic hardware that requires parallel computing and transmission of large amounts of data. Under the same level of computing power, existing optical computing devices have orders of magnitude advantages in energy consumption and delay compared to AI chips in the digital field such as NVIDIA GPUs. These advantages are sufficient to promote the development of photonic computing accelerators to meet the growing demand for data at the hardware level. On this basis, using the optical chip manufacturing technology in integrated photonics and the compatibility with electronic circuits, the resulting optoelectronic hybrid computing architecture has high programmability and integrability, which is more conducive to the fusion of real-time computing in the optical domain and traditional integrated circuits in AI operations, thereby applying the advantages of optical computing to the current device manufacturing process and optimizing and innovating existing architectures.

[0003] Due to the relative lack of storage, nonlinearity, and other functions, researchers mainly use light to accelerate logical operations, especially matrix operations of neural networks in neuromorphic hardware. In order to apply the advantages of high bandwidth and high multiplexing of light, researchers have proposed three directions for implementing optical matrix operations based on space division multiplexing, wavelength division multiplexing, and time division multiplexing. Existing optical matrix operation schemes have the following problems: (1) the size of the system required by the device is linearly or even quadratically related to the size of the matrix being operated, and the scalability is poor; in addition, the increase in the number of system devices greatly increases the chip footprint and power consumption; (2) different on-chip architectures are usually required for different types of matrices, and many schemes can only implement operations between a fixed number of matrices and other matrices, and cannot quickly and flexibly implement any type of operation between matrices of any size, any form, and any value, lacking flexibility; (3) many time division multiplexing schemes require various types of integrators (such as photodetectors) to integrate high-frequency signals in the time domain with high accuracy, which requires high response and bandwidth of the integrator, greatly increasing the difficulty of design and manufacture. SUMMARY

[0004] Therefore, to solve one of the above problems, an optical convolution accelerator and a control method thereof are provided to improve scalability and flexibility, reduce complexity and power consumption.

[0005] In one aspect, an optical convolution accelerator is provided, comprising a light source, an intensity modulator, a micro-ring modulator and a photodetector, wherein,

[0006] The light source is configured to generate an optical carrier signal.

[0007] The intensity modulator is configured to modulate the optical carrier signal according to a first linear sequence in time domain to obtain an optical intensity signal; the first linear sequence carries to-be-processed information.

[0008] The micro-ring modulator is configured to modulate a second linear sequence of convolution kernel elements into the optical intensity signal in a time division multiplexing manner to obtain a modulation sub-signal, and superimpose modulation sub-signals of different symbol periods by time delay to obtain an output signal; the output signal carries a convolution operation result of the first linear sequence and the second linear sequence.

[0009] The photodetector is configured to detect the output signal.

[0010] Optionally, the micro-ring modulator comprises a first 2*2 multimode interferometer, a first Mach-Zehnder electro-optic intensity modulator, a second 2*2 multimode interferometer and a micro-ring structure, a first input end of the first 2*2 multimode interferometer is connected to an output end of the intensity modulator, two output ends of the first 2*2 multimode interferometer are respectively connected to two input ends of the first Mach-Zehnder electro-optic intensity modulator, two output ends of the first Mach-Zehnder electro-optic intensity modulator are respectively connected to two input ends of the second 2*2 multimode interferometer, a first output end of the second 2*2 multimode interferometer is connected to the photodetector, and a second output end of the second 2*2 multimode interferometer is connected to a second input end of the first 2*2 multimode interferometer through the micro-ring structure.

[0011] Optionally, the intensity modulator comprises a second Mach-Zehnder electro-optic intensity modulator.

[0012] In another aspect, a control method of an optical convolution accelerator is provided, comprising:

[0013] Arranging to-be-processed information into a first linear sequence in time domain according to a processing order, and modulating an intensity modulator with the first linear sequence to obtain an optical intensity signal;

[0014] The convolution kernel elements are arranged into a second linear sequence according to an execution sequence, and the micro-ring modulator is modulated in a time-division multiplexing manner to obtain a modulation sub-signal, and the modulation sub-signals of different symbol periods are superimposed by delay control; when the first linear sequence modulation is completed, the superimposed modulation sub-signal is controlled to be output, and the photoelectric detector detects the output signal;

[0015] The convolution operation result is determined according to the output signal.

[0016] Optionally, the information to be processed includes image information, and the information to be processed is arranged into a first linear sequence in a time domain according to a processing sequence, and specifically includes:

[0017] The image information is converted into an input matrix.

[0018] The input matrix is divided into blocks according to the size of the convolution kernel elements, and the elements in each block are arranged into a first linear sequence in a time domain according to the processing sequence.

[0019] Optionally, the delay is controlled by a micro-ring structure, and when the first linear sequence includes a type of information to be processed, the delay period of the micro-ring structure is the same as one symbol period.

[0020] Optionally, the delay is controlled by a micro-ring structure, and when the first linear sequence includes N types of information to be processed, the delay period of the micro-ring structure contains N symbol periods, and N is a natural number greater than 1.

[0021] On the other hand, an embodiment of the present application provides a control system of an optical convolution accelerator, comprising:

[0022] A first module is configured to arrange information to be processed into a first linear sequence in a time domain according to a processing sequence, and modulate an intensity modulator by using the first linear sequence to obtain a light intensity signal.

[0023] A second module is configured to arrange convolution kernel elements into a second linear sequence according to an execution sequence, modulate a micro-ring modulator in a time-division multiplexing manner to obtain a modulation sub-signal, and superimpose the modulation sub-signals of different symbol periods by delay control; when the first linear sequence modulation is completed, the superimposed modulation sub-signal is controlled to be output, and a photoelectric detector detects the output signal.

[0024] A third module is configured to determine a convolution operation result according to the output signal.

[0025] On the other hand, an embodiment of the present application provides a control device of an optical convolution accelerator, comprising:

[0026] At least one processor;

[0027] At least one memory is configured to store at least one program.

[0028] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0029] In another aspect, an embodiment of the present application provides a computer readable storage medium, which stores a processor executable program, and the processor executable program is used for executing the above method when executed by a processor.

[0030] Implementing an embodiment of the present application includes the following advantages:

[0031] (1) Since only a single device is used to complete large-scale matrix operations, and by adjusting the sequence length and micro-ring coupling rate, vector dot product operations of any length can be realized, and it is extended to matrix multiplication. Therefore, compared with the space division multiplexing device array scheme in which the number of devices increases quadratically with the size of matrix operation, the device has higher scalability.

[0032] (2) The device can realize matrix multiplication and convolution operation of different scales by only changing the modulation rate of the data sequence and its mapping method without changing the device structure and peripherals, and its configuration flexibility is much better than other schemes.

[0033] (3) The scheme only uses time division multiplexing data sequence input Mach-Zehnder modulator and micro-ring to realize matrix operation, without using wavelength or spatial dimension, which reduces the design and manufacturing complexity of the system device, and greatly reduces the power consumption and delay of the entire system. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 is a structural schematic diagram of an optical convolution accelerator provided by an embodiment of the present application;

[0035] Figure 2 is a structural schematic diagram of a micro-ring modulator provided by an embodiment of the present application;

[0036] Figure 3 is a step flowchart of a control method of an optical convolution accelerator provided by an embodiment of the present application;

[0037] Figure 4 is an operation flowchart of an optical convolution accelerator provided by an embodiment of the present application;

[0038] Figure 5 is an experimental result diagram of extracting an image edge by using an optical convolution accelerator provided by an embodiment of the present application;

[0039] Figure 6 is an operation flowchart of another optical convolution accelerator provided by an embodiment of the present application;

[0040] Figure 7 is a structural block diagram of a control system of an optical convolution accelerator provided by an embodiment of the present application;

[0041] Figure 8 is a structural block diagram of a control device of an optical convolution accelerator provided by an embodiment of the present application. DETAILED DESCRIPTION

[0042] The present application will be further described below in conjunction with the drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description, and the order between the steps is not limited in any way, and the execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0043] As shown in Figure 1 , an embodiment of the present application provides an optical convolution accelerator, comprising a light source 1-1, an intensity modulator 1-2, a micro-ring modulator 1-3 and a photodetector 1-4; wherein,

[0044] The light source 1-1 is configured to generate an optical carrier signal.

[0045] The intensity modulator 1-2 is configured to modulate the optical carrier signal according to a first linear sequence in the time domain to obtain an optical intensity signal; the first linear sequence carries to-be-processed information.

[0046] The micro-ring modulator 1-3 is configured to modulate the optical intensity signal according to a second linear sequence of convolution kernel elements in a time division multiplexing manner to obtain a modulated sub-signal, and to superimpose the modulated sub-signals of different symbol periods by delaying to obtain an output signal; the output signal carries a convolution operation result of the first linear sequence and the second linear sequence.

[0047] The photodetector 1-4 is configured to detect the output signal.

[0048] Specifically, the light source 1-1 can adopt a distributed feedback laser, and the photodetector 1-4 can adopt a photodiode.

[0049] Optionally, referring to Figure 2The micro-ring modulator 1-3 comprises a first 2x2 multimode interferometer 1-3-1, a first Mach-Zehnder electro-optic intensity modulator 1-3-2, a second 2x2 multimode interferometer 1-3-3 and a micro-ring structure 1-3-4. The first input end A of the first 2x2 multimode interferometer 1-3-1 is connected to the output end of the intensity modulator. The two output ends of the first 2x2 multimode interferometer 1-3-1 are respectively connected to the two input ends of the first Mach-Zehnder electro-optic intensity modulator 1-3-2. The two output ends of the first Mach-Zehnder electro-optic intensity modulator 1-3-2 are respectively connected to the two input ends of the second 2x2 multimode interferometer 1-3-3. The first output end C of the second 2x2 multimode interferometer 1-3-3 is connected to the photodetector. The second output end D of the second 2x2 multimode interferometer 1-3-3 is connected to the second input end B of the first 2x2 multimode interferometer through the micro-ring structure 1-3-4.

[0050] It should be noted that the micro-ring modulator can be integrated on a thin film lithium niobate substrate.

[0051] The input vector signal E of the time-division multiplexed optical amplitude modulation in The operation result E stored in the ring input from the input end A and the B port r ′ The interleaving and superposition are performed by the first 2x2 multimode interferometer:

[0052]

[0053] After the interleaving and superposition by the first 2x2 multimode interferometer, the optical output of the upper arm output end E of the first 2x2 multimode interferometer is E up , the optical output of the lower arm output end F is E down , T represents the optical intensity ratio of the upper and lower arms of the first 2x2 multimode interferometer, i represents the imaginary part of the complex number, and t represents time.

[0054] The other input vector (i.e., the second linear sequence, the vector weight signal in the weighted superposition) is input from the internal electrical signal input end of the structure in a push-pull mode in a manner that is time-sequenced and has the same symbol rate as the input vector signal. The two arms of different lengths change the optical path Δl x of the two arms of the first Mach-Zehnder electro-optic intensity modulator in opposite directions by changing the refractive index of the lithium niobate thin film waveguide. m , l m + Δl m is the length of the lower arm optical waveguide of the first Mach-Zehnder electro-optic intensity modulator, α is the decay rate of the optical intensity with the waveguide length (m -1 ), and n gLet λ be the refractive index of lithium niobate and λ be the current wavelength of light, thus changing the phase of the light signal in both arms:

[0055]

[0056] The optical output of the upper arm output terminal H after modulation by the 2×2 Mach-Zehnder electro-optic intensity modulator is E. ′ up The light output of the lower arm output terminal M is E. ′ down The second 2×2 multimode interferometer interweaves and integrates the optical signals on both arms of the first Mach-Zehnder electro-optic intensity modulator, enabling it to convert changes in the phase of the optical signal into changes in the intensity of the light signal by superimposing opposite phases.

[0057]

[0058] After the interleaving and superposition of the second 2×2 multimode interferometer, the optical output signal at output port C is E. out The optical signal coupled into the microring from port D is E. r .

[0059] After the input vector signal is modulated by the weighted signal, a multiplication operation is completed. Subsequently, the result travels from port D through a micro-loop structure and then to port B, where it is time-aligned with the vector signal input from port A in the next symbol period for accumulation.

[0060]

[0061] Passing through a length of l r After the micro-ring is delayed and attenuated, the optical signal E representing the current calculation result is... r ′ The signal is output from port B in the next symbol cycle. To achieve strict timing alignment, the symbol period T of the input signal... B With the micro-ring delay T in the whole structure s A strict match is required to deduce the input signal symbol rate (note: the speed of light in a vacuum is c) as follows:

[0062]

[0063] Finally, after all the input vector data has been fed into the microring modulator structure, the energy stored in the microring is the result of this vector dot product operation. The D-port signal E represents the result of the operation. r With port E of B r ′The micro-ring is always in a loop, so when the operation result is obtained, the signal encoded in amplitude in the micro-ring needs to be introduced into the output port, and the result of this operation is read from the output port C using a photodiode and subsequent data processing is performed.

[0064] Optionally, referring to Figure 1 , the intensity modulator comprises a second Mach-Zehnder electro-optic intensity modulator.

[0065] Specifically, the second Mach-Zehnder electro-optic intensity modulator can adopt a commercial Mach-Zehnder intensity modulator.

[0066] In a specific embodiment, after the optical signal is emitted from the distributed feedback laser, the input vector data is modulated by the Mach-Zehnder intensity modulator first, then the weight vector data is modulated by the micro-ring modulator chip, and the multiplication result of the two is superimposed, and the final operation result is output by the port, and after being converted into an electrical signal by a photodiode, it enters a digital computer for subsequent data processing.

[0067] Referring to Figure 3 , the embodiment of the present application provides a control method of an optical convolution accelerator, comprising:

[0068] S100, arranging the information to be processed according to a processing order into a first linear sequence in time domain, and modulating an intensity modulator by using the first linear sequence to obtain an optical intensity signal;

[0069] S200, arranging the convolution kernel elements according to an execution order into a second linear sequence, modulating a micro-ring modulator in a time division multiplexing manner to obtain a modulated sub-signal, and superimposing the modulated sub-signals of different symbol periods by delay control; when the first linear sequence modulation is completed, the superimposed modulated sub-signal is output, and a photodetector is controlled to detect the output signal;

[0070] S300, determining a convolution operation result according to the output signal.

[0071] Specifically, referring to Figure 4The to-be-processed information to be convolved is rearranged in a certain order into a linear sequence in time domain. Then, the linear sequence of electrical signals in time domain is modulated onto an optical carrier through a Mach-Zehnder electro-optical intensity modulator to generate an optical intensity signal in which matrix element information is encoded into optical amplitude information, and the optical intensity signal is input into a first input port in the micro-ring modulator structure. By loading a linear sequence of convolution kernel elements into the electrical signal input port of the micro-ring modulator structure in a time-division multiplexing manner at the same time, the value of the weight vector element at the corresponding position is modulated into the optical input signal in the same symbol period through the Mach-Zehnder intensity modulator in the micro-ring modulator structure, so as to realize the multiplication operation on the time sequence of the injected optical signal. Then, the feedback loop optical amplitude of the feedback port of the micro-ring in the micro-ring modulator structure is adjusted by adjusting the micro-ring coupling rate of the micro-ring modulator through the internal modulator, and the overlap of optical signals in different symbol periods in time sequence is realized after the multiplication operation through the micro-ring delay, so as to realize the accumulation operation in the micro-ring modulator structure in the optical domain. After the multiplication and addition operation of all elements of the input vector and the weight vector is completed, the dot product operation result of the two vectors is output from the output port.

[0072] In one specific embodiment, the matrix multiplication is realized by using the structure in Figure 1 . According to formula (1), any two-dimensional convolution operation can be converted into a multiplication operation between a matrix and a vector by element rearrangement, taking a 2*2 convolution kernel as an example. Figure 4 In the structure in Figure 1 , the data is first converted into a matrix multiplication according to the manner in formula (1), and is input into the micro-ring modulator device in a time sequence linear sequence, so as to realize the convolution operation, as shown in formula (6).

[0073]

[0074] Optionally, the to-be-processed information includes image information, and the to-be-processed information is arranged into a first linear sequence in time domain according to a processing order, and specifically includes:

[0075] S110, converting the image information into an input matrix;

[0076] S120, dividing the input matrix according to the size of the convolution kernel elements, and arranging the elements in each block into a first linear sequence in time domain according to a processing order.

[0077] It should be noted that the size of the convolution kernel elements is determined according to actual application, and the embodiment does not make specific limitation. The processing order is determined according to actual application, and the embodiment does not make specific limitation, such as from top to bottom, from left to right, etc.

[0078] Optionally, the delay is controlled by a micro-ring structure. When the first linear sequence includes a type of information to be processed, the delay period of the micro-ring structure is the same as one symbol period.

[0079] In one specific embodiment, see Figure 5 An image edge extraction experiment was conducted using a single on-chip thin-film lithium niobate microring modulator to verify the convolution processing capability of the proposed scheme, achieving edge extraction of a black and white pixel image. The thin-film lithium niobate microring modulator used in the experiment has a device area of ​​3.4 mm × 0.7 mm and a bandwidth exceeding 67 GHz. To synchronize with the time delay of the feedback loop within the microring modulator, a data sequence with a bit rate of 18.35 Gb / s was selected for demonstration. The experimental results are shown below. Figure 5 As shown, the famous Lenna image was selected and transformed into an image like... Figure 5 (a) shows a 33×33 pixel black and white image as the input matrix, whose gray levels are represented by normalized integers from 0 to 255. Additionally, [the following text is incomplete and likely refers to a different image format] was selected. Figure 5 The 2×2 matrix in the central small image serves as the convolution kernel. Since the photodiode at the receiving end of the experimental system only detects light intensity, for convolution kernels with negative values, it is usually obtained by subtracting two convolution kernels containing only positive values. After convolving the two, the 101 pixels with the highest absolute gray values ​​in the convolution result are extracted and represented as the edges of the image. Figure 5 As shown in (c), the white pixels are the image edge portions extracted after calculation by the experimental system. For comparison, Figure 5 (b) shows the white image edge portion extracted by the digital computer after performing the same operations as the experimental system described above. Figure 5 (d) shows the confusion matrix comparing the digital convolution results with the experimental convolution results. It can be seen that although edge pixels only account for about 10% of all pixels in the entire image, the experimental system can still accurately extract them. The edge pixel matching rate between the experimental system and the digital computer is as high as 89%, demonstrating the feasibility and accuracy of this optical convolution processing accelerator based on a monolithic thin-film lithium niobate microring modulator in image edge extraction.

[0080] Optionally, the delay is controlled by a micro-ring structure. When the first linear sequence includes N types of information to be processed, the delay period of the micro-ring structure contains N symbol periods with the same period, where N is a natural number greater than 1.

[0081] In one specific embodiment, this invention also proposes an in-ring multi-bit parallel modulation scheme based on the convolution operation model of the micro-ring modulator. By increasing the modulation rate without changing the number of devices or peripheral chip equipment, a matrix operation module can be used to simultaneously perform multiplication / convolution operations on multiple matrices, thereby increasing the system's computational density and parallelism. Taking the in-ring 2-bit parallel scheme as an example, using... Figure 5 The flowchart of parallel convolution operation for the 33*33 Lenn black and white pixel image in (a) is as follows: Figure 6 As shown. With Figure 4 Unlike the flowchart in the previous example, the multi-bit parallel scheme within the ring no longer treats the entire micro-ring delay period as a single symbol period. Instead, it incorporates multiple symbol periods into the micro-ring delay period, allowing k different symbols to be transmitted simultaneously within the ring. Since the micro-ring modulator and the internal modulator modulate only once every k symbols in the micro-ring modulator structure, these k symbols are independent of each other, thus achieving parallel modulation. Figure 6 The blue and orange sections represent two sets of data processed in parallel within the ring. Because the modulation rate is twice the original 18.35 Gb / s, they can be computed in parallel without interference within the micro-ring. Therefore, two convolutional kernels can be used simultaneously to extract features from this Lenna image. This multi-bit parallel scheme does not require additional peripheral equipment, so... Figure 6 System architecture diagram of matrix operation module and Figure 4 The same applies. However, because this scheme requires parallel storage and computation of k symbols within the micro-ring, the modulation rate becomes k times the original, and the modulation data also needs to be... Figure 6 The blue and orange modules are interwoven in the same way, and the convolution results of the two convolution kernels will also alternate at the output port.

[0082] The embodiment of the present application only uses a single device to realize the acceleration of convolution operation in the optical domain. The processing time and error of a 33x33 Lenna black and white pixel image are compared with the standard convolution operation result calculated by a digital computer, as shown in Table 1. First, Table 1 uses the experimental system and the convolution operation result obtained after the convolution kernel of the digital computer, and the RMSE of the effective part (since the negative pixels are invalid when drawing and extracting features, the effective part is the part with non-negative results) is used to represent the calculation error. The error of the convolution operation result of the optical matrix module is only 0.126, which shows that the optical matrix operation module in the embodiment of the present application has good accuracy in convolution operation. In addition, the processing time of the two Lenna images is compared. The optical matrix operation module needs to complete (33-1)x(33-1)=1024 four-dimensional vector dot product operations to complete the feature extraction of the Lenna image, and each dot product operation needs at least 4+1=5 symbol periods (an extra period is used to read the result and clear zero). Since the convolution kernel has negative numbers, the convolution results of two positive convolution kernels need to be subtracted to obtain the final convolution result. Therefore, the optical convolution layer with two convolution kernels has a processing time of That is, the feature extraction operation of the Lenna image only needs 0.56μs. After applying the multi-bit parallel scheme, the overall calculation time will also be halved with the number of parallel bits. Taking a 2-bit parallel scheme as an example, since two multiplication and addition operations can be performed at the same time, the processing time of the Lenna image is reduced to 1 / 2, i.e. 0.28μs. The convolution operation in the digital computer is completed by the MATLAB program, and the processor is Intel Core i9-10980 CPU. In order to reduce accidental errors, after using the same convolution kernel to perform one hundred convolution operations on the 33x33 pixel Lenna grayscale image, the average value of the processing time is taken, and the result is 51.575μs. It can be seen that the calculation speed of the convolution operation of the pixel image based on the optical matrix operation module in the embodiment of the present application is nearly 100 times higher than that of the traditional digital computer. The high bandwidth and low delay of the optical convolution accelerator make it have a wide application scenario.

[0083] Table 1

[0084] Evaluation criteria RMSE Processing time (microseconds, ps) Digital computer / 51.575 Optical matrix operation module 0.126 0.56 2-bit scheme optical matrix operation module / 0.28

[0085] Implementing the embodiment of the present application includes the following beneficial effects:

[0086] (1) Since large-scale matrix operations can be completed using only a single device, and vector dot product operations of arbitrary length can be achieved by adjusting the sequence length and micro-ring coupling rate, and extended to matrix multiplication, this device has higher scalability than spatial multiplexing device array schemes where the number of devices increases quadratically with the scale of matrix operations.

[0087] (2) This device can perform matrix multiplication and convolution operations of different scales by simply changing the modulation rate of the data sequence and its mapping method without changing the device structure and peripheral equipment. Its configuration flexibility is far better than other solutions.

[0088] (3) This scheme only uses time-division multiplexed data sequence input to Mach-Zehnder modulator and micro-ring to realize matrix operation. It does not require wavelength or spatial dimension, which reduces the complexity of system equipment design and manufacturing, and also greatly reduces the power consumption and delay of the entire system.

[0089] See Figure 7 This invention provides a control system for an optical convolution accelerator, comprising:

[0090] The first module is used to arrange the information to be processed into a first linear sequence in the time domain according to the processing order, and to use the first linear sequence to modulate the intensity modulator to obtain the light intensity signal;

[0091] The second module is used to arrange the convolution kernel elements into a second linear sequence according to the execution order, and modulate the micro-ring modulator in a time-division multiplexing manner to obtain the modulation sub-signal. The modulation sub-signals with different symbol periods are superimposed by delay control. When the first linear sequence modulation is completed, the superimposed modulation sub-signal is output and the photodetector is controlled to detect the output signal.

[0092] The third module is used to determine the result of the convolution operation based on the output signal.

[0093] It is evident that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0094] See Figure 8 This invention provides a control device for an optical convolution accelerator, comprising:

[0095] At least one processor;

[0096] At least one memory for storing at least one program;

[0097] When at least one program is executed by at least one processor, the at least one processor performs the method described above.

[0098] The memory, as a kind of non-transient computer readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs.The memory can include high-speed random access memory, and can also include non-transient memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transient solid-state memory device.In some embodiments, the memory can optionally include remote memory that is remotely arranged relative to the processor, which can be connected to the processor through a network.Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0099] It can be seen that the contents in the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0100] In addition, the present application also discloses a computer program product or a computer program, which is stored in a computer readable storage medium.The processor of the computer device can read the computer program from the computer readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned method.Similarly, the contents in the above method embodiments are applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0101] The present application also provides a computer readable storage medium, which stores a program executable by a processor, and the program executable by the processor is used to implement the above-mentioned method when executed by the processor.

[0102] It is to be understood that all or some of the steps, systems, etc. in the methods disclosed above can be performed by software, firmware, hardware, and / or any suitable combination thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a micro-processing unit, as hardware, or as an integrated circuit, such as an application- specific integrated circuit. Such software can be distributed on computer readable media, which can comprise computer storage media (or non-transitory media), and communication media (or transitory media). As is well known to those of ordinary skill in the art, computer storage media includes all computer-readable media in which data, computer executable instructions, or other computer readable data is / are publicized, embodied, or otherwise accessed. Computer storage media does not include communication media unless the communication media facilitates access to computer readable data. By way of example, and not limitation, computer storage media can include random- access memory (RAM), read-only memory (ROM), EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage, or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, as is well known to those of ordinary skill in the art, communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" means a signal that has one or more of its characteristics changed or set in a manner so as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as wireless networks, cellular telephone connections, RF data links, Bluetooth connections, and the like.

[0103] The above description is that of the preferred embodiments of the application. Various modifications and changes can be made thereto without departing from the spirit of the application, which is defined by the scope of the following claims.

Claims

1. A method of controlling an optical convolution accelerator, characterized by, include: The information to be processed is arranged into a first linear sequence in the time domain according to the processing order, and the first linear sequence is modulated onto the optical carrier by an intensity modulator to obtain an optical intensity signal that encodes the information to be processed into optical amplitude information. The convolution kernel elements are arranged into a second linear sequence according to the execution order, and modulated by a micro-ring modulator in a time-division multiplexing manner. Within the same symbol period, the values ​​of the weight vector elements at the corresponding positions are sequentially modulated into the optical intensity signal to obtain the modulation sub-signal. The optical amplitude of the feedback loop of the micro-ring is adjusted by adjusting the micro-ring coupling rate, and the modulation sub-signals of different symbol periods are superimposed in time by delay control. When the first linear sequence modulation is completed, the superimposed modulation sub-signal is output, and the photodetector is controlled to detect the output signal. The micro-ring modulator includes a micro-ring structure, and the delay is controlled by the micro-ring structure. When the first linear sequence includes N types of information to be processed, the delay period of the micro-ring structure contains N symbol periods, where N is a natural number greater than 1. The result of the convolution operation is determined based on the output signal.

2. The control method according to claim 1, characterized by, The information to be processed includes image information, and arranging the information to be processed into a first linear sequence in the time domain according to the processing order specifically includes: Convert image information into an input matrix; The input matrix is ​​divided into blocks according to the size of the convolution kernel elements, and the elements in each block are arranged into a first linear sequence in the time domain according to the processing order.

3. The control method according to claim 1, characterized by, The delay is controlled by a micro-ring structure. When the first linear sequence includes a type of information to be processed, the delay period of the micro-ring structure is the same as that of a symbol period.

4. An optical convolution accelerator characterized by, Control is performed using the control method described in any one of claims 1-3, comprising a light source, an intensity modulator, a micro-ring modulator, and a photodetector; wherein, The light source is used to generate optical carrier signals; The intensity modulator is used to modulate the optical carrier signal according to a first linear sequence in the time domain to obtain an optical intensity signal; the first linear sequence carries information to be processed. The micro-ring modulator is used to modulate the light intensity signal by using a second linear sequence of convolution kernel elements in a time-division multiplexing manner to obtain a modulation sub-signal, and to superimpose the modulation sub-signals with different symbol periods by delay to obtain an output signal; the output signal carries the convolution result of the first linear sequence and the second linear sequence; The photodetector is used to detect the output signal.

5. The optical convolution accelerator of claim 1, wherein, The microring modulator includes a first 2×2 multimode interferometer, a first Mach-Zehnder electro-optic intensity modulator, a second 2×2 multimode interferometer, and a microring structure. The first input terminal of the first 2×2 multimode interferometer is connected to the output terminal of the intensity modulator. The two output terminals of the first 2×2 multimode interferometer are respectively connected to the two input terminals of the first Mach-Zehnder electro-optic intensity modulator. The two output terminals of the first Mach-Zehnder electro-optic intensity modulator are respectively connected to the two input terminals of the second 2×2 multimode interferometer. The first output terminal of the second 2×2 multimode interferometer is connected to the photodetector. The second output terminal of the second 2×2 multimode interferometer is connected to the second input terminal of the first 2×2 multimode interferometer through the microring structure.

6. The optical convolution accelerator of claim 1, wherein, The intensity modulator comprises a second Mach-Zehnder electro-optic intensity modulator.

7. A control system for an optical convolution accelerator, characterized by, The method comprises: The first module is configured to arrange the to-be-processed information into a first linear sequence in time domain according to a processing sequence, and modulate the first linear sequence onto an optical carrier through an intensity modulator to obtain an optical intensity signal in which the to-be-processed information is encoded as optical amplitude information. The second module is configured to arrange elements of a convolution kernel into a second linear sequence according to an execution sequence, modulate a micro-ring modulator in a time-division multiplexing manner, modulate values of weight vector elements at corresponding positions into the optical intensity signal in sequence within a same symbol period to obtain a modulation sub-signal, adjust a feedback loop light amplitude of the micro-ring by adjusting a micro-ring coupling rate, and superimpose the modulation sub-signals of different symbol periods in time sequence by delay control; when the first linear sequence is modulated, the superimposed modulation sub-signal is outputted, and an output signal is detected by a photoelectric detector; the micro-ring modulator comprises a micro-ring structure, and the delay is controlled through the micro-ring structure; when the first linear sequence comprises N types of to-be-processed information, a delay period of the micro-ring structure comprises N symbol periods, and N is a natural number greater than 1. The third module is configured to determine a convolution operation result according to the output signal.

8. A control device for an optical convolution accelerator, characterized by The method comprises: At least one processor; At least one memory configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor is caused to implement the method according to any one of claims 1-3.

9. A computer readable storage medium having stored therein a program that is executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is configured to implement the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Photon matrix multiplication device and method for neural network

    CN116484931A

  • Time-multiplexed photonic computer

    US20240061316A1