A neural network computing system and method based on an optical matrix computing chip
Patent Information
- Application Number
- CN202410316791.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-03-19
AI Technical Summary
[0005]鉴于此,本发明实施例提供了一种基于光矩阵计算芯片的神经网络计算系统和方法,以消除光矩阵计算芯片的噪声问题,解决计算精度和计算速度低的问题
[0031]本发明提供一种基于光学矩阵计算芯片的神经网络计算系统和方法,该系统包括:FPGA芯片、处理器模块、光矩阵计算芯片、光调制器、输入数模转换器、输出模数转换器和光电探测器。基于该系统的神经网络计算方法包括:由处理器模块将预训练的神经网络模型中线性层计算部分和卷积层计算部分分别载入FPGA芯片和光矩阵计算芯片中进行计算。FPGA芯片中二值线性计算模块包括多层二值网络线性层和单个全精度线性层。卷积层部分计算在处理器模块转换为多个矩阵块组成的矩阵计算。处理器模块接收二值线性计算模块和光矩阵计算芯片计算后返回的数据进行汇算处理,输出神经网络模型计算的总结果。本发明通过将全精度卷积层与二值线性层融合,提高了神经网络模型与硬件适配程度,能够抵抗光计算芯片的噪声干扰,提高计算精度和计算效率。
Smart Images

Figure CN118446251B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of optical computing and artificial intelligence, and in particular to a neural network computing system and method based on an optical matrix computing chip. Background Technology
[0002] An optical matrix computing chip is a type of hardware capable of performing efficient matrix calculations using optics. Optical chips are constructed using cascaded Mach-Zehnder interferometers. By arranging the Mach-Zehnder interferometers through triangular or square decomposition, they can be used to perform calculations on neural networks. However, optical calculations are susceptible to errors such as noise, requiring compensation and calibration to improve computational accuracy.
[0003] An FPGA (Field Programmable Gate Array) is an integrated circuit chip that can be programmed to perform specific digital circuit functions, playing a crucial role in neural network computation. However, FPGA resources are limited, making it impossible to directly deploy large models. Considering resource constraints, employing multiple computations or specific computational modules would increase computational overhead.
[0004] Binary neural networks can achieve high-speed, low-power neural network computation by binarizing weights and activation values, making them ideal for deployment on edge computing hardware, especially reconfigurable hardware such as FPGAs. Although binary neural networks have the advantages of high computational speed and memory saving, the accuracy of the network model is reduced due to the limitation of its expressive power caused by the oversimplified binary weights. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a neural network computing system and method based on an optical matrix computing chip to eliminate the noise problem of the optical matrix computing chip and solve the problems of low computing accuracy and computing speed.
[0006] One aspect of the present invention provides a neural network computing system based on an optical matrix computing chip, the system comprising:
[0007] The FPGA chip includes a binary linear calculation module, a weighted digital-to-analog converter controller, an input digital-to-analog converter controller, and an output analog-to-digital converter controller; the binary linear calculation module includes a continuous multi-layer binary network linear layer and a single full-precision linear layer;
[0008] The processor module is connected to the input and output terminals of the binary linear calculation module, the weight digital-to-analog converter controller, the input digital-to-analog converter controller, and the output analog-to-digital converter controller. The processor module is used to acquire a pre-trained neural network model, where the linear layers include binary linear layers. It loads the parameters of the linear layers of the neural network model into the binary linear calculation module, transmits the parameters of the convolutional layers in the neural network model to the weight digital-to-analog converter controller, transmits the first data to be processed from the convolutional layers in the neural network model to the input digital-to-analog converter controller, transmits the second data to be processed from the linear layers in the neural network model to the binary linear calculation module for calculation, receives the data returned after calculation by the binary linear calculation module and the optical matrix calculation chip, processes it, and outputs the total result of the neural network model calculation.
[0009] A weighted digital-to-analog converter is used to receive the convolutional layer parameters of the neural network model transmitted by the weighted digital-to-analog converter according to the scheduling instructions of the weighted digital-to-analog converter controller, and convert them into convolutional parameter analog voltages for loading into the optical matrix computing chip;
[0010] An input digital-to-analog converter is used to receive the first data transmitted by the input digital-to-analog converter controller, and convert the first data into a first analog electrical signal according to the scheduling instructions of the input digital-to-analog converter controller;
[0011] An optical modulator is used to receive the first analog electrical signal transmitted by the input digital-to-analog converter and convert the first analog electrical signal into an input optical signal for input to the optical matrix computing chip.
[0012] The optical matrix computing chip is used to load the convolution parameter analog voltage, receive the input optical signal, perform calculations, and output the calculation result optical signal.
[0013] A photodetector is used to receive the optical signal of the calculation result output by the optical matrix calculation chip, and convert it into a second analog electrical signal for output;
[0014] An output analog-to-digital converter is used to receive the second analog electrical signal and convert it into a third digital signal, and output it to the output analog-to-digital converter controller. The output analog-to-digital converter controller schedules the third digital signal back to the processor module.
[0015] In some embodiments of the present invention, the processor module includes a SOC, a CPU, or an MCU.
[0016] In some embodiments of the present invention, the photodetector is a balanced photodetector.
[0017] In some embodiments of the present invention, the optical matrix computing chip includes three optical modules for mapping unitary matrices and diagonal matrices.
[0018] Another aspect of the present invention provides a computation method based on the above-described neural network computation system based on an optical matrix computation chip, the method being executed in a processor module, comprising:
[0019] A pre-trained neural network model is obtained, wherein the linear layers in the neural network model include binary linear layers; the parameters of the neural network model are read, the linear layer parameters are loaded into the binary linear calculation module in the FPGA chip, and the convolutional layer parameters are loaded into the weight digital-to-analog converter controller in the FPGA chip, so as to control the weight digital-to-analog converter to convert the convolutional layer parameters into convolutional parameter analog voltages and load them into the optical matrix calculation chip;
[0020] Read the input data of the neural network model;
[0021] The convolutional layer calculation in the neural network model is converted into a matrix calculation composed of multiple matrix blocks. The first data to be processed after conversion is transmitted to the input digital-to-analog converter controller to control the input digital-to-analog converter to convert the first data into a first analog electrical signal, which is then converted into an input optical signal by the optical modulator to be input into the optical matrix calculation chip for calculation. The calculation result optical signals of all matrix blocks are output in sequence.
[0022] The optical signal receiving the calculation result is converted by a photodetector and an output analog-to-digital converter, and then the complete matrix calculation result is restored by the third digital signal returned by the output analog-to-digital converter controller.
[0023] The second data to be processed in the linear layer of the neural network model is transmitted to the binary linear calculation module for calculation, and the calculation result returned by the binary linear calculation module is received.
[0024] The system receives and processes the data returned by the binary linear calculation module and the optical matrix calculation chip, and outputs the total result calculated by the neural network model.
[0025] In some embodiments of the present invention, the neural network model includes: a full-precision convolutional layer, an input binarization layer, a binary linear layer, and a full-precision linear layer.
[0026] In some embodiments of the present invention, the binary linear layer includes an XNOR linear layer.
[0027] In some embodiments of the present invention, the neural network model includes: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a batch normalization layer, a first XNOR binary linear layer, a second XNOR binary linear layer, and a full-precision linear layer.
[0028] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0029] Another aspect of the present invention provides a computer program product, including a computer program / instructions, characterized in that the computer program / instructions, when executed by a processor, implement the steps of the method described above.
[0030] The beneficial effects of the present invention are at least as follows:
[0031] This invention provides a neural network computing system and method based on an optical matrix computing chip. The system includes an FPGA chip, a processor module, an optical matrix computing chip, an optical modulator, an input digital-to-analog converter, an output analog-to-digital converter, and a photodetector. The neural network computing method based on this system includes: the processor module loading the linear layer calculation part and the convolutional layer calculation part of a pre-trained neural network model into the FPGA chip and the optical matrix computing chip respectively for calculation. The binary linear calculation module in the FPGA chip includes multiple binary network linear layers and a single full-precision linear layer. The convolutional layer calculation is converted into matrix calculation composed of multiple matrix blocks in the processor module. The processor module receives the data returned after calculation by the binary linear calculation module and the optical matrix computing chip, performs summation processing, and outputs the total result of the neural network model calculation. This invention improves the hardware compatibility of the neural network model by fusing the full-precision convolutional layer and the binary linear layer, can resist noise interference from the optical computing chip, and improves calculation accuracy and efficiency.
[0032] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.
[0033] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description
[0034] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, are not intended to limit the scope of the invention. The components in the drawings are not drawn to scale but are merely illustrative of the principles of the invention. For ease of illustration and description of certain parts of the invention, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to the invention. In the drawings:
[0035] Figure 1 This is a structural diagram of a neural network computing system based on an optical matrix computing chip according to an embodiment of the present invention.
[0036] Figure 2 This is a flowchart of the computation method of a neural network computing system based on an optical matrix computing chip according to an embodiment of the present invention.
[0037] Figure 3 This is a structural diagram of a neural network model according to another embodiment of the present invention.
[0038] Figure 4 This is a diagram of the neural network model structure used for testing the MNIST dataset according to an embodiment of the present invention.
[0039] Figure 5 This is a diagram showing the noise interference test results of the optical matrix calculation chip under the MNIST dataset according to an embodiment of the present invention.
[0040] Figure 6 This is a diagram of the neural network model structure used for testing the FashionMNIST dataset according to another embodiment of the present invention.
[0041] Figure 7 This is a graph showing the noise interference test results of the optical matrix calculation chip under the FashionMNIST dataset according to another embodiment of the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0043] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0044] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0045] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.
[0046] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.
[0047] An FPGA (Field Pogrammable Gate Array) is an integrated circuit chip that allows users to reprogram its logic functions and connections according to their needs. The basic structure of an FPGA chip includes programmable input / output units, configurable logic blocks, a digital clock management module, routing resources, and low-level embedded functional units. FPGAs utilize six-bit lookup tables composed of SRAM to implement combinational logic, enabling the implementation of arbitrary digital circuits when resources are sufficient. The logic of an FPGA is implemented by loading programming data into internal static memory cells. The values stored in these memory cells determine the logic functions of the logic units and the connection methods between modules or between modules and I / O, ultimately determining the functions that the FPGA can achieve. FPGAs allow for an unlimited number of programming iterations, offering high flexibility and reconfigurability.
[0048] An optical matrix chip is an integrated circuit chip that utilizes optical principles for computation and processing. It uses optical signals instead of traditional electrical signals to achieve data storage, transmission, and processing at the optical level. Optical matrix chips are used for signal transmission and computational operations within the optical domain. Their working principle involves controlling the amplitude, phase, or frequency of the optical signal to achieve operations similar to logic operations, addition, and multiplication in traditional digital circuits. The advantages of optical matrix chips include high bandwidth, low power consumption, and strong anti-interference capabilities.
[0049] Binary linear networks are a special type of neural network in which the weights and activation values are quantized as binary values (usually +1 or -1). In binary linear networks, the connections between neurons and the computation process are performed in a binary manner, thereby reducing computational complexity and storage requirements. The use of binary computation in binary linear networks can significantly reduce computational complexity, making them suitable for deployment and operation on resource-constrained devices.
[0050] One embodiment of the present invention provides a neural network computing system based on an optical matrix computing chip, the system comprising:
[0051] The FPGA chip includes a binary linear computation module, a weighted digital-to-analog converter controller, an input digital-to-analog converter controller, and an output analog-to-digital converter controller. The binary linear computation module consists of a continuous multi-layer binary network linear layer and a single full-precision linear layer.
[0052] The processor module is connected to the input and output terminals of the binary linear computation module, the weight digital-to-analog converter controller, the input digital-to-analog converter controller, and the output analog-to-digital converter controller. The processor module acquires the pre-trained neural network model, which includes binary linear layers. It loads the linear layer parameters of the neural network model into the binary linear computation module and transmits the convolutional layer parameters to the weight digital-to-analog converter controller. It transmits the first data to be processed from the convolutional layers of the neural network model to the input digital-to-analog converter controller. It transmits the second data to be processed from the linear layers of the neural network model to the binary linear computation module for calculation. It receives the data returned from the binary linear computation module and the optical matrix computation chip after calculation, processes it, and outputs the overall result of the neural network model calculation.
[0053] The weighted digital-to-analog converter is used to receive the convolutional layer parameters of the neural network model transmitted by the weighted digital-to-analog converter controller according to the scheduling instructions of the weighted digital-to-analog converter controller, and convert them into convolutional parameter analog voltages for loading into the optical matrix computing chip.
[0054] An input digital-to-analog converter is used to receive the first data transmitted by the input digital-to-analog converter controller and convert the first data into a first analog electrical signal according to the scheduling instructions of the input digital-to-analog converter controller.
[0055] An optical modulator is used to receive the first analog electrical signal transmitted by the input digital-to-analog converter and convert the first analog electrical signal into an input optical signal input optical matrix computing chip.
[0056] The optical matrix computing chip is used to load convolution parameter analog voltage, receive input optical signals, perform calculations, and output the calculation result optical signal.
[0057] A photodetector is used to receive the optical signal of the calculation result output by the optical matrix calculation chip and convert it into a second analog electrical signal for output.
[0058] The output analog-to-digital converter receives the second analog electrical signal and converts it into a third digital signal, which is then output to the output analog-to-digital converter controller. The output analog-to-digital converter controller schedules the third digital signal back to the processor module.
[0059] The optical matrix computing chip is composed of multiple cascaded Mach-Zehnder interferometers (MZIs). Unitary matrix multiplication is achieved by arranging the MZIs in a triangular or square structure using triangular or square decomposition. Singular value decomposition (SVD) is used to decompose the matrix into the product of two unitary matrices and a diagonal matrix, enabling arbitrary matrix multiplication and completing matrix calculations for neural networks. Optical signals are less susceptible to environmental electromagnetic radiation and induction compared to electrical signals.
[0060] In some embodiments of the present invention, the processor module includes a SOC, a CPU, or an MCU.
[0061] SOC stands for System on Chip, a chip that integrates various functional modules, including processor cores, memory units, input / output interfaces, clock management, analog interfaces, etc., forming a complete computer system. An SOC can contain multiple processor cores, hardware accelerators, peripheral controllers, and various internal buses and communication structures. SOCs are characterized by high integration, low power consumption, and high performance, and can be customized to meet specific application requirements.
[0062] CPU, or Central Processing Unit, is a key component of a computer system, responsible for executing program instructions and processing computational operations.
[0063] MCU, or Microcontroller Unit, is a small computer system chip that integrates a processor core, memory (including flash memory and RAM), input / output interfaces, timers, and various peripheral controllers. Microcontrollers are typically used in embedded systems to perform specific control, monitoring, and processing tasks.
[0064] In some embodiments of the present invention, the photodetector is a balanced photodetector.
[0065] A balanced photodetector is a photodetector based on differential technology, similar in principle to a balanced bridge, which can improve the sensitivity and anti-interference capability of the photodetector. Balanced photodetectors are commonly used in high-speed optical communication, lidar, optical remote sensing, and other fields to improve system performance and reliability.
[0066] In some embodiments of the present invention, the optical matrix computing chip includes three optical modules for mapping unitary matrices and diagonal matrices.
[0067] Specifically, the three optical modules are mapped to three computational matrices, namely the unitary matrix U, the diagonal matrix ∑, and the unitary matrix V. * .
[0068] Another aspect of the present invention provides a computation method based on the above-described neural network computing system based on an optical matrix computing chip, the method being executed in a processor module, including steps S101 to S106:
[0069] Step S101: Obtain the pre-trained neural network model. The linear layers in the neural network model include binary linear layers. Read the parameters of the neural network model, load the linear layer parameters into the binary linear calculation module in the FPGA chip, and load the convolutional layer parameters into the weight digital-to-analog converter controller in the FPGA chip to control the weight digital-to-analog converter to convert the convolutional layer parameters into convolutional parameter analog voltages and load them into the optical matrix calculation chip.
[0070] Specifically, loading the convolution parameter analog voltage into the optical matrix computing chip involves applying the convolution parameter analog voltage to the optical phase shifter within the optical matrix computing chip. The core component of the optical matrix computing chip is multiple Mach-Zehnder interferometers (MZIs), each consisting of two optical beam splitters and two optical phase shifters.
[0071] Step S102: Read the input data of the neural network model.
[0072] Step S103: The convolutional layer calculation in the neural network model is converted into a matrix calculation composed of multiple matrix blocks. The first data to be processed after conversion is transmitted to the input digital-to-analog converter controller to control the input digital-to-analog converter to convert the first data into a first analog electrical signal, which is then converted into an input optical signal by the optical modulator to be input to the optical matrix calculation chip for calculation. The calculation result optical signal of all matrix blocks is output in sequence.
[0073] Step S104: After the optical signal of the received calculation result is converted by the photodetector and the output analog-to-digital converter, the third digital signal returned by the output analog-to-digital converter controller is used to restore the complete matrix calculation result.
[0074] Step S105: Transmit the second data to be processed in the linear layer of the neural network model to the binary linear calculation module for calculation, and receive the calculation result returned by the binary linear calculation module.
[0075] Step S106: Receive the data returned by the binary linear calculation module and the optical matrix calculation chip after calculation, process it, and output the total result of the neural network model calculation.
[0076] The binary linear computation module is used to perform binary computation on the data to be processed in the linear layer. Traditional neural network training methods are difficult to run directly in this system. Therefore, a pre-trained neural network model is used. This model is trained in the computer before being used for computation in this system.
[0077] In some embodiments, the neural network model includes: a full-precision convolutional layer, an input binarization layer, a binary linear layer, and a full-precision linear layer.
[0078] In some embodiments, the binary linear layer includes an XNOR linear layer.
[0079] The XNOR linear layer introduces the XNOR operation into the linear layer of a neural network. XNOR is a bitwise comparison operation between two binary numbers, resulting in a logical value of 1 or 0. In the XNOR linear layer, the input data is typically first converted to binary form through binarization, and then XNOR is performed with the weights to achieve efficient computation.
[0080] In some embodiments, the neural network model includes: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a batch normalization layer, a first XNOR binary linear layer, a second XNOR binary linear layer, and a full-precision linear layer.
[0081] Another embodiment of the present invention provides a neural network computing system based on an optical matrix computing chip, wherein the structure of the neural network model is as follows: Figure 3 As shown, the neural network model includes convolutional layers, input binarization layers, XNOR binary linear layers, and full-precision linear layers. The neural network model in this invention is characterized by the convolutional layers using high-precision weights partially or entirely, while the linear layers, except for the last one which uses full-precision linear layers, consist of XNOR linear layers combined with input binarization layers. The input binary activation can be added or removed based on the accuracy and noise resistance requirements of the actual application.
[0082] In some embodiments, the neural network model may also include other neural network layers such as activation function layers, pooling layers, and / or batch normalization layers to further increase the complexity of the model or simplify the number of parameters, while still obtaining features that resist partial noise, reduce model size, and speed up computation from the high-precision convolution and binarization mixture on which it is based.
[0083] like Figure 1 The diagram shows the architecture of a neural network computing system based on an optical matrix computing chip. This example system uses a ZYNQ series FPGA, specifically the XC7Z100, which includes an armv7 architecture SoC and runs a Linux system. The system's execution methods include:
[0084] Before running neural network inference, the neural network model is first trained on a high-performance computer, and then the model parameters are saved to the SOC's storage module. The SOC reads the model parameters and transmits them to the weight DAC controller and XNOR linear layer through the AXI-FULL bus. The weight DAC controller applies the corresponding optical phase shifter voltage to the phase shifter through the AD5767 DAC, thus completing the model preparation.
[0085] When encountering a convolutional layer, the input data is read into the SOC. The SOC converts the convolutional part into matrix calculations and divides it into matrix blocks of appropriate size for the optical matrix calculation chip. The matrix blocks are then sent to the input DAC controller via the AXI-STREAM bus. The DAC controller controls the output voltage of the AD9767 DAC and controls the optical modulator to perform amplitude and phase modulation. After waiting for the electrical processing delay and the optical processing delay, the data is converted from optical signal to electrical signal by a balanced photodetector, and then read into the ADC controller by an AD9238 ADC. After that, the data is transmitted back to the SOC via the AXI-STREAM bus. This cycle continues until all block matrix transmissions are completed. The SOC then restores the block matrix to complete the matrix calculation, thus completing the convolutional layer calculation.
[0086] When encountering a linear layer, the SOC directly inputs the data to the binary linear layer via AXI-STREAM. After waiting for the required computation delay, the data is returned to the SOC via the AXI-STREAM bus. Due to the hardware resource saving of the XNOR linear layer, all binary linear layers and the last full-precision linear layer can be implemented in the FPGA for computation. In this way, the result can be obtained with only one transmission. By reasonably allocating the ratio of convolutional layers and binary neural network layers, hardware resources can be fully utilized to complete the neural network task at an extremely fast speed.
[0087] This invention integrates the linear layers of a binary network with the convolutional layers of a full-precision network. The binary network linear layers use binarized parameters to reduce memory usage and resist some noise. The full-precision network convolutional layers are no different from ordinary neural networks, ensuring the accuracy of the network. Using an optical matrix computing chip enables faster computation, achieving a balance between accuracy and computational speed. The binarization layer can be directly deployed on an FPGA, and the convolutional layers only require the addition of an ADC (analog-to-digital converter) controller and a DAC (digital-to-analog converter) controller, achieving extremely fast inference speed while consuming only a small amount of hardware resources.
[0088] According to the principles of this invention, two different network structures can be used to apply the MNIST and Fashion-MNIST datasets respectively. The calculation results are then observed through computer simulation to verify the effectiveness of the invention. During the test, Gaussian noise with a standard deviation of 0.001 was applied to the balanced photodetector, a beam splitter with a beam splitting error of 3.54% was applied, and Gaussian noise with a standard deviation of 0-0.05 was applied to the optical phase shifter. The maximum optical matrix calculation chip size during the test was 100 ports.
[0089] In one embodiment of the present invention, the neural network model structure used for the MNIST dataset test is as follows: Figure 4 As shown, Figure 4 This includes a binary linear layer model and a full-precision linear layer model for comparison. Compared to the binary linear layer model, the full-precision linear layer model is completely identical in structure except that all XNOR binary linear layers are replaced with full-precision linear layers and the input binary activation is removed.
[0090] In the absence of noise, the accuracy of the linear layer binary network is 97.23%, and the accuracy of the linear layer full-precision network is 97.73%.
[0091] After adding noise, such as Figure 5 As shown, under Gaussian noise of an optical phase shifter with a standard deviation of 0.03, the accuracy of the linear layer full-precision network is 34.4%, while the linear layer binary network still maintains an accuracy of 94.94%, which is 60.54% higher than that of the linear layer full-precision network. In terms of parameter space, the linear layer full-precision network occupies a total space of 94244 bytes, while the linear layer binary network occupies a total space of 9500 bytes, which is 84744 bytes less than that of the linear layer full-precision network, compressing the parameter space by approximately 9.92 times.
[0092] Under the same linear binary network, the computational speed of convolutional layers is compared between CPU and optical matrix computing chip. The highest operating frequency of existing FPGAs is 1GHz, which uses a 1GHz electro-optic modulator and balanced photodetector for modulation and sampling, along with a 1GHz ADC (analog-to-digital converter) and DAC (digital-to-analog converter). In this example, using a 100-port optical matrix computing chip, a convolutional layer with 1 input channel, 4 output channels, and a kernel size of 2*2 (1, 4, 2*2) requires approximately 27*27 / (100 / 4) ≈ 30 computations. A convolutional layer with 4 input channels, 25 output channels, and a kernel size of 5*5 (4, 25, 5*5) requires 9*9 = 81 computations. The total computation time for 111 operations is 111 ÷ 10^6. 9 =1.11×10- 7Seconds. Using a computer to perform the same amount of computation, a (1, 4, 2*2) convolutional layer requires 27*27*4*4 multiplications and 27*27*4*3 additions; a (4, 25, 5*5) convolutional layer requires 9*9*25*25*4 multiplications and 9*9*25*99 additions, totaling 423,387 calculations. Assuming a computer performs 5 billion calculations per second, the time would be 423,387 ÷ (5*10^5)^25. 9 ) = 8.46 × 10 -5 The speed of calculation for full-precision convolutional layers is significantly increased by approximately 762 times under ideal conditions, as evidenced by the optical matrix computing chip.
[0093] In another embodiment of the invention, the neural network model structure used is tested using the FashionMNIST dataset as follows: Figure 6 As shown. To increase the complexity of the comparison models and enhance the feasibility of the approach, this example increased the model size for testing. The output channels of the second convolutional layer in both models were increased from 25 to 50, and the number of parameters in the linear layers was also increased, reaching five layers. Compared to the binary linear layer model, the full-precision linear layer model added two ReLU activation layers.
[0094] In the absence of noise, the accuracy of the linear layer binary network is 80.61%, and the accuracy of the linear layer full-precision network is 84.46%.
[0095] After adding noise, such as Figure 7 As shown, under Gaussian noise of an optical phase shifter with a standard deviation of 0.03, the linear layer full-precision network achieves an accuracy of 39.98%, while the linear layer binary network achieves an accuracy of 55.76%, which is 15.78% higher than the linear layer full-precision network. In terms of parameter space, the linear layer full-precision network occupies 1,905,864 bytes, while the linear layer binary network occupies only 72,194 bytes, reducing the space occupied by 1,833,670 bytes compared to the linear layer full-precision network, a compression of approximately 26.40 times.
[0096] Under the same linear binary network, the computational speed of convolutional layers is compared on the CPU and optical matrix computing chip. Using the same hardware specifications as the previous example tested with the MNIST dataset, and employing a 100-port optical chip, the (1, 4, 2*2) convolutional layer requires approximately 27*27 / (100 / 4) ≈ 30 optical matrix computing chip operations, and the (4, 25, 5*5) convolutional layer requires 9*9 = 81 optical matrix computing chip operations. The total time for 111 operations is 111 ÷ 10. 9 =1.11×10 -7Seconds. Using a computer to perform the same amount of computation, a (1, 4, 2*2) convolutional layer requires 27*27*4*4 multiplications and 27*27*4*3 additions; a (4, 50, 5*5) convolutional layer requires 9*9*50*25*4 multiplications and 9*9*50*99 additions, totaling 805,950 calculations. Assuming a computer performs 5 billion calculations per second, the time would be 805,950 ÷ (5*10^5)^25. 9 ) = 1.61 × 10 -4 In seconds, under ideal conditions, the speed would be increased by approximately 1450 times.
[0097] In summary, this invention provides a neural network computing system and method based on an optical matrix computing chip. The system includes an FPGA chip, a processor module, an optical matrix computing chip, an optical modulator, an input digital-to-analog converter, an output analog-to-digital converter, and a photodetector. The neural network computing method based on this system includes: the processor module loading the linear layer calculation part and the convolutional layer calculation part of a pre-trained neural network model into the FPGA chip and the optical matrix computing chip respectively for computation. The binary linear calculation module in the FPGA chip includes multiple binary network linear layers and a single full-precision linear layer. The convolutional layer calculation is converted into matrix calculation composed of multiple matrix blocks in the processor module. The processor module receives the data returned after computation by the binary linear calculation module and the optical matrix computing chip, performs summation processing, and outputs the total result of the neural network model computation. This invention improves the hardware compatibility of the neural network model by fusing the full-precision convolutional layer and the binary linear layer, can resist noise interference from the optical computing chip, and improves computational accuracy and efficiency.
[0098] Corresponding to the above method, the present invention also provides an apparatus comprising a computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and the apparatus performing the steps of the method as described above when the computer instructions are executed by the processor.
[0099] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.
[0100] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.
[0101] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0102] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0103] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A neural network computing system based on an optical matrix computing chip, characterized in that, The system includes: The FPGA chip includes a binary linear calculation module, a weighted digital-to-analog converter controller, an input digital-to-analog converter controller, and an output analog-to-digital converter controller; the binary linear calculation module includes a continuous multi-layer binary network linear layer and a single full-precision linear layer; The processor module is connected to the input and output terminals of the binary linear calculation module, the weight digital-to-analog converter controller, the input digital-to-analog converter controller, and the output analog-to-digital converter controller. The processor module is used to acquire a pre-trained neural network model, where the linear layers include binary linear layers. It loads the parameters of the linear layers of the neural network model into the binary linear calculation module, transmits the parameters of the convolutional layers in the neural network model to the weight digital-to-analog converter controller, transmits the first data to be processed from the convolutional layers in the neural network model to the input digital-to-analog converter controller, transmits the second data to be processed from the linear layers in the neural network model to the binary linear calculation module for calculation, receives the data returned after calculation by the binary linear calculation module and the optical matrix calculation chip, processes it, and outputs the total result of the neural network model calculation. The weighted digital-to-analog converter is used to receive the convolutional layer parameters of the neural network model transmitted by the weighted digital-to-analog converter controller according to the scheduling instructions of the weighted digital-to-analog converter controller, and convert them into convolutional parameter analog voltages for loading into the optical matrix computing chip; An input digital-to-analog converter is used to receive the first data transmitted by the input digital-to-analog converter controller, and convert the first data into a first analog electrical signal according to the scheduling instructions of the input digital-to-analog converter controller; An optical modulator is used to receive the first analog electrical signal transmitted by the input digital-to-analog converter and convert the first analog electrical signal into an input optical signal for input to the optical matrix computing chip. The optical matrix computing chip is used to load the convolution parameter analog voltage, receive the input optical signal, perform calculations, and output the calculation result optical signal. A photodetector is used to receive the optical signal of the calculation result output by the optical matrix calculation chip, and convert it into a second analog electrical signal for output; An output analog-to-digital converter is used to receive the second analog electrical signal and convert it into a third digital signal, and output it to the output analog-to-digital converter controller. The output analog-to-digital converter controller schedules the third digital signal back to the processor module.
2. The neural network computing system based on an optical matrix computing chip according to claim 1, characterized in that, The processor module includes a SOC, CPU, or MCU.
3. The neural network computing system based on an optical matrix computing chip according to claim 1, characterized in that, The photodetector is a balanced photodetector.
4. The neural network computing system based on an optical matrix computing chip according to claim 1, characterized in that, The optical matrix computing chip includes three optical modules for mapping unitary matrices and diagonal matrices.
5. A computational method for a neural network computation system based on an optical matrix computation chip as described in claim 1, characterized in that, This method is used to execute in the processor module and includes: A pre-trained neural network model is obtained, wherein the linear layers in the neural network model include binary linear layers; the parameters of the neural network model are read, the linear layer parameters are loaded into the binary linear calculation module in the FPGA chip, and the convolutional layer parameters are loaded into the weight digital-to-analog converter controller in the FPGA chip, so as to control the weight digital-to-analog converter to convert the convolutional layer parameters into convolutional parameter analog voltages and load them into the optical matrix calculation chip; Read the input data of the neural network model; The convolutional layer calculation in the neural network model is converted into a matrix calculation composed of multiple matrix blocks. The first data to be processed after conversion is transmitted to the input digital-to-analog converter controller to control the input digital-to-analog converter to convert the first data into a first analog electrical signal, which is then converted into an input optical signal by the optical modulator to be input into the optical matrix calculation chip for calculation. The calculation result optical signals of all matrix blocks are output in sequence. The optical signal receiving the calculation result is converted by a photodetector and an output analog-to-digital converter, and then the complete matrix calculation result is restored by the third digital signal returned by the output analog-to-digital converter controller. The second data to be processed in the linear layer of the neural network model is transmitted to the binary linear calculation module for calculation, and the calculation result returned by the binary linear calculation module is received. The system receives and processes the data returned by the binary linear calculation module and the optical matrix calculation chip, and outputs the total result calculated by the neural network model.
6. The calculation method of the neural network computing system based on the optical matrix computing chip according to claim 5, characterized in that, The neural network model includes: a full-precision convolutional layer, an input binarization layer, a binary linear layer, and a full-precision linear layer.
7. The calculation method of the neural network computing system based on the optical matrix computing chip according to claim 6, characterized in that, The binary linear layer includes an XNOR linear layer.
8. The calculation method of the neural network computing system based on the optical matrix computing chip according to claim 5, characterized in that, The neural network model includes: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a batch normalization layer, a first XNOR binary linear layer, a second XNOR binary linear layer, and a full-precision linear layer.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 5 to 8.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Convolution calculation method and convolution operation circuit
CN111324858A
Convolution calculation method and device based on photon calculation chip
CN114723019A