Register component for optoelectronic hybrid neural network computing architecture and processor
Patent Information
- Application Number
- CN202610677123.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-21
AI Technical Summary
[0005]本申请提供一种用于光电混合神经网络计算架构的寄存器组件及处理器,以解决现有光电混合计算中光电转换的时序协同差、转换精度低、数据冲突多、适配性不足等问题
[0016]由上述内容可知,本申请提供一种用于光电混合神经网络计算架构的寄存器组件及处理器,所述寄存器组件包括Mod寄存器模块,所述Mod寄存器模块被配置为读取前一层神经网络输出的所述特征图数据,对所述特征图数据进行处理,并输出处理后的所述特征图数据;PS寄存器模块,所述PS寄存器模块被配置为读取所述特征图数据对应的所述权重数据,对所述权重数据进行处理,并输出处理后的所述权重数据;PD寄存器模块,所述PD寄存器模块被配置为接收MZI阵列光芯片输出的所述结算结果,并将所述结算结果向外输出。本申请通过上述方案解决了现有光电混合计算中光电转换的时序协同差、转换精度低、数据冲突多、适配性不足等问题。
Smart Images

Figure CN122616626A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of neural network technology, and in particular to a register component and processor for a hybrid optoelectronic neural network computing architecture. Background Technology
[0002] With the rapid development of artificial intelligence and deep learning technologies, the scale and complexity of neural network models continue to grow, placing unprecedented demands on the performance of computing hardware. Against this backdrop, achieving high-density, large-scale neural network deployments faces significant challenges, with the inherent computationally intensive nature of linear layers (such as convolutional layers) becoming one of the core bottlenecks for hardware acceleration. Traditional neural network processors based on electronic computing, such as graphics processing units and tensor processing units, while effectively supporting model training and inference to a certain extent, are still constrained by the inherent limitations of the von Neumann architecture. The memory wall problem inherent in this architecture leads to frequent data transfer between memory and processor, becoming a major constraint on performance and energy efficiency, significantly hindering improvements in computational throughput and optimization of energy efficiency.
[0003] To overcome the physical limitations of electronic computing, optical computing technology emerged. Leveraging the inherent advantages of photon propagation—such as ultra-high bandwidth, ultra-fast transmission, and extremely low power consumption and latency—it provides a new technological path for highly parallel linear computation. Optical computing chips have shown outstanding potential in linear operations such as matrix-vector multiplication, significantly accelerating the dense linear computation portions of neural networks. However, optical signals are relatively weak in nonlinear characteristics, and existing optical chips are still in the early research and laboratory verification stages in realizing nonlinear computation functions; simultaneously, effective storage of optical signals also faces technical challenges. Therefore, current technological approaches generally tend to adopt a hybrid optoelectronic computing architecture, utilizing the optical domain for highly parallel linear computation, while the electrical domain handles nonlinear activation function processing, data storage, and overall system control, thereby achieving the effect of electrically controlled light and collaborative computing.
[0004] Although the optoelectronic hybrid computing architecture conceptually integrates the advantages of both light and electricity, a series of key technical deficiencies still exist in practical engineering implementation, especially in the optoelectronic conversion and coordination stages, which restrict the full realization of the overall system performance. First, in terms of timing coordination, existing solutions lack precise timing synchronization and control mechanisms. The working rhythm and data flow between the optical computing unit and the electronic computing unit often fail to coordinate effectively, leading to asynchrony between key steps such as feature map data transmission, weight parameter configuration, and calculation result reception. This timing mismatch easily triggers data conflicts and pipeline stalls, severely hindering the execution efficiency of computing tasks. Second, in terms of optoelectronic conversion interface design, the adaptability of existing conversion devices is insufficient. Feature map data typically requires high-speed, streaming transmission to match the high throughput characteristics of optical computing, while neural network weight data often requires high-precision, stable configuration to ensure computational accuracy. Current single conversion rate or interface design cannot simultaneously accommodate the different speed and accuracy requirements of these two types of data, resulting in compromises in interface bandwidth or configuration accuracy, thereby limiting the overall system performance. Furthermore, in terms of data scheduling and management, existing technologies suffer from low efficiency in data path and cache resource utilization. Due to the lack of effective cache isolation and scheduling strategies, frequent access conflicts occur during data reading, writing, and transfer, increasing unnecessary waiting time and power consumption. At the same time, the cache hierarchy is not fully utilized to reduce the number of data transfers, resulting in the system's effective throughput failing to meet theoretical expectations. Summary of the Invention
[0005] This application provides a register component and processor for a hybrid optoelectronic neural network computing architecture to solve problems such as poor timing coordination, low conversion accuracy, numerous data conflicts, and insufficient adaptability in existing optoelectronic hybrid computing.
[0006] In a first aspect, this application provides a register component for a hybrid optoelectronic neural network computing architecture, the register component comprising: The Mod register module is configured to read the feature map data output from the previous layer of the neural network, process the feature map data, and output the processed feature map data. The PS register module is configured to read the weight data corresponding to the feature map data, process the weight data, and output the processed weight data. The PD register module is configured to receive the settlement result output by the MZI array optical chip and output the settlement result to the outside.
[0007] Preferably, the Mod register module is further configured as follows: The feature map data is processed differently based on the values of different bits and then output.
[0008] Preferably, the PS register module is further configured as follows: The weight data is controlled and processed according to the values of different bits and then output to the outside.
[0009] Preferably, the PS register module is further configured as follows: The Mod register module and the PD register module are controlled to determine the timing relationship between the processed feature map data and the weight data arriving at the MZI array optical chip.
[0010] Secondly, this application also provides a processor for a hybrid optoelectronic neural network computing architecture, the processor comprising any of the register components described above, the processor comprising: The MZI array optical chip is configured to perform calculations based on the feature map data and weight data input by the register component, and output the calculation results to the outside via the register component. Off-chip DRAM, which is configured to store feature map data to be calculated, weight data to be calculated, and the calculation results output by the MZI array optical chip; A DMA controller configured to control and access the off-chip DRAM.
[0011] Preferably, the processor further includes: A signal conversion module is disposed between the register component and the MZI array optical chip, and the signal conversion module is configured to perform signal mode conversion processing on the input data.
[0012] Preferably, the signal conversion module is further configured to: The feature map data and weight data input to the register component are subjected to signal mode conversion, and the converted data is transmitted to the MZI array optical chip. The calculation results output by the MZI array optical chip are converted into signal modes, and the converted calculation results are transmitted to the register component.
[0013] Preferably, the signal conversion module includes: A high-speed digital-to-analog / analog-to-digital converter unit, the high-speed digital-to-analog / analog-to-digital converter unit being configured to perform signal mode conversion on the feature map data and / or the calculation results at a first frequency; A low-speed digital-to-analog converter is configured to perform signal mode conversion on the weighted data according to a second frequency; the frequency value of the first frequency is greater than the frequency value of the second frequency.
[0014] Preferably, the processor further includes: A buffer manager, comprising BANK A and BANK B, is configured to perform a ping-pong operation on the data calculation results by switching between BANKs.
[0015] Preferably, the processor further includes: A calculation unit is configured to perform linear calculations based on data input from the buffer manager and temporarily store the calculated data. A nonlinear module, configured to perform nonlinear calculations based on data input from the buffer manager; The processor is also configured to: The top-level control selects the computing unit, the nonlinear module, or the MZI array optical chip to perform relevant calculations.
[0016] As described above, this application provides a register component and processor for a hybrid optoelectronic neural network computing architecture. The register component includes a Mod register module configured to read the feature map data output from the previous layer of the neural network, process the feature map data, and output the processed feature map data; a PS register module configured to read the weight data corresponding to the feature map data, process the weight data, and output the processed weight data; and a PD register module configured to receive the settlement result output from the MZI array optical chip and output the settlement result. This application solves the problems of poor timing coordination, low conversion accuracy, numerous data conflicts, and insufficient adaptability in existing optoelectronic hybrid computing through the above solution. Attached Figure Description
[0017] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a register component for a hybrid optoelectronic neural network computing architecture according to this application; Figure 2 This is a schematic diagram of a processor for an optoelectronic hybrid neural network computing architecture according to this application. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0020] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0021] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0022] Figure 1 This is a schematic diagram of a register component for a hybrid optoelectronic neural network computing architecture according to this application.
[0023] Figure 2 This is a schematic diagram of a processor for an optoelectronic hybrid neural network computing architecture according to this application.
[0024] See Figure 1 as well as Figure 2 As can be seen, this embodiment provides a register component and processor for a hybrid optoelectronic neural network computing architecture. This embodiment uses the overall architecture of the register component and the processor as an example for technical demonstration. Specifically, in this embodiment, this technology is based on a traditional neural network processor architecture, constructing a photoelectric hybrid neural network processor architecture, and achieving precise control of photoelectric conversion through the design of a dedicated control register module. For example... Figure 1As shown, this optoelectronic hybrid neural network processor controls the entire hybrid system through instructions. First, the entire system is initialized, and then the feature map data and instruction data are cached in the on-chip SRAM of the DMA controller. Second, the PS (Phase Shifter) register module decompresses the weight data and configures the refractive index of the MZI phase shifter to establish a preset matrix operation relationship. At the same time, the Mod (Modulator) register module reads the feature map data from the BANK in the buffer manager, performs DC or AC processing on the signal, and transmits it to the MZI array optical chip through a high-speed digital-to-analog / analog-to-digital converter (high-speed DAC). Finally, the PD (Photodetector) register module further processes the calculation results after the MZI array optical chip (Mach-Zehnder Interferometer, MZI) calculation and writes them into the BANK cache in the buffer manager for the next data processing or peripheral reading.
[0025] in: MZI array optical chip: Based on the traditional CPU+NPU heterogeneous system, the original NPU computing tasks are decomposed. The MZI array optical chip (optical domain) is responsible for linear calculations such as convolutional layers and fully connected layers; the electrical domain is responsible for nonlinear calculations such as activation functions, pooling layers, and normalization.
[0026] Dedicated register module: Design three digital control registers, Mod, PD and PS, which are responsible for feature map data processing and transmission, calculation result reception and processing and writing to the buffer BANK, weight configuration and timing scheduling functions, respectively.
[0027] Dual-buffered design: Dual-buffered SRAM is integrated on-chip, supporting ping-pong operation to achieve pipelined parallelism of data loading and conversion, avoiding data conflicts.
[0028] This embodiment aims to solve the problems of poor timing coordination, low conversion accuracy, numerous data conflicts, and insufficient adaptability in existing optoelectronic hybrid computing. It provides an optoelectronic conversion control technology with precise control, efficient coordination, and strong anti-interference capability, realizing seamless connection between linear optical domain computing and nonlinear electrical domain computing, and improving the throughput, energy efficiency ratio, and stability of neural network computing.
[0029] Furthermore, in some embodiments, a fine-grained register design is also provided: 1. Mod register module (16-bit), as shown in Table 1.
[0030] Core function: Controls feature map data to be transmitted to the optical chip via a high-speed DAC, supporting multi-mode data encoding; Key fields: ① Idle status bit: Indicates the modulator's operating status (idle / running); ②On-chip cache BANK selection bit: Switches the dual-buffered BANK data source to achieve ping-pong operation; ③ Trigger control field: Supports two modes: "Instant trigger" and "PS-configured delay trigger", adapting to different computing scenarios; ④ Modulation mode field: Supports DC mode (direct data transmission) and single-cycle, double-cycle, and triple-cycle AC modes (output alternating a and -a sequences), reducing the impact of thermal effects on optical devices.
[0031] Table 1 Operating parameters of the Mod register module
[0032] 2. PD register module (16-bit), as shown in Table 2.
[0033] Core function: Controls the optical chip output data to be written back to the NPU buffer via a high-speed ADC, supporting data noise reduction processing; Key fields: ① Idle status bit: Indicates the working status of the photodetector (idle / operating); ②BANK selection bit: Mutually exclusive with the BANK selection bit of the Mod module to avoid data conflicts; ③ Trigger control field: Works in sync with the Mod module's trigger control timing to ensure data synchronization; ④ Detection mode field: Supports DC mode (directly stores raw data) and AC differential processing mode (subtracts the positive and negative half-cycle sampled values to suppress common-mode noise); ⑤ Data length selection bit: Supports "match to Mod data length" or "full BANK storage" to adapt to different data volume requirements.
[0034] Table 2 Operating parameters of PD register module
[0035]
[0036] 3. PS register module (16-bit), as shown in Table 3.
[0037] Core functions: Controlling the transmission and timing of weighted data, and coordinating the work of Mod and PD modules; Key fields: ① Idle status bit: Indicates the operating status of the phase modulator (idle / running); ② Reuse control bits: Support weight reuse (avoiding duplicate configuration) and updates (loading new weights); ③ Zeroing control bit: After configuration, the voltage can be selected to be zero (to reduce static power consumption) or kept (to stabilize the optical path state). ④ Time unit field: Supports timing configuration at the millisecond, microsecond, nanosecond, and picosecond levels to adapt to the response speed of different optical devices; ⑤ Duration field: Controls the duration of voltage application for weight configuration, adapting to the characteristic requirements of optical devices such as phase change materials.
[0038] Table 3 PS Register Module Operating Parameters
[0039]
[0040] This embodiment has the following advantages: Precise timing coordination: Through the fine-grained timing configuration of dedicated registers, seamless integration of photoelectric conversion and computation is achieved, avoiding pipeline interruptions and improving system throughput; Excellent conversion efficiency: The dual-rate conversion units are adapted to the different needs of feature map data and weight data respectively. High-speed conversion supports GHz-level transmission, and high-precision low-speed DAC conversion ensures the accuracy of weight configuration. Strong anti-interference capability: AC coding and differential processing mechanisms effectively suppress thermal effects and common-mode noise, which can effectively reduce conversion errors; Wide adaptability: Supports multiple modulation modes, cache configurations and timing granularities, adapting to different convolutional kernel sizes, neural network structures and optical chip types; Outstanding energy efficiency: Compared with traditional pure electric computing platforms, the energy consumption of the photoelectric conversion process is significantly reduced. Combined with the high parallelism advantage of optical computing, the overall system energy efficiency ratio is improved by orders of magnitude.
[0041] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the contents of this disclosure, thereby enabling those skilled in the art to better utilize the embodiments.
Claims
1. A register component for a hybrid optoelectronic neural network computing architecture, wherein the architecture is an on-chip in-memory computing architecture, characterized in that, The register component includes: The Mod register module is configured to read the feature map data output from the previous layer of the neural network, process the feature map data, and output the processed feature map data. The PS register module is configured to read the weight data corresponding to the feature map data, process the weight data, and output the processed weight data. The PD register module is configured to receive the settlement result output by the MZI array optical chip and output the settlement result to the outside.
2. The register component for a hybrid optoelectronic neural network computing architecture according to claim 1, characterized in that, The Mod register module is also configured to: The feature map data is processed differently based on the values of different bits and then output.
3. A register component for a hybrid optoelectronic neural network computing architecture according to claim 1, characterized in that, The PS register module is also configured to: The weight data is controlled and processed according to the values of different bits and then output to the outside.
4. A register component for a hybrid optoelectronic neural network computing architecture according to claim 1, characterized in that, The PS register module is also configured to: The Mod register module and the PD register module are controlled to determine the timing relationship between the processed feature map data and the weight data arriving at the MZI array optical chip.
5. A processor for a hybrid optoelectronic neural network computing architecture, the processor comprising the register component of any one of claims 1 to 4, characterized in that, The processor includes: The MZI array optical chip is configured to perform calculations based on the feature map data and weight data input by the register component, and output the calculation results to the outside via the register component. Off-chip DRAM, which is configured to store feature map data to be calculated, weight data to be calculated, and the calculation results output by the MZI array optical chip; A DMA controller configured to control and access the off-chip DRAM.
6. A processor for a hybrid optoelectronic neural network computing architecture according to claim 5, characterized in that, The processor also includes: A signal conversion module is disposed between the register component and the MZI array optical chip, and the signal conversion module is configured to perform signal mode conversion processing on the input data.
7. A processor for a hybrid optoelectronic neural network computing architecture according to claim 6, characterized in that, The signal conversion module is also configured to: The feature map data and weight data input to the register component are subjected to signal mode conversion, and the converted data is transmitted to the MZI array optical chip. The calculation results output by the MZI array optical chip are converted into signal modes, and the converted calculation results are transmitted to the register component.
8. A processor for a hybrid optoelectronic neural network computing architecture according to claim 7, characterized in that, The signal conversion module includes: A high-speed digital-to-analog / analog-to-digital converter unit, the high-speed digital-to-analog / analog-to-digital converter unit being configured to perform signal mode conversion on the feature map data and / or the calculation results at a first frequency; A low-speed digital-to-analog converter is configured to perform signal mode conversion on the weighted data according to a second frequency; the frequency value of the first frequency is greater than the frequency value of the second frequency.
9. A processor for a hybrid optoelectronic neural network computing architecture according to claim 5, characterized in that, The processor also includes: A buffer manager, comprising BANK A and BANK B, is configured to perform a ping-pong operation on the data calculation results by switching between BANKs.
10. A processor for a hybrid optoelectronic neural network computing architecture according to claim 9, characterized in that, The processor also includes: A calculation unit is configured to perform linear calculations based on data input from the buffer manager and temporarily store the calculated data. A nonlinear module, configured to perform nonlinear calculations based on data input from the buffer manager; The processor is also configured to: The top-level control selects the computing unit, the nonlinear module, or the MZI array optical chip to perform relevant calculations.