Online training system and method for neuromorphic hardware array

By coordinating the fully parallel bidirectional drive architecture and the field-programmable gate array control unit, efficient online training of neuromorphic hardware arrays was achieved, solving the parallel drive and synchronization problems of existing systems and improving the current signal acquisition speed and dynamic range measurement capability.

CN122290679APending Publication Date: 2026-06-26QIANYUAN NATIONAL LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QIANYUAN NATIONAL LABORATORY
Filing Date
2026-04-02
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing neuromorphic hardware array testing and training systems cannot achieve fully parallel voltage driving and weight updates, have poor timing synchronization, and cannot meet the nanosecond-level synchronization requirements of spiking neural networks. Furthermore, their dynamic range is limited, making it impossible to take into account real-time measurements over a wide current range.

Method used

Employing a fully parallel bidirectional drive architecture, the multi-channel digital-to-analog converter chip provides independent programmable bias voltages or programming pulses to the column drains of the neuromorphic hardware array. Combined with the field-programmable gate array control unit, nanosecond-level synchronization timing is generated, and current signals are converted in parallel to realize parallel combinational multiplication operations and synchronous weight updates.

Benefits of technology

It enables efficient online training of neuromorphic hardware arrays, meets the stringent requirements of spiking neural networks for nanosecond-level synchronization, and improves the measurement capabilities of current signal acquisition speed and dynamic range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290679A_ABST
    Figure CN122290679A_ABST
Patent Text Reader

Abstract

This application discloses an online training system and method for neuromorphic hardware arrays. The system includes: a chip carrier board for carrying the neuromorphic hardware array under test and providing lead-out interfaces for all its electrodes; a test motherboard including: a voltage drive assembly for simultaneously applying independently programmable bias voltages or programming pulses to the drains of each column; a gate gating unit for selecting one or more rows of gates to activate the memory cells of the corresponding row to participate in forward computation or weight update; a multi-channel parallel signal readout chain for dividing the current signals output from the sources of each row into multiple groups, aggregating the current signals in each group into one analog signal, and converting the multiple analog signals into digital current signals in parallel; and a field-programmable gate array control unit for generating synchronization timing to control the voltage drive assembly to apply bias voltages or programming pulses, control the gate gating unit to select a specified row, and control the multi-channel parallel signal readout chain to acquire digital current signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semiconductor testing technology, and in particular to an online training system and method for neuromorphic hardware arrays. Background Technology

[0002] With the rapid development of artificial intelligence technology, neuromorphic computing has become an important development direction in the post-Moore's Law era due to its low power consumption and high parallelism. Among them, the "sensor-memory-computing integration" architecture based on neuromorphic hardware such as memristor arrays and optoelectronic sensor arrays can realize core neural network operations such as vector matrix multiplication at the physical level, providing a brand-new hardware platform for edge computing and online learning. This type of hardware is usually composed of large-scale cross arrays. The conductance value of the storage unit (such as memristor) at each cross point represents the weight of the neural network. Forward computation can be completed by applying voltage to the column terminals of the array and reading current from the row terminals, while online updates of the weights can be achieved by applying programming pulses.

[0003] However, existing testing and training systems mainly use the readout architecture of traditional CMOS image sensors, whose core mechanism relies on a serial scanning method of "row gating and column-level readout". At the hardware implementation level, such systems sequentially activate specific rows through row decoders, transmit the current or charge of the selected rows to a shared column bus, and then the column-level readout circuit completes the quantization. Although this architecture performs well in the field of standardized imaging, it is difficult to meet the requirements of online training of neuromorphic hardware: First, it cannot achieve fully parallel voltage driving and weight updates. The serial scanning method cannot provide independent "write" bias voltages for hundreds or thousands of columns in the array at the same time, resulting in the inability to physically implement efficient parallel combination multiplication and weight updates; Second, the timing synchronization is poor. The traditional rolling shutter scanning method causes significant delays in the time axis of each unit in the array, which cannot meet the stringent requirements of spiking neural networks for nanosecond-level photoelectric pulse synchronization; Third, the dynamic range is limited. Traditional readout circuits are usually optimized for specific integration times, making it difficult to take into account the wide dynamic range of real-time measurement from femtoampere-level dark current to milliampere-level conduction current. Summary of the Invention

[0004] In view of this, this application provides an online training system and method, storage medium, and computer device for neuromorphic hardware arrays. It achieves efficient online training of the neuromorphic hardware array under test through a fully parallel bidirectional drive architecture: It employs a high-density fully parallel voltage drive combination composed of multiple multi-channel digital-to-analog converter chips, with each output channel connected one-to-one to a column of drains in the neuromorphic hardware array under test. This allows for the simultaneous application of independently programmable bias voltages or programming pulses to all columns of drains, thereby directly supporting parallel combined multiplication operations and synchronous weight updates at the physical level, completely overcoming the limitations of traditional serial scanning. The limitations of column-level drive capability are described; at the same time, the field-programmable gate array control unit is used to generate nanosecond-level precision synchronization timing, strictly coordinating voltage drive, row gating and current acquisition, and establishing a "gating and acquisition" timing relationship in the forward computation stage, ensuring that the excitation and response of each storage unit of the neuromorphic hardware array under test remain highly consistent on the time axis, meeting the stringent requirements of spiking neural networks for nanosecond-level synchronization; in addition, the current signals output from each row source are grouped, aggregated and converted into digital current signals in parallel through a multi-channel parallel signal readout chain, which can strictly guarantee the acquisition speed of digital current signals.

[0005] According to one aspect of this application, an online training system for neuromorphic hardware arrays is provided, comprising: A chip carrier board is used to carry the neuromorphic hardware array under test and to provide lead-out interfaces for all electrodes of the neuromorphic hardware array under test, wherein the electrodes include drain, gate and source. The test motherboard, electrically connected to the chip carrier board, includes: The voltage-driven assembly consists of multiple multi-channel digital-to-analog converter chips. Each output channel is connected one-to-one with a column of drains of the neuromorphic hardware array under test. Under the control of the field-programmable gate array control unit, it is used to simultaneously apply independently programmable bias voltages or programming pulses to each column of drains to realize forward computation or weight updates during online training. The gate selection unit is connected to each row of gates of the neuromorphic hardware array under test, and is used to select one or more rows of gates under the control of the field programmable gate array control unit to activate the memory cells of the corresponding row to participate in forward calculation or weight update. A multi-channel parallel signal readout chain, whose input is connected to each row of the source of the neuromorphic hardware array under test, is used to divide the current signals output from each row of the source into multiple groups under the control of the field-programmable gate array control unit, converge the current signals in each group into one analog signal, and convert the multiple analog signals into digital current signals in parallel, so that the field-programmable gate array control unit can calculate the prediction error and determine the weight update parameters. The field-programmable gate array control unit is connected to the voltage drive combination, the gate gating unit, and the multi-channel parallel signal readout chain, respectively, to generate synchronous timing. During the forward computation phase, it controls the voltage drive combination to apply a bias voltage, controls the gate gating unit to select a specified row, and controls the multi-channel parallel signal readout chain to acquire digital current signals after selection. During the weight update phase, it controls the gate gating unit to select a specified row and controls the voltage drive combination to apply programming pulses.

[0006] According to another aspect of this application, an online training method for neuromorphic hardware arrays is provided, comprising: The field-programmable gate array control unit generates an input voltage vector based on training samples. It controls the voltage drive combination to write the voltage corresponding to each element in the input voltage vector into the register of each corresponding digital-to-analog converter channel in the voltage drive combination through a parallel bus, and triggers all digital-to-analog converter channels to output voltage simultaneously, so as to apply the bias voltage corresponding to the input voltage vector to each column drain of the neuromorphic hardware array under test. The field-programmable gate array control unit controls the gate selection unit to select the gate of a specified row of the neuromorphic hardware array under test, so as to activate the memory cell of the corresponding row, and the activated memory cell outputs an analog current signal based on Ohm's law according to the bias voltage applied to the column and its own conductance value. After the designated row gate is selected, the field-programmable gate array control unit triggers multiple analog-to-digital converters in the multi-channel parallel signal readout chain to perform analog-to-digital conversion operations in parallel. The multi-channel parallel signal readout chain groups and aggregates the analog current signals output by each memory cell in the specified row gate into multiple analog buses, and converts the analog current signals of each analog bus into digital current signals in parallel, which serve as the actual output of the training samples. The field-programmable gate array control unit compares the actual output with the expected output of the training sample, calculates the error gradient corresponding to each storage unit in the neuromorphic hardware array under test based on the comparison result, and determines the weight update parameters of each storage unit according to the error gradient. The field-programmable gate array (FPGA) control unit updates the parameters according to the weights, writes the corresponding programming pulse parameters into the registers of the corresponding digital-to-analog converter (DAC) channels in the voltage drive combination via a parallel bus, and triggers all DAC channels to output programming pulses simultaneously. Simultaneously, it controls the gate selection unit to select the gate of the specified row, applying the programming pulse to the corresponding row's memory cell to update the conductance value of each memory cell. Then, it returns to the step where the FPGA control unit controls the gate selection unit to select the gate of the specified row of the neuromorphic hardware array under test, until the training of the neuromorphic hardware array under test meets the preset convergence condition, resulting in a trained neuromorphic hardware array under test.

[0007] According to another aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described online training method for neuromorphic hardware arrays.

[0008] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described online training method for neuromorphic hardware arrays.

[0009] By employing the above technical solutions, this application provides an online training system and method, storage medium, and computer device for neuromorphic hardware arrays. Through a fully parallel bidirectional drive architecture, it achieves efficient online training of the neuromorphic hardware array under test. It utilizes a high-density, fully parallel voltage drive combination composed of multiple multi-channel digital-to-analog converter chips. Each output channel is connected one-to-one with a column of drains in the neuromorphic hardware array under test, enabling the simultaneous application of independently programmable bias voltages or programming pulses to all column drains. This directly supports parallel combined multiplication operations and synchronous weight updates at the physical level, completely overcoming the limitations of traditional serial multiplication. Row scanning limits column-level driving capabilities; meanwhile, by using a field-programmable gate array (FPGA) control unit to generate nanosecond-level precision synchronization timing, voltage driving, row gating, and current acquisition are strictly coordinated, and a "gating-after-acquisition" timing relationship is established in the forward computation stage, ensuring that the excitation and response of each storage unit of the neuromorphic hardware array under test remain highly consistent on the time axis, meeting the stringent requirements of spiking neural networks for nanosecond-level synchronization; in addition, by using a multi-channel parallel signal readout chain to group and aggregate the current signals output from each row source and convert them into digital current signals in parallel, the acquisition speed of digital current signals can be strictly guaranteed.

[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This illustration shows a schematic diagram of the structure of an online training system for neuromorphic hardware arrays provided in an embodiment of this application; Figure 2 A flowchart illustrating an online training method for neuromorphic hardware arrays provided in an embodiment of this application is shown. Figure 3 A flowchart illustrating another online training method for neuromorphic hardware arrays provided in an embodiment of this application is shown. Figure 4 A schematic diagram of the device structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation

[0012] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0013] This embodiment provides an online training system for neuromorphic hardware arrays, such as... Figure 1 As shown, the system includes: A chip carrier board is used to carry the neuromorphic hardware array under test and to provide lead-out interfaces for all electrodes of the neuromorphic hardware array under test, wherein the electrodes include drain, gate and source. The test motherboard, electrically connected to the chip carrier board, includes: The voltage-driven assembly consists of multiple multi-channel digital-to-analog converter chips. Each output channel is connected one-to-one with a column of drains of the neuromorphic hardware array under test. Under the control of the field-programmable gate array control unit, it is used to simultaneously apply independently programmable bias voltages or programming pulses to each column of drains to realize forward computation or weight updates during online training. The gate selection unit is connected to each row of gates of the neuromorphic hardware array under test, and is used to select one or more rows of gates under the control of the field programmable gate array control unit to activate the memory cells of the corresponding row to participate in forward calculation or weight update. A multi-channel parallel signal readout chain, whose input is connected to each row of the source of the neuromorphic hardware array under test, is used to divide the current signals output from each row of the source into multiple groups under the control of the field-programmable gate array control unit, converge the current signals in each group into one analog signal, and convert the multiple analog signals into digital current signals in parallel, so that the field-programmable gate array control unit can calculate the prediction error and determine the weight update parameters. The field-programmable gate array control unit is connected to the voltage drive combination, the gate gating unit, and the multi-channel parallel signal readout chain, respectively, to generate synchronous timing. During the forward computation phase, it controls the voltage drive combination to apply a bias voltage, controls the gate gating unit to select a specified row, and controls the multi-channel parallel signal readout chain to acquire digital current signals after selection. During the weight update phase, it controls the gate gating unit to select a specified row and controls the voltage drive combination to apply programming pulses.

[0014] This application provides an online training system for neuromorphic hardware arrays, mainly comprising a chip carrier board and a test motherboard. The separate design of the chip carrier board and the test motherboard is fundamental to the entire system. The chip carrier board, as the physical carrier of the neuromorphic hardware array under test, plays a crucial role in physically extracting high-density signals. Neuromorphic hardware arrays typically contain hundreds or thousands of cross-points, each corresponding to a memory cell. Each memory cell requires connection to three electrodes: drain, gate, and source. The drain is used to apply a voltage signal, the gate is used to control row selection, and the source is used to output a current signal. Since different models or sizes of neuromorphic hardware arrays differ in packaging and electrode layout, integrating all test circuits onto a single board would require redesigning the entire circuit board each time the neuromorphic hardware array under test is replaced, resulting in extremely high costs. Therefore, this system adopts a modular architecture with a separate chip carrier board and test motherboard. The chip carrier board only contains electrode lead-out interfaces and connection cables, without any active devices. This allows for testing of different specifications of neuromorphic hardware arrays by simply replacing the lower-cost chip carrier board, while the expensive test motherboard can be reused. The chip carrier board is connected to the test motherboard via a high-density flexible cable, which transmits all drain, gate, and source signals of the neuromorphic hardware array to the test motherboard in their entirety, providing a physical path for subsequent driving and readout operations.

[0015] The parallel, independent programming capability of the voltage-driven assembly is core to achieving physical-level forward computation and weight updates. This assembly consists of multiple multi-channel digital-to-analog converter (DAC) chips, each DAC channel corresponding one-to-one with a column of drains in the neuromorphic hardware array under test. During the forward computation phase of the neural network, the field-programmable gate array (FPGA) control unit generates an input voltage vector based on training samples, where each element corresponds to the required bias voltage for a column of drains. The FPGA simultaneously writes these voltage values ​​into the registers of all DAC channels via a high-speed parallel bus, triggering all DACs to output voltages simultaneously. Since all column drains simultaneously receive their independent bias voltages, each column of memory cells in the neuromorphic hardware array under test begins to output current according to Ohm's law; that is, the current output by each memory cell is equal to the bias voltage applied to that column multiplied by the current conductance value (i.e., weight) of that memory cell. This fully parallel operation allows the entire neuromorphic hardware array under test to complete a full vector-matrix multiplication operation within one clock cycle, which is unmatched by traditional serial scanning architectures. During the weight update phase, the role of the voltage-driven assembly is transformed into applying programming pulses. At this point, the FPGA updates the weight parameters calculated by the training algorithm, writes the corresponding amplitude and width programming pulse parameters into each DAC channel, and triggers all DACs to output simultaneously, thereby achieving synchronous adjustment of the conductance values ​​of each column of the storage cells in the neuromorphic hardware array under test, and completing the physical update of the weights. In a specific embodiment, the voltage drive combination may include at least 32 16-channel digital-to-analog converter chips, providing no less than 512 independent analog voltage output channels, with an output voltage range of no less than ±15V and a resolution of no less than 0.1mV; it also includes a high-speed serial peripheral interface bus for concurrent control of all digital-to-analog converter chips by the field-programmable gate array control unit, realizing synchronous update of the voltage of all output channels.

[0016] The row gating function of the gate gating unit ensures the accuracy of each forward computation or weight update operation. In the neuromorphic hardware array under test, memory cells are arranged in rows and columns, with all memory cells in each row sharing the same gate line. When an enable voltage is applied to the gate of a specified row, all memory cells in that row are activated, allowing their source current to be output; memory cells in unselected rows remain off and do not affect the output. The gate gating unit is connected to the gate of each row of the neuromorphic hardware array under test. Under the control of the FPGA, the address of the row to be activated is converted into a gating signal, and the gating voltage generated by the DAC is applied to the gate of the specified row. During the forward computation phase, the FPGA can gating the row to be computed, causing the memory cells in that row to output current according to the bias voltages applied to each column. These currents are then sent to a multi-channel parallel signal readout chain for acquisition. During the weight update phase, the FPGA also needs to gating the row whose weights need to be updated, allowing the programming pulse to act on the memory cells in that row, changing their conductance values. Since forward computation and weight updates are usually required row by row during training, the fast and accurate gating capability of the gate gating unit directly affects the efficiency of the entire online training.

[0017] The multi-channel parallel signal readout chain achieves efficient conversion from analog current to digital signal, providing a data foundation for FPGA error calculation. Specifically, when a memory cell in a specified row is activated, the bias voltage applied to each column and the conductance of the memory cell in that row jointly determine the analog current signal output by each memory cell. These analog current signals enter the multi-channel parallel signal readout chain through the sources of each row, and the number is the same as the number of rows of the neuromorphic hardware array under test (e.g., 512 rows). If a separate readout circuit were configured for each channel, the hardware cost would be extremely high. Therefore, this system adopts a "grouping and convergence, parallel conversion" strategy. In a specific embodiment, each parallel readout channel of the multi-channel parallel signal readout chain integrates a four-level adaptive range transimpedance amplifier circuit; this circuit automatically switches between four preset current ranges through a low-leakage analog switch, covering a measurement range from ±500pA (resolution ≤1fA) to ±0.5mA, to achieve single-shot full-coverage accurate measurement of the device from femtoampere-level dark current to milliampere-level conduction current.

[0018] The Field-Programmable Gate Array (FPGA) control unit acts as the brain of the entire system, coordinating the collaborative work of various modules through phased synchronous timing to form a complete online training closed loop. Specifically, the FPGA is connected to the voltage drive combination, gate gating unit, and multi-channel parallel signal readout chain, responsible for generating synchronous timing signals with nanosecond-level precision. During training, the FPGA clearly divides the operation into two stages: the forward computation stage and the weight update stage. In the forward computation stage, the FPGA first controls the voltage drive combination to simultaneously apply the bias voltage corresponding to the training sample to all column drains, and simultaneously controls the gate gating unit to select the specified row gate, activating the memory cell of that row. After the row is selected, the FPGA immediately controls the multi-channel parallel signal readout chain to start acquisition, converting the analog current signal output of that row into a digital current signal. The acquired digital current signal serves as the actual output, which is used by the training algorithm inside the FPGA (running in the processing system) to compare with the expected output, calculate the error gradient, and determine the weight update parameters for each memory cell. Upon entering the weight update phase, the FPGA first controls the gate selection unit to reselect the rows whose weights need updating. Then, it controls the voltage drive combination to simultaneously apply corresponding programming pulses to the drains of each column according to the previously calculated weight update parameters. The purpose of the programming pulses is to change the conductance value of the memory cells, i.e., update the weights of the neural network. After one weight update is completed, the FPGA returns to the forward computation phase and continues to repeat the above process. This iterative cycle continues until the training of the entire array meets the preset convergence condition, ultimately resulting in a trained neuromorphic hardware array.

[0019] By applying the technical solution of this embodiment, efficient online training of the neuromorphic hardware array under test is achieved through a fully parallel bidirectional drive architecture: a high-density fully parallel voltage drive combination composed of multiple multi-channel digital-to-analog converter chips is adopted, with each output channel connected one-to-one with a column drain of the neuromorphic hardware array under test. This allows for the simultaneous application of independently programmable bias voltages or programming pulses to all column drains, thereby directly supporting parallel combined multiplication operations and synchronous weight updates at the physical level, completely overcoming the limitations of traditional serial scanning on column-level drive capabilities. Simultaneously, a field-programmable gate array control unit is used to generate nanosecond-level precision synchronization timing, strictly coordinating voltage drive, row gating, and current acquisition, and establishing a "gating-after-acquisition" timing relationship during the forward computation stage. This ensures that the excitation and response of each memory unit of the neuromorphic hardware array under test remain highly consistent on the time axis, meeting the stringent requirements of spiking neural networks for nanosecond-level synchronization. Furthermore, a multi-channel parallel signal readout chain groups and aggregates the current signals output from each row source and converts them into digital current signals in parallel, strictly guaranteeing the acquisition speed of the digital current signals.

[0020] Optionally, in this embodiment, the multi-channel parallel signal readout chain includes: a first-stage multiplexer group comprising multiple multiplexers, each multiplexer having its input connected to multiple rows of sources of the neuromorphic hardware array under test; a second-stage multiplexing network having its input connected to the outputs of each of the first-stage multiplexer group, used to group and converge source signals into multiple analog buses; multiple transimpedance amplifiers, each transimpedance amplifier having its input connected to one of the analog buses, used to convert analog current signals on the corresponding analog bus into analog voltage signals; and multiple analog-to-digital converters, each analog-to-digital converter having its input connected to the output of one of the transimpedance amplifiers, used to convert the analog voltage signals output by each transimpedance amplifier into digital current signals in parallel.

[0021] In this embodiment, the first-stage multiplexer group serves as the entry point for the entire multi-channel parallel signal readout chain, undertaking the task of initially compressing massive amounts of source signals. The neuromorphic hardware array under test typically contains hundreds or thousands of rows of sources, each capable of outputting an analog current signal upon activation. If a separate readout circuit were configured for each path, the hardware cost would be extremely high, and the wiring complexity would increase dramatically. Therefore, this embodiment introduces a first-stage multiplexer group containing multiple multiplexers (MUXs), each MUX's input connected to multiple rows of sources simultaneously. A multiplexer is essentially an array of electronic switches that, under control, can time-division multiplex one of the multiple input signals to its output. For example, a 16:1 MUX can connect to 16 rows of sources and output these 16 rows of signals sequentially over 16 time slices. Through this stage of processing, the original hundreds or thousands of source signals are initially aggregated into a smaller number of signals, laying the foundation for subsequent grouping processing and significantly reducing the number of channels in subsequent circuits.

[0022] The second-stage multiplexing network further groups and busifies the signals based on the primary aggregation, forming several parallel analog buses. Specifically, the number of signals output by the first-stage multiplexer group is still relatively large, and these signals are independent, not yet forming a structure conducive to parallel acquisition. The second-stage multiplexing network receives the output signals from each multiplexer group and reorganizes these signals according to a preset grouping strategy through another layer of multiplexing switches. For example, signals from different multiplexer groups can be combined according to certain rules to form several independent analog buses. Each analog bus aggregates current signals from multiple rows of sources, but only the current signal from one row of sources is selected for output at any given time. This two-stage structure of "primary aggregation followed by grouping and busing" allows the system to flexibly map the source signals of a large array onto a small number of parallel readout channels with limited hardware resources, creating conditions for subsequent multi-channel parallel acquisition. Each analog bus is essentially a physical transmission line carrying the currently selected current signal.

[0023] Transimpedance amplifiers act as a bridge between analog current and analog voltage signals, converting analog current into analog voltage and providing a foundation for wide dynamic range measurements. Specifically, the signal on each analog bus is essentially an analog current, while subsequent analog-to-digital converters (ADCs) typically require an analog voltage input to function properly; therefore, current-to-voltage conversion is essential. The transimpedance amplifier is the core device for this conversion. Its basic operating principle is to generate an output voltage by passing the input current through a feedback resistor. The output voltage is proportional to the input current, and the proportionality coefficient is determined by the resistance value of the feedback resistor. In this embodiment, each transimpedance amplifier's input terminal is connected one-to-one with one analog bus, so the number of transimpedance amplifiers operating simultaneously corresponds to the number of analog buses. This one-to-one connection ensures that the analog current signal on each analog bus can be converted independently and in real-time, avoiding crosstalk between channels. More importantly, the gain (i.e., conversion coefficient) of the transimpedance amplifier can be changed by switching the feedback resistor, providing a hardware foundation for subsequent automatic range switching and wide dynamic range measurements covering picoampere to milliampere levels.

[0024] After conversion by the transimpedance amplifier, the current information on each analog bus is presented in the form of voltage. However, these voltages are still analog quantities and cannot be directly processed by digital logic circuits. The role of the analog-to-digital converter (ADC) is to quantize the continuously changing analog voltages into discrete digital values. In this embodiment, the input terminal of each ADC is connected one-to-one with the output terminal of a transimpedance amplifier, forming multiple completely independent parallel readout channels. When the field-programmable gate array (FPGA) control unit issues a data acquisition command, all ADCs can start conversion simultaneously, thereby achieving parallel data acquisition. This parallel operation mode enables the system to simultaneously acquire the current information on all analog buses within a very short time window and convert it into digital current signals. These digital current signals are the actual output results of the neural network's forward computation. They are transmitted to the FPGA control unit in real time for subsequent error calculation and weight update decisions, thus forming a complete online training closed loop.

[0025] This application embodiment utilizes a two-stage aggregation structure consisting of a first-stage multiplexer group and a second-stage multiplexing network to group and aggregate hundreds or thousands of analog current signals from the source output into a smaller number of analog buses. This significantly reduces the number of subsequent signal processing channels, substantially lowering hardware costs and system wiring complexity. Furthermore, by connecting multiple transimpedance amplifiers to multiple analog buses in a one-to-one correspondence and connecting multiple analog-to-digital converters to multiple transimpedance amplifiers in a one-to-one correspondence, parallel conversion of multiple analog signals is achieved. This avoids the timing bottleneck of traditional single-channel serial acquisition and effectively improves the array's scan frame rate.

[0026] Optionally, in this embodiment of the application, each transimpedance amplifier automatically selects the target current range from multiple preset current ranges based on the magnitude of the analog current on the corresponding analog bus using an analog switch.

[0027] In this embodiment, each transimpedance amplifier is connected to an analog bus in a one-to-one correspondence, enabling it to independently select its range based on the actual current conditions of the analog bus it is responsible for, avoiding measurement distortion caused by a "one-size-fits-all" range setting. Specifically, before each formal acquisition, the transimpedance amplifier assesses the current strength on the analog bus using a predicted value and compares the predicted value with the threshold values ​​of multiple preset current ranges to determine the target current range most suitable for the current signal strength. Subsequently, the transimpedance amplifier controls the conduction state of the analog switch to switch the internal feedback resistor network to the resistance value corresponding to the target current range, thereby completing the gain configuration. After the above measurement, comparison, and switching, the transimpedance amplifier converts the analog current signal into an analog voltage signal at the matched range, allowing weak currents to be amplified to a sufficient amplitude with high gain, while large currents are controlled within a safe range with low gain. This achieves a wide dynamic range single-shot full-coverage measurement from picoamperes to milliamperes without sacrificing accuracy.

[0028] Optionally, in this embodiment, the field-programmable gate array control unit is a system-on-a-chip architecture integrating a programmable logic section and a processing system section. The programmable logic section is connected to the voltage drive assembly, the gate gating unit, and the multi-channel parallel signal readout chain, and is used to generate synchronous timing to control the voltage drive assembly to apply a bias voltage or programming pulse, control the gate gating unit to select a specified row, and control the multi-channel parallel signal readout chain to collect digital current signals after selection. The processing system section is connected to the programmable logic section and is used to run a training algorithm, calculate weight update parameters based on the digital current signals collected by the multi-channel parallel signal readout chain, and calculate programming pulse parameters corresponding to each column drain based on the weight update parameters. The programmable logic section controls the voltage drive assembly to apply corresponding programming pulses to the corresponding column drains.

[0029] In this embodiment, the field-programmable gate array (FPGA) control unit adopts a system-on-a-chip (SoC) architecture that integrates the programmable logic section and the processing system section. This heterogeneous design fundamentally solves the contradiction between real-time hardware control and complex algorithm computation. The programmable logic section is essentially a logic array that can be customized by the user. It can execute multiple tasks in parallel in a purely hardware manner with a response speed of nanoseconds, making it very suitable for generating precise timing control signals. The processing system section is equivalent to a processor core embedded in the chip (such as an ARM architecture CPU), which is good at running complex software algorithms, such as backpropagation training algorithms. Integrating the two on the same chip avoids the board-level communication latency caused by the dual-chip architecture of "FPGA + independent processor" in traditional solutions, and enables hardware control and algorithm computation to directly exchange data through the internal high-speed bus, laying a physical foundation for the subsequent realization of low-latency online training closed loop.

[0030] The programmable logic unit (PLU) is responsible for generating nanosecond-level precision synchronization timing, directly driving the voltage drive combination, gate gating unit, and signal readout chain to work together. Because the forward computation and weight updates of neuromorphic hardware have extremely high timing precision requirements—for example, in a spiking neural network, a pulse arrival time deviation exceeding 100 nanoseconds can lead to training failure—timing control must be performed by hardware logic rather than software. The PLU is directly connected to these three core hardware modules. Based on the timing state machine configured within the field-programmable gate array (FPGA) control unit, it precisely determines when to trigger the voltage drive combination output voltage, when to gating the gate of a specified row, and how many nanoseconds to delay before initiating the multi-channel parallel signal readout chain's acquisition operation. This "gating-before-acquisition" timing relationship is particularly critical because only after the row is activated will the source have a stable current output; premature acquisition will result in invalid data, while excessive delay will reduce training efficiency. The PLU implements this precise timing control through a hardware counter, ensuring that each forward computation or weight update is completed within the correct time window.

[0031] The processing system acts as the algorithmic brain of the system, running the training algorithm and calculating weight update parameters. Simultaneously, it indirectly controls the hardware to execute updates through the programmable logic unit, forming a complete online training closed loop. After the programmable logic unit completes a forward calculation and acquires digital current signals, these signals can be directly transmitted to the memory of the processing system unit via the chip's internal high-speed bus, without needing an external interface, thus significantly reducing data transmission latency. The processing system unit runs a lightweight neural network training algorithm. It first uses the acquired digital current signal as the actual output of the current training sample, compares it with the pre-stored expected output, and calculates the error. Then, through backpropagation or other training rules, it calculates the error gradient corresponding to each memory cell in the tested neuromorphic hardware array, thereby determining the weight update parameters (i.e., conductance changes) that need to be adjusted for each memory cell. Based on this, the processing system unit converts these weight update parameters into programmable pulse parameters (including pulse amplitude, width, polarity, etc.) that the voltage-driven combination can recognize, as these parameters determine the direction and magnitude of the memristor conductance change. Finally, the processing system transmits these programming pulse parameters to the programmable logic unit via an internal bus. The programmable logic unit then generates precise synchronization timing and controls the voltage drive combination to simultaneously output these programming pulses at the corresponding column drains. Simultaneously, it controls the gate gating unit to select the row that needs updating, thus completing a full weight update. After one iteration, the processing system can immediately begin the next iteration, repeating this cycle to achieve independent online training completely independent of the host computer.

[0032] This application's embodiments achieve a deep integration of real-time hardware control and complex algorithm computation by designing the field-programmable gate array (FPGA) control unit as an on-chip architecture integrating programmable logic and processing systems. The programmable logic generates synchronous timing sequences with nanosecond precision, directly driving the voltage drive combination, gate gating unit, and signal readout chain to work collaboratively, ensuring a strict "gating-after-acquisition" timing relationship during forward computation and weight updates. The processing system runs the training algorithm, directly acquiring the acquired digital current signal via an internal high-speed bus, performing error calculation and weight update parameter generation locally, and having the programmable logic execute the update in real time. This heterogeneous architecture completely eliminates the board-level communication delay between the FPGA and the external processor in traditional solutions, constructing a microsecond-level hardware closed loop from signal acquisition to weight updates. This enables independent online training completely independent of the host computer, significantly improving the training efficiency and real-time response capability of neuromorphic hardware.

[0033] Optionally, in this embodiment, the gate selection unit includes a digital-to-analog converter, a digital decoder, and a high-voltage multiplexer; the digital-to-analog converter is used to generate a gating voltage or pulse under the control of the field-programmable gate array control unit; the digital decoder is used to control the high-voltage multiplexer to apply the gating voltage or pulse to one or more rows of gates of the neuromorphic hardware array under test under the control of the field-programmable gate array control unit.

[0034] In this embodiment, the gate gating unit consists of three core components: a digital-to-analog converter (DAC), a digital decoder, and a high-voltage multiplexer. These three components work together to convert and distribute the FPGA control commands into actual gate gating signals. Specifically, the DAC generates programmable gating voltages or pulse waveforms, the decoder converts the row address commands from the FPGA into gating control signals, and the high-voltage multiplexer acts as a switching matrix at the execution end, accurately outputting the gating signals to the target row gate. This collaborative structure allows the gate gating unit to flexibly adapt to the requirements of different training stages for gating signal forms (simple DC voltages or complex pulse sequences), and enables fast and accurate gating of arbitrary rows in large-scale arrays, providing a reliable row-level control foundation for subsequent forward computation and weight updates.

[0035] Specifically, the digital-to-analog converter (DAC) acts as a signal source in the entire gate gating unit, generating gating voltages or pulses of arbitrary waveforms and amplitudes according to the FPGA's control instructions. Specifically, during the forward computation phase, the DAC outputs a stable DC gating voltage to activate the memory cells in a specified row, enabling them to respond to the bias voltage applied to the column and output current. During the weight update phase, the DAC outputs programming pulses with specific amplitude, width, and polarity. The precise parameters of these pulses directly affect the direction and amplitude of the memory cell conductance changes, which is crucial for achieving accurate weight updates.

[0036] The digital decoder and the high-voltage multiplexer work together to assign addresses and physically output gating signals, ensuring that the rows specified by the FPGA are accurately selected and the correct gating signals are applied. Specifically, the digital decoder receives the binary encoding of the row address sent by the FPGA and decodes it into the corresponding gating control signal for the high-voltage multiplexer. For example, when the 5th row needs to be selected, the digital decoder converts the address signal into a control code that closes the 5th channel of the high-voltage multiplexer. The high-voltage multiplexer is a gating network composed of a high-voltage analog switch array. Its input is connected to the output of the digital-to-analog converter (DAC), and its outputs are connected to the gates of each row of the array. When the control signal output by the digital decoder turns on a channel within the high-voltage multiplexer, the gating voltage or pulse generated by the DAC is applied to the corresponding row gate through that channel. Because the high-voltage multiplexer is designed to withstand high voltage, it can withstand the high-voltage pulses required for memristor programming, ensuring the reliability of signal transmission during the weight update phase. By combining a digital decoder with a high-voltage multiplexer, the FPGA can achieve arbitrary single or multiple row selection of hundreds or thousands of gates by simply outputting a simple row address signal. This simplifies the control interface and ensures the selection speed and accuracy in large-scale arrays.

[0037] In one specific embodiment, the online training system for neuromorphic hardware arrays supports dual-modal operation to balance visualization debugging and high-speed training: in the first mode, the collected data is uploaded to the host computer, where the host computer software performs array response visualization imaging and runs a complex deep learning framework for algorithm training; in the second mode, the training algorithm is directly embedded and runs in the processing system of the field-programmable gate array control unit to achieve low-latency hardware-in-the-loop training.

[0038] Furthermore, as Figure 1 In terms of specific system implementation, this application provides an online training method for neuromorphic hardware arrays, such as... Figure 2 As shown, the method includes: Step 101: The field-programmable gate array control unit generates an input voltage vector based on the training samples, controls the voltage drive combination to write the voltage corresponding to each element in the input voltage vector into the register of each corresponding digital-to-analog converter channel in the voltage drive combination through a parallel bus, and triggers all digital-to-analog converter channels to output voltage simultaneously, so as to apply the bias voltage corresponding to the input voltage vector to each column drain of the neuromorphic hardware array under test.

[0039] Step 102: The field-programmable gate array control unit controls the gate selection unit to select the gate of a specified row of the neuromorphic hardware array under test, so as to activate the memory cell of the corresponding row, and so that each activated memory cell outputs an analog current signal based on Ohm's law according to the bias voltage applied to its column and its own conductance value.

[0040] Step 103: After the gate of the specified row is selected, the field programmable gate array control unit triggers multiple analog-to-digital converters in the multi-channel parallel signal readout chain to perform analog-to-digital conversion operations in parallel.

[0041] Step 104: The multi-channel parallel signal readout chain groups and aggregates the analog current signals output by each memory cell in the specified row gate into multiple analog buses, and converts the analog current signals of each analog bus into digital current signals in parallel, which are used as the actual output of the training samples.

[0042] Step 105: The field-programmable gate array control unit compares the actual output with the expected output of the training sample, calculates the error gradient corresponding to each storage unit in the neuromorphic hardware array under test based on the comparison result, and determines the weight update parameters of each storage unit according to the error gradient.

[0043] Step 106: The field-programmable gate array control unit updates the parameters according to the weights, writes the corresponding programming pulse parameters into the registers of the corresponding digital-to-analog converter channels in the voltage drive combination via a parallel bus, and triggers all digital-to-analog converter channels to output programming pulses simultaneously; at the same time, it controls the gate selection unit to select the gate of the specified row, so that the programming pulse is applied to the memory cell of the corresponding row to update the conductance value of each memory cell, and returns to the step of the field-programmable gate array control unit controlling the gate selection unit to select the gate of the specified row of the neuromorphic hardware array under test, until the training of the neuromorphic hardware array under test meets the preset convergence condition, and the trained neuromorphic hardware array under test is obtained.

[0044] This application provides an online training method for neuromorphic hardware arrays. First, the field-programmable gate array (FPGA) control unit (FPGA) generates an input voltage vector with the same column dimension as the neuromorphic hardware array under test, based on the numerical characteristics of the current training samples. Each element in this vector corresponds to the bias voltage value required for a column of drains. Since driving the entire array requires applying bias voltages to all columns simultaneously, the FPGA control unit utilizes the high bandwidth of the parallel bus to write these voltage values ​​simultaneously into the registers of each digital-to-analog converter (DAC) channel in the voltage driving combination. After writing, the FPGA control unit issues a synchronization trigger signal, causing all DAC channels to start digital-to-analog conversion and output voltage at the same time, thereby simultaneously applying bias voltages corresponding one-to-one with the input voltage vector to each column of drains in the array. This fully parallel writing and triggering mechanism ensures that the input signal can act on the entire array within the same time window, creating the necessary conditions for subsequent parallel multiplication operations.

[0045] Once the bias voltage is stably applied to the drains of each column, the field-programmable gate array (FPGA) control unit then controls the gate selection unit to select the gate of the specified row. The gate selection unit acts as a row selection switch, simultaneously turning on all memory cells in the selected row. At this time, the bias voltage on each column acts on the corresponding memory cell in that row. According to Ohm's law, the current output by each memory cell is exactly equal to the column bias voltage multiplied by the current conductance of that memory cell. Since all column voltages are applied simultaneously, all row memory cells are activated simultaneously. The entire array can complete a full vector-matrix multiplication operation—the product of the input voltage vector and the weight matrix—within one time period. The result is presented as an analog current signal at the source of each row for subsequent readout circuitry to acquire.

[0046] After issuing a row gating command, the field-programmable gate array (FPGA) control unit can precisely wait for an extremely short time window. This window is long enough for the row gating signal to stabilize and for the memory cell output current to reach a steady state before immediately triggering multiple analog-to-digital converters (ADCs) in the multi-channel parallel signal readout chain to initiate parallel analog-to-digital conversion. This "gating-after-triggering" timing relationship is crucial: if triggered too early, the row is not yet fully activated, and the acquired current signal will be too small or even zero; if triggered too late, training time will be wasted, reducing overall efficiency. The FPGA control unit generates nanosecond-level precision synchronization timing through its internal programmable logic, ensuring that the acquisition command is issued precisely at the optimal moment after the row gating signal has stabilized, so that the results of each forward calculation can be accurately captured.

[0047] When the analog current signals output from each memory cell in a specified row enter the multi-channel parallel signal readout chain, they are first aggregated by a first-stage multiplexer group, merging the analog current signals from multiple source cells into a smaller number of signals. Then, a second-stage multiplexing network further integrates these initially aggregated signals into several analog buses according to a preset grouping strategy, with each analog bus corresponding to an independent readout channel. In each readout channel, a transimpedance amplifier converts the analog current signal on the analog bus into an analog voltage signal, and then a corresponding analog-to-digital converter converts the analog voltage signal into a digital current signal in parallel. This "grouping and aggregation first, then parallel conversion" architecture allows the system to acquire current information from all analog buses simultaneously with limited hardware resources, avoiding the speed bottleneck of traditional serial acquisition. The final output digital current signal is the actual output value of the current training sample.

[0048] The Field Programmable Gate Array (FPGA) control unit compares the acquired actual output with the pre-stored expected output of the training sample, obtaining the error value through subtraction. Subsequently, the backpropagation algorithm running inside the FPGA control unit propagates the error layer by layer from the output layer back to each memory cell in the array, calculating the contribution of each memory cell to the final error, i.e., the error gradient. The sign of the error gradient determines whether the conductance of that memory cell should increase or decrease, while the magnitude of the error gradient determines the adjustment range. Based on this, the algorithm converts these error gradients into specific weight update parameters, typically represented by the amplitude, width, and polarity of the programming pulse. These parameters directly guide subsequent physical weight update operations.

[0049] Furthermore, the field-programmable gate array (FPGA) control unit updates the calculated weight parameters and writes the corresponding programming pulse parameters into the digital-to-analog converter (DAC) channel registers corresponding to the drains of each column via a parallel bus, triggering all DAC channels to simultaneously output programming pulses. Unlike bias voltage, programming pulses are short-duration voltage signals with specific amplitude and width, capable of changing the conductivity state inside the memory cells, thereby achieving precise adjustment of the conductance value. Simultaneously with outputting programming pulses, the FPGA control unit again controls the gate selection unit to select the same row gate as the one used in the forward calculation, allowing the programming pulses to act on that row of memory cells, completing the physical update of the weights. After one weight update is completed, the FPGA control unit automatically returns to the step of selecting the specified row gate and begins the next round of forward calculation and weight update, iterating repeatedly. When the recognition accuracy of the entire array for all training samples reaches a preset threshold, or the error decreases to an acceptable range, the training process terminates, and the resulting neuromorphic hardware array under test possesses the ability to complete the target task.

[0050] In the embodiments of this application, optionally, as shown... Figure 3 As shown, step 104 includes: Step 104-1: The analog current signals output by each memory cell in the specified row gate are initially aggregated by the first-stage multiplexer group.

[0051] Step 104-2: The analog current signals after primary aggregation are grouped and aggregated into multiple analog buses through the second-level multiplexing network.

[0052] Step 104-3: Through multiple transimpedance amplifiers connected one-to-one with the multiple analog buses, the analog current signals on each analog bus are simultaneously converted into analog voltage signals.

[0053] Step 104-4: Through multiple analog-to-digital converters connected one-to-one with the multiple transimpedance amplifiers, the analog voltage signals are simultaneously and in parallel converted into digital current signals.

[0054] Optionally, for each transimpedance amplifier, step 104-3, "simultaneously converting the analog current signals on each analog bus into analog voltage signals," includes: measuring the magnitude of the analog current on the corresponding analog bus, comparing the measured analog current value with the threshold values ​​of multiple preset current ranges, and selecting a matching target current range based on the comparison result; switching its gain configuration to the gain range corresponding to the target current range by controlling the conduction state of the analog switch; and converting the analog current signals on the analog bus into analog voltage signals under the switched gain range.

[0055] In this embodiment, each transimpedance amplifier is dedicated to one analog bus, rapidly measuring the analog current signal on the bus via its internal sampling circuitry. This measurement process is typically extremely short, completed within nanoseconds before the actual conversion. The transimpedance amplifier internally stores threshold values ​​for multiple preset current ranges; for example, the smallest range corresponds to 0 to 500 picoamps, the middle range to 500 picoamps to 50 nanoamps, and the largest range to microamps to milliamps. It compares the measured current value with these threshold values ​​one by one to determine which range's coverage area the current signal falls within, thus selecting the most suitable target current range for the signal strength. Specifically, the selection is based on the smallest range that can cover the current current value, as a smaller range provides higher measurement accuracy, but it must be ensured that the upper limit of that range is not exceeded, which could lead to signal saturation.

[0056] After selecting the target current range, the transimpedance amplifier controls an analog switch to change the connection state of its internal feedback resistor network, thus physically switching the gain. The gain of the transimpedance amplifier is determined by the resistance value of its feedback resistor. The larger the feedback resistor, the higher the output voltage generated by the same input current, i.e., the higher the gain, suitable for measuring weak currents; the smaller the feedback resistor, the lower the gain, suitable for measuring large currents. To enable switching between multiple ranges, the transimpedance amplifier has multiple feedback resistors with different resistance values ​​pre-set inside, each corresponding to a range. These resistors are connected to the core circuit of the transimpedance amplifier via analog switches. An analog switch is an electronic switching element that can be controlled to turn on or off by an electrical signal. When the transimpedance amplifier determines the target current range, it outputs a corresponding control signal to turn on the analog switch corresponding to the target current range, connecting the feedback resistor of that range to the circuit, while ensuring that the analog switches of other ranges remain off, disconnecting other resistors. The entire switching process is completed automatically by hardware, with an extremely fast switching speed, completing gain configuration within microseconds.

[0057] At this point, the transimpedance amplifier's internal circuitry is configured with the gain state best suited to the current signal strength. It takes the analog current signal from the analog bus as input, passes through the feedback resistor in the connection circuit, and outputs the corresponding analog voltage signal. Since the target current range is pre-matched based on the actual signal strength, weak currents can be amplified to a sufficient amplitude by high gain to ensure that the subsequent analog-to-digital converter can distinguish minute changes; large currents are controlled within a safe range by low gain to avoid signal saturation or damage to subsequent circuits. Through this three-step process of "measuring and comparing first, then switching the gain, and finally converting the output," the transimpedance amplifier can automatically adapt to a wide dynamic range signal from picoamperes to milliamperes in a single measurement, without manual intervention or multiple measurements. This allows for precise capture of the current signal across the entire operating range of the storage unit while ensuring accuracy.

[0058] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, and to fully illustrate the specific implementation process of this embodiment, another online training method for neuromorphic hardware arrays is provided, which includes: Taking the training of a 512×512 memristor array to recognize the handwritten digit "8" as an example, we assume that each intersection of this array is a memristor (memory cell), and its conductance value represents the weight of the neural network.

[0059] I. Forward computation phase: First, the image of the number "8" is transformed into an input voltage vector. Assume that the image of "8," after preprocessing, is mapped into a 512-dimensional input voltage vector, for example, [0.1V, 0.2V, …, 2.5V, …]. Each element of this vector corresponds to the voltage value that needs to be applied to the drain electrodes of a column in the array.

[0060] The FPGA (Field-Programmable Gate Array) first generates control commands based on this input voltage vector, and then writes these 512 voltage values ​​into the registers of each of the 512 DAC (Digital-to-Analog Converter) channels in the voltage drive combination via a parallel bus. After writing, the FPGA sends a synchronization trigger signal, and all DAC channels start working simultaneously, converting the digital voltage values ​​into analog voltages and outputting them simultaneously to their respective column drains. At this point, each of the 512 columns of the entire array has its own independent bias voltage applied.

[0061] Next, the specified row is selected to complete the multiplication operation. The FPGA then controls the gate selection unit to select the gate of the first row (assuming the first row is trained first). The digital decoder in the gate selection unit converts the row address into a control signal, which drives the high-voltage multiplexer to apply the selection voltage generated by the DAC to the gate of the first row, activating the 512 memristors in that row.

[0062] Each memristor is currently performing a simple physical calculation: Output current = column voltage × memristor conductance (Ohm's Law). For example, if the voltage of the first column is 0.1V, and the current conductance of the memristor in the first column of the first row is 0.5μS (micro Siemens), then it outputs a current of 50nA. Thus, 512 memristors in one row simultaneously output 512 current signals, which flow from their respective sources and converge on the source line of the first row.

[0063] Furthermore, parallel acquisition converts the current into a digital signal. After the row selection signal stabilizes, the FPGA immediately triggers multiple ADCs (analog-to-digital converters) in the multi-channel parallel signal readout chain to start the conversion in parallel.

[0064] The 512 analog current signals collected on the source lines first enter the first-stage multiplexer group. Each 16:1 MUX performs primary aggregation of the currents from the 16 rows of sources, forming 32 signals. Next, the second-stage multiplexing network further groups and aggregates these signals into four analog buses. These four buses are connected to four transimpedance amplifiers (TIAs), each converting the analog current signals on its bus into analog voltage signals. Since the current signals can be very weak (pA level), the TIAs can automatically switch ranges based on the current magnitude. If the current is only tens of picoamps, it can switch to a high-gain setting, amplifying the signal to a range that the ADC can recognize. Finally, the four ADCs (each corresponding to a TIA) simultaneously convert the analog voltage signals into digital current signals. Thus, one row gating operation yields 512 digitized current values, which represent the neural network's first guess at the number "8".

[0065] II. Error Calculation and Weight Update: The FPGA uses the 512 acquired digital current signals as the "actual output" and compares them with the "expected output." For the number "8," the expected output might be a vector where only the 8th output terminal has a high current, while the others have low current. The FPGA subtracts the two to obtain the error vector.

[0066] The processing unit (PS) inside the FPGA runs a backpropagation algorithm, propagating this error back to the entire network. It calculates the contribution of each memristor to the final error (i.e., the error gradient) and determines, based on the gradient, whether the conductance of each memristor should be increased or decreased, and by how much. This information is ultimately converted into programming pulse parameters. For example, for a memristor that needs increased conductance, a +3V positive pulse lasting 10μs might be applied; for one that needs decreased conductance, a -2V negative pulse lasting 5μs might be applied.

[0067] Based on these pulse parameters, the FPGA again writes the programming pulse parameters into the registers of all DAC channels via the parallel bus, triggering all DACs to simultaneously output these programming pulses. Simultaneously, the FPGA again controls the gate selection unit to select the first row of gates (the same row as the one calculated previously), applying the programming pulses to the memristors in the first row. These pulses change the internal physical state of the memristors, precisely increasing or decreasing their conductance, thus achieving a physical update of the weights.

[0068] At this point, the weight of the first row has been updated.

[0069] III. Iterative Loops and Final Convergence: After updating the first row, the FPGA will automatically return and begin processing the second row, the third row, and so on, until all 512 rows have completed one forward computation and weight update. This constitutes one complete training cycle.

[0070] Then, the FPGA repeats all the steps from "applying voltage" to "updating weights" using the image of the number "8" again (or using the next training sample). This process is repeated for dozens, hundreds, or even thousands of rounds.

[0071] During this process, the conductance of the memristors in the array is gradually adjusted so that when the input voltage vector of the digit "8" is input, the current vector output by the array gets closer and closer to the desired output (i.e., the current at the 8th output terminal is much greater than that at the other terminals). When the FPGA detects that the recognition accuracy has reached a preset threshold (e.g., 95%), or the error has dropped to an acceptable range, the training process terminates. At this point, the 512×512 memristor array has completed training and becomes a dedicated neural network hardware capable of recognizing the handwritten digit "8", which can be directly used for inference tasks.

[0072] It should be noted that other corresponding descriptions of the functional units involved in the online training method for neuromorphic hardware arrays provided in this application embodiment can be found in the following references. Figure 1 The corresponding descriptions in the system will not be repeated here.

[0073] This application also provides a computer device, which may specifically be a personal computer, a server, a network device, etc. Figure 4 As shown, the computer device includes a bus, a processor, memory, and a communication interface, and may also include an input / output interface and a display device. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores location information. The network interface allows communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0074] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0075] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0076] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0077] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0078] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0079] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0080] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An online training system for neuromorphic hardware arrays, characterized in that, include: A chip carrier board is used to carry the neuromorphic hardware array under test and to provide lead-out interfaces for all electrodes of the neuromorphic hardware array under test, wherein the electrodes include drain, gate and source. The test motherboard, electrically connected to the chip carrier board, includes: The voltage-driven assembly consists of multiple multi-channel digital-to-analog converter chips. Each output channel is connected one-to-one with a column of drains of the neuromorphic hardware array under test. Under the control of the field-programmable gate array control unit, it is used to simultaneously apply independently programmable bias voltages or programming pulses to each column of drains to realize forward computation or weight updates during online training. The gate selection unit is connected to each row of gates of the neuromorphic hardware array under test, and is used to select one or more rows of gates under the control of the field programmable gate array control unit to activate the memory cells of the corresponding row to participate in forward calculation or weight update. A multi-channel parallel signal readout chain, whose input is connected to each row of the source of the neuromorphic hardware array under test, is used to divide the current signals output from each row of the source into multiple groups under the control of the field-programmable gate array control unit, converge the current signals in each group into one analog signal, and convert the multiple analog signals into digital current signals in parallel, so that the field-programmable gate array control unit can calculate the prediction error and determine the weight update parameters. The field-programmable gate array control unit is connected to the voltage drive combination, the gate gating unit, and the multi-channel parallel signal readout chain, respectively, to generate synchronous timing. During the forward computation phase, it controls the voltage drive combination to apply a bias voltage, controls the gate gating unit to select a specified row, and controls the multi-channel parallel signal readout chain to acquire digital current signals after selection. During the weight update phase, it controls the gate gating unit to select a specified row and controls the voltage drive combination to apply programming pulses.

2. The system according to claim 1, characterized in that, The multi-channel parallel signal readout chain includes: The first-level multiplexer group includes multiple multiplexers, and the input of each multiplexer is connected to the multi-row source of the neuromorphic hardware array under test. The second-stage multiplexing network has its input connected to the output of each of the first-stage multiplexer groups, and is used to group and aggregate source signals into multiple analog buses. Multiple transimpedance amplifiers, each with its input terminal connected to a corresponding analog bus, are used to convert the analog current signal on the corresponding analog bus into an analog voltage signal. Multiple analog-to-digital converters, each with its input terminal connected to the output terminal of one of the transimpedance amplifiers, are used to convert the analog voltage signals output by each transimpedance amplifier into digital current signals in parallel.

3. The system according to claim 2, characterized in that, Each of the transimpedance amplifiers automatically selects the target current range from multiple preset current ranges based on the analog current magnitude on the corresponding analog bus via an analog switch.

4. The system according to claim 1, characterized in that, The field-programmable gate array control unit is a system-on-a-chip architecture that integrates programmable logic and processing systems. The programmable logic section is connected to the voltage drive assembly, the gate selection unit, and the multi-channel parallel signal readout chain to generate synchronous timing, control the voltage drive assembly to apply bias voltage or programming pulse, control the gate selection unit to select a specified row, and control the multi-channel parallel signal readout chain to acquire digital current signals after selection. The processing system is connected to the programmable logic unit and is used to run the training algorithm, calculate the weight update parameters based on the digital current signals collected by the multi-channel parallel signal readout chain, and calculate the programming pulse parameters corresponding to each column drain based on the weight update parameters. The programmable logic unit controls the voltage drive combination to apply the corresponding programming pulses to the corresponding column drains.

5. The system according to claim 1, characterized in that, The gate selection unit includes a digital-to-analog converter, a digital decoder, and a high-voltage multiplexer; The digital-to-analog converter is used to generate gating voltages or pulses under the control of the field-programmable gate array control unit; The digital decoder is used, under the control of the field-programmable gate array control unit, to control the high-voltage multiplexer to apply the gating voltage or pulse to one or more rows of gates of the neuromorphic hardware array under test.

6. An online training method for neuromorphic hardware arrays, characterized in that, include: The field-programmable gate array control unit generates an input voltage vector based on training samples. It controls the voltage drive combination to write the voltage corresponding to each element in the input voltage vector into the register of each corresponding digital-to-analog converter channel in the voltage drive combination through a parallel bus, and triggers all digital-to-analog converter channels to output voltage simultaneously, so as to apply the bias voltage corresponding to the input voltage vector to each column drain of the neuromorphic hardware array under test. The field-programmable gate array control unit controls the gate selection unit to select the gate of a specified row of the neuromorphic hardware array under test, so as to activate the memory cell of the corresponding row, and the activated memory cell outputs an analog current signal based on Ohm's law according to the bias voltage applied to the column and its own conductance value. After the designated row gate is selected, the field-programmable gate array control unit triggers multiple analog-to-digital converters in the multi-channel parallel signal readout chain to perform analog-to-digital conversion operations in parallel. The multi-channel parallel signal readout chain groups and aggregates the analog current signals output by each memory cell in the specified row gate into multiple analog buses, and converts the analog current signals of each analog bus into digital current signals in parallel, which serve as the actual output of the training samples. The field-programmable gate array control unit compares the actual output with the expected output of the training sample, calculates the error gradient corresponding to each storage unit in the neuromorphic hardware array under test based on the comparison result, and determines the weight update parameters of each storage unit according to the error gradient. The field-programmable gate array (FPGA) control unit updates the parameters according to the weights, writes the corresponding programming pulse parameters into the registers of the corresponding digital-to-analog converter (DAC) channels in the voltage drive combination via a parallel bus, and triggers all DAC channels to output programming pulses simultaneously. Simultaneously, it controls the gate selection unit to select the gate of the specified row, applying the programming pulse to the corresponding row's memory cell to update the conductance value of each memory cell. Then, it returns to the step where the FPGA control unit controls the gate selection unit to select the gate of the specified row of the neuromorphic hardware array under test, until the training of the neuromorphic hardware array under test meets the preset convergence condition, resulting in a trained neuromorphic hardware array under test.

7. The method according to claim 6, characterized in that, The multi-channel parallel signal readout chain groups and aggregates the analog current signals output by each memory cell in the specified row gate into multiple analog buses, and converts the analog current signals of each analog bus into digital current signals in parallel, including: The analog current signals output by each memory cell in the specified row gate are initially aggregated by the first-stage multiplexer group. The analog current signals, after primary aggregation, are grouped and aggregated into multiple analog buses through a second-stage multiplexing network; Multiple transimpedance amplifiers, each corresponding to one of the multiple analog buses, simultaneously convert the analog current signals on each analog bus into analog voltage signals. Multiple analog-to-digital converters, each corresponding to one of the multiple transimpedance amplifiers, simultaneously and in parallel convert each analog voltage signal into a digital current signal.

8. The method according to claim 7, characterized in that, For each transimpedance amplifier, the simultaneous conversion of analog current signals on each analog bus into analog voltage signals includes: Measure the magnitude of the analog current on the corresponding analog bus, compare the measured analog current value with the threshold of multiple preset current ranges, and select the matching target current range based on the comparison result. By controlling the conduction state of the analog switch, its own gain configuration is switched to the gain level corresponding to the target current range. At the switched gain level, the analog current signal on the analog bus is converted into an analog voltage signal.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 6 to 8.

10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 6 to 8.